According to @_avichawla, Jev-style scoring directly ranks fixed labels, cutting decoding overhead and parsing errors versus structured output and text gen.