According to @_avichawla, Jev-style scoring directly ranks fixed labels, cutting decoding overhead and parsing errors versus structured output and text gen.
Some results have been hidden because they may be inaccessible to you
Show inaccessible results