benchmark author demo
One judge call, or twelve dimension scores?
Compares one direct Jev question per row against 12–14 Jev-scored dimensions with locally fitted weights across three classification tasks.
Notes
20,139 rows, 34.1M input tokens, 9m14s, $1.43. Decomposition helped on Japanese NLI (0.837 → 0.908) and hurt on security-filter false positives (1.5% direct vs 37.2% dimensions); dimensions cost 1.6–2.3x tokens.
author demo — Result shown by its author. About this label