benchmark source-backed
llm-rankers: checking Jev's calibration with TREC judgments
The ielab document-ranking toolkit gained a jev module, with code that uses TREC human relevance judgments to test whether Jev's probabilities are actually calibrated.
Notes
Code for the calibration check; results in the repository.
source-backed — Public repository, docs or live artifact. About this label