jevland

benchmark author demo

Jev as a reranker: graded-relevance eval

Evaluates Jev reranking against bge-m3 embedding search over an Agent Skills Hub catalog, with judge-bias and robustness analysis.

33,047 skills, 9,831 labelled pairs, 164 queries: fusing bge-m3 with Jev lifts NDCG@10 from 0.774 to 0.864; standalone Jev rerank scores 0.028 below the embedding baseline.

author demo — Result shown by its author. About this label

← Back to the directory