- axis
- methodology
- verification
- human
- whatItTests
- Open platform ranking LLMs via crowdsourced pairwise human-preference votes (over 240K votes at time of publication), using statistical ranking methods (Bradley-Terry/Elo) to produce a leaderboard.
- discriminates
[reputation-elo]
- saturation
- n-a
- patternLabFit
- Direct methodological ancestor of any ELO/pairwise-ranking pattern in Pattern Lab; validates that pairwise human votes agree with expert raters before trusting a ranking derived from votes.
- notes
- Distinct from JudgeBench (which scores judge correctness against gold labels) and MAST (failure taxonomy): this is the primary source for Bradley-Terry/Elo pairwise-ranking methodology itself, the mechanism behind reputation-elo style ranking.