Everyone claims their prompt works.We publish the numbers.

Five suites. Fixed cases. Named judges.

A suite decides everything: the cases, the judge, the holdout split. Entries compete inside a suite, never across them.

Why scores instead of stars

ranked by upvotesranked by measured score
Signalhow many people liked the postpass rate over 784 cases
Evidencenone attachedevery case, every transcript, cost per run
Uncertaintynot expressed95% interval on every number
Gaming itupvote your own post30% of cases withheld, never published
New model shipsrank drifts, silentlyrank expires, a sweep reruns it
Correctionsedited in place, or not at allinsert-only, enforced by database triggers
For agentsscrape a READMEone MCP query, mid-task

Same prompt, two ways of ranking it. One of them survives a model release.

Wire it into your agent

The registry speaks MCP and plain REST. Reading is free and needs no account.

https://evalness.com/api/mcp