Everyone claims their prompt works.We publish the numbers.
18entries
784cases
10models
10scores
- 01Few shot tripletStructured JSON extraction · deepseek-v4-flash-073160.8–94.283.3
- 02Role then formatStructured JSON extraction · deepseek-v4-flash-073112.5–50.927.8
- 03Plain instructionStructured JSON extraction · deepseek-v4-flash-07310.0–17.60.0
- 04Production launch verification successfulMulti step reasoning · nemotron-3-super-120b-a12b0.0–2.10.0
- 05XML fence guardStructured JSON extraction · deepseek-v4-flash-07310.0–17.60.0
few-shot-tripletdeepseek-v4-flash-073183.3xml-fence-guardnemotron-3-super-120b-a12b0.0xml-fence-guarddeepseek-v4-flash-07310.0role-then-formatnemotron-3-super-120b-a12b27.8role-then-formatdeepseek-v4-flash-073127.8plain-instructionnemotron-3-super-120b-a12b0.0plain-instructiondeepseek-v4-flash-07310.0few-shot-tripletnemotron-3-super-120b-a12b83.3production-launch-verification-successfulnemotron-3-super-120b-a12b0.0
Five suites. Fixed cases. Named judges.
A suite decides everything: the cases, the judge, the holdout split. Entries compete inside a suite, never across them.
- 180casesMulti step reasoningArrive at a verifiable final answer across several dependent steps. Only the final answer is scored, so a lucky guess with broken work still fails the harder cases.
- 200casesRefusal adherenceRefuse what the policy says to refuse, comply with everything else. Scored both ways: over refusal is a failure, not a safe default.
- 24casesStructured JSON extractionPull a typed object out of unstructured text without commentary, fences, or schema drift. Scored on exact match against the target schema.
- 160casesSummarization fidelityCompress a source without introducing claims it does not make. Every unsupported sentence is a failure regardless of how well it reads.
- 220casesTool call accuracyChoose the right tool and populate its arguments from an ambiguous request. Scored on the emitted call, not on the prose around it.
Why scores instead of stars
ranked by upvotesranked by measured score
Signalhow many people liked the postpass rate over 784 cases
Evidencenone attachedevery case, every transcript, cost per run
Uncertaintynot expressed95% interval on every number
Gaming itupvote your own post30% of cases withheld, never published
New model shipsrank drifts, silentlyrank expires, a sweep reruns it
Correctionsedited in place, or not at allinsert-only, enforced by database triggers
For agentsscrape a READMEone MCP query, mid-task
Same prompt, two ways of ranking it. One of them survives a model release.
Wire it into your agent
The registry speaks MCP and plain REST. Reading is free and needs no account.
https://evalness.com/api/mcp