Methodology
How a number on this site is produced, and what it does not tell you. If any of this is wrong, the registry is worthless, so it is written down rather than implied.
Purpose
Evalness is a public registry for prompting techniques: a place to submit a prompt or skill, argue about it in the open, and settle the argument with a measured score instead of an upvote. The ranking is the referee; the discussion is the reason to come back. Everyone else ranks prompts by opinion — here opinion and proof sit on the same page.
It is built in three parts. Rankings are live: any published entry is run against a suite on a named model version and reported with an interval, a holdout, and a cost. Discussion is being built — threaded, versioned, and attached to the entry it argues about, so a claim and its evidence never drift apart. A verified skill library is the destination: skills that pass an adversarial review — accuracy measured, capabilities declared and checked — become public and are exposed over a machine API so an agent can pull a technique it can trust rather than a snippet it cannot.
The claim a badge makes is deliberately narrow. Not harmless, which cannot be proven, but passed these checks, on these cases, on this date, requesting these capabilities. An unverifiable promise is the same failure as an unmeasured score, and this registry exists to avoid exactly that.
Scores
A score is the share of cases an entry passed, on one suite, against one named model version, at one point in time. It is reported with a 95 percent Wilson confidence interval, which behaves correctly near 0 and 100 percent where the normal approximation produces impossible bounds.
Two entries are only meaningfully different when their intervals do not overlap. The compare page states this verdict explicitly rather than leaving you to eyeball two numbers.
Holdout
Thirty percent of every suite is withheld and never published. Assignment is deterministic from the case key, so republishing a suite cannot leak a withheld case, and the same case always lands on the same side of the split.
Both the public and holdout score are shown. A large gap between them means the entry was tuned to the visible cases and will not generalise.
Decay
A score belongs to one model version. When a provider ships a new version, the old score is frozen, not migrated, and the entry is marked stale until it is rerun. Retired versions keep their scores forever.
Immutability
Prompts, cases, per-case results, and published scores are insert-only. This is enforced by database triggers, not by application code, so a console session or a future migration cannot quietly rewrite history. A correction is a new row; the old one is marked superseded and stays visible.
Cost
Every run reports the spend it actually incurred, computed from token counts and the provider rate for that version. Three caps apply: per run, per user per day, and a global daily ceiling. A run that would cross any of them is refused before it starts.
Independence
Nothing on this registry can be bought. There is no paid placement, no sponsored score, and no provider influence over cases, judges, or rank. A paid tier buys volume, privacy, and verification, never position.
Model providers do not fund sweeps and cannot see holdout cases. If that ever changes, the change will be written here before it takes effect.
What this does not tell you
A suite is a proxy, not the truth. A high score means an entry did well on these cases with this judge, not that it is the best prompt for your task. The cases are published so you can decide whether the proxy resembles your problem.