Summarization fidelity
Compress a source without introducing claims it does not make. Every unsupported sentence is a failure regardless of how well it reads.
Multi step reasoningRefusal adherenceStructured JSON extractionSummarization fidelityTool call accuracy
model minimax-m3model nemotron-3-ultra-550b-a55bmodel nemotron-3-super-120b-a12bmodel deepseek-v4-flash-0731model claude-fable-5model gpt-5.6-solmodel gemini-3.2-promodel llama-4.2-405bmodel claude-opus-4.8model mistral-large-33 entries · 160 cases
#EntryScore · 95% CI 60–99ValueHoldoutCostState
01source-span-cite@marek · v2 · n=112Reading this board
- Order
- Entries are ordered by their public score on the selected model version. The interval bar shows the 95 percent confidence range, so two entries whose bars overlap are not meaningfully apart.
- Holdout
- Thirty percent of every suite is withheld and never published. A large gap between the public and holdout score is the signal that an entry was tuned to the visible cases.
- Freshness
- An entry is fresh while its newest run targets the current version of the selected model. Anything older is marked and queued for a rerun.