Multi step reasoning
Arrive at a verifiable final answer across several dependent steps. Only the final answer is scored, so a lucky guess with broken work still fails the harder cases.
Multi step reasoningRefusal adherenceStructured JSON extractionSummarization fidelityTool call accuracy
model minimax-m3model nemotron-3-ultra-550b-a55bmodel nemotron-3-super-120b-a12bmodel deepseek-v4-flash-0731model claude-fable-5model gpt-5.6-solmodel gemini-3.2-promodel llama-4.2-405bmodel claude-opus-4.8model mistral-large-33 entries · 180 cases
#EntryScore · 95% CI 57–95ValueHoldoutCostState
01verify-then-answer@marek · v2 · n=126Reading this board
- Order
- Entries are ordered by their public score on the selected model version. The interval bar shows the 95 percent confidence range, so two entries whose bars overlap are not meaningfully apart.
- Holdout
- Thirty percent of every suite is withheld and never published. A large gap between the public and holdout score is the signal that an entry was tuned to the visible cases.
- Freshness
- An entry is fresh while its newest run targets the current version of the selected model. Anything older is marked and queued for a rerun.