Tool call accuracy
Choose the right tool and populate its arguments from an ambiguous request. Scored on the emitted call, not on the prose around it.
Multi step reasoningRefusal adherenceStructured JSON extractionSummarization fidelityTool call accuracy
model minimax-m3model nemotron-3-ultra-550b-a55bmodel nemotron-3-super-120b-a12bmodel deepseek-v4-flash-0731model claude-fable-5model gpt-5.6-solmodel gemini-3.2-promodel llama-4.2-405bmodel claude-opus-4.8model mistral-large-33 entries · 220 cases
#EntryScore · 95% CI 64–98ValueHoldoutCostState
01signature-echo@dara · v2 · n=154Reading this board
- Order
- Entries are ordered by their public score on the selected model version. The interval bar shows the 95 percent confidence range, so two entries whose bars overlap are not meaningfully apart.
- Holdout
- Thirty percent of every suite is withheld and never published. A large gap between the public and holdout score is the signal that an entry was tuned to the visible cases.
- Freshness
- An entry is fresh while its newest run targets the current version of the selected model. Anything older is marked and queued for a rerun.