Plain instruction
The obvious prompt anyone would write first. Included as the baseline every other entry in this suite is measured against.
Prompt
// no system prompt
// user
Extract the data as JSON: {{document}}
Score history
Model versionVersionScore · CIHoldoutState
mistral-large-3v161.3 53.8–68.359.7fresh
gpt-5.6-solv170.2 62.9–76.669.4fresh
gemini-3.2-prov165.5 58.0–72.265.3fresh
llama-4.2-405bv157.1 49.6–64.456.9fresh
claude-opus-4.8v167.9 60.5–74.566.7fresh
claude-fable-5v170.8 63.6–77.269.4fresh
current scorev1
61.3
95% CI 53.8 to 68.3 · n=168
Holdout score59.7
Judgeprogrammatic
Cost per run$0.661
Median latency1.9s
LicenseCC-BY-4.0
Author@system
0 endorsements
Sign in to endorse this entry.
Across models
mistral-large-361.3fresh
gpt-5.6-sol70.2fresh
gemini-3.2-pro65.5fresh
llama-4.2-405b57.1fresh
claude-opus-4.867.9fresh
claude-fable-570.8fresh
Versions
v1Baseline61.3