Plain instruction

The obvious prompt anyone would write first. Included as the baseline every other entry in this suite is measured against.

Compare

Prompt

// no system prompt // user Extract the data as JSON: {{document}}

Score history

Model versionVersionScore · CIHoldoutState
mistral-large-3v161.3 53.868.359.7fresh
gpt-5.6-solv170.2 62.976.669.4fresh
gemini-3.2-prov165.5 58.072.265.3fresh
llama-4.2-405bv157.1 49.664.456.9fresh
claude-opus-4.8v167.9 60.574.566.7fresh
claude-fable-5v170.8 63.677.269.4fresh
current scorev1
61.3
95% CI 53.8 to 68.3 · n=168
Holdout score59.7
Judgeprogrammatic
Cost per run$0.661
Median latency1.9s
LicenseCC-BY-4.0
Author@system
0 endorsements

Sign in to endorse this entry.

Across models

mistral-large-361.3fresh
gpt-5.6-sol70.2fresh
gemini-3.2-pro65.5fresh
llama-4.2-405b57.1fresh
claude-opus-4.867.9fresh
claude-fable-570.8fresh

Versions

v1Baseline61.3