Anthropic · launched 22 Sept 2026
Claude Opus 5.5
Anthropic's newest Opus model. Tokens cost 20% less than Opus 5 and cache reads 60% less, with output generated more than 30% faster.
Read the official announcement- Vendor
- Anthropic
- Launch date
- 22 Sept 2026
- Input price
- $4.00 / 1M tokens
- Output price
- $20 / 1M tokens
- Cached input
- $0.20 / 1M tokens
- Batch discount
- 50%
- Context window
- 1M tokens
- Max output
- 128K tokens
- Input types
- Text, Image, File
- API model ID
- claude-opus-5-5
- OpenRouter ID
- anthropic/claude-opus-5.5
Input price
$4.00
was $5.00 · 20% cheaper
Output price
$20
was $25 · 20% cheaper
Blended price
$8.00
was $10 · 20% cheaper
- Terminal-Bench 4.0xhigh effort · $7.35 per attempt · reported by Anthropic66.4%
- FrontierCode v1.1medium effort · $0.80 per task · reported by Anthropic54.64%
- CursorBench 4.0max effort · $13.4 per task · reported by Anthropic57.8%
- GDPval-AA v2.1max effort · $8.92 per task · reported by Anthropic1846 Elo
- AutomationBenchmax effort · $1.37 per task · reported by Anthropic40%
- WANDRmax effort · $37.9 per attempt · reported by Anthropic72.3%
- Terminal-Bench 4.0xhigh effort · $7.35 per attempt · reported by Anthropic66.4%
- FrontierCode v1.1medium effort · $0.80 per task · reported by Anthropic54.6%
- CursorBench 4.0max effort · $13.4 per task · reported by Anthropic57.8%
- AA-Briefcase v1.1max effort · $21 per task · reported by Anthropic1822 Elo
Price compared with other models
| Model | Launched | Input / 1M | Output / 1M | Cached / 1M | Blended / 1M | vs Claude Opus 5.5 |
|---|---|---|---|---|---|---|
| GPT-6 LunaOpenAI | 22 Sept 2026 | $0.10 | $0.50 | $0.01 | $0.20 | 98% cheaper |
| GPT-5.6 LunaOpenAI | 9 Jul 2026 | $0.20 | $1.20 | $0.02 | $0.45 | 94% cheaper |
| Claude Sonnet 5Anthropic | 30 Jun 2026 | $2.00 | $10 | $0.20 | $4.00 | 50% cheaper |
| Claude Sonnet 5.5Anthropic | 28 Sept 2026 | $2.00 | $10 | $0.20 | $4.00 | 50% cheaper |
| GPT-6 SolOpenAI | 22 Sept 2026 | $2.00 | $10 | $0.20 | $4.00 | 50% cheaper |
| Claude Opus 5.5Anthropic | 22 Sept 2026 | $4.00 | $20 | $0.20 | $8.00 | Baseline |
| GPT-5.6 SolOpenAIPrice before the GPT-6 launch, from OpenAI's announcement. OpenRouter now lists $2 / $10. | 9 Jul 2026 | $4.00 | $20 | – | $8.00 | Same price |
| Claude Opus 5Anthropic · previous generation | 24 Jul 2026 | $5.00 | $25 | $0.50 | $10 | 25% more |
| Claude Fable 5Anthropic | 9 Jun 2026 | $10 | $50 | $1.00 | $20 | 2.5× the price |
| Claude Fable 5.1Anthropic | 1 Sept 2026 | $10 | $50 | $0.25 | $20 | 2.5× the price |
| GPT-6 AstraOpenAI | 4 Sept 2026 | $10 | $50 | $1.00 | $20 | 2.5× the price |
USD per million tokens, checked 2 Oct 2026 via OpenRouter. Comparison uses the blended price (3 input tokens for every output token).
Official benchmark results
Reported by Anthropic on 22 Sept 2026 · exact values
| Benchmark | Claude Opus 5.5 | Claude Fable 5.1 | Claude Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0Agentic coding | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 (main)Agentic coding | 54.4% | 50.3% | 48% | 53.3% | 47.5% |
| CursorBench 4.0Agentic coding | 57.8% | 51.8% | 46.6% | – | 41.7% |
| GDPval-AA v2.1Knowledge work | 1846 Elo | 1735 Elo | 1708 Elo | 1542 Elo | 1588 Elo |
| AutomationBenchBusiness workflows | 40% | 31.4% | 26.9% | 41.4% | 28.8% |
| Humanity's Last ExamMultidisciplinary reasoning · with tools | 67.7% | 65.6% | 63.6% | 57.2% | – |
| Terminal-Bench-Science 0.1Agentic scientific research | 58.7% | 52.6% | 29% | 64.6% | 22.4% |
| OSWorld 2.1Computer use · partial | 81.8% | 80.7% | 74% | – | – |
| ChartographyVisual chart recognition · with tools | 89% | 88.4% | 83.4% | – | – |
- Bold marks the best reported score in each row.
- Claude Opus 5.5 results use adaptive thinking at max effort unless noted.
- Terminal-Bench 4.0 shows each model's highest score: Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort. GPT-6 Astra and GPT-5.6 Sol figures there are as reported by OpenAI.
- A dash means the lab did not report a result for that model.
Reported by Anthropic on 28 Sept 2026 · exact values
| Benchmark | Claude Sonnet 5.5 | Claude Sonnet 5 | Claude Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0Agentic coding | 70.6% | 10.3% | 66.4% | – |
| FrontierCode 1.1 (main)Agentic coding · Sonnet 5.5 at max effort, 52.1% at xhigh | 46.2% | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0Agentic coding | 55.5% | 34.1% | 57.8% | – |
| GDPval-AA v2.1Knowledge work | 1844 Elo | 1449 Elo | 1846 Elo | 1487 Elo |
| AA-Briefcase v1.1Knowledge work | 1811 Elo | 1359 Elo | 1822 Elo | 1483 Elo |
| Humanity's Last ExamMultidisciplinary reasoning · with tools | 64.5% | 54.9% | 67.7% | – |
| OSWorld 2.1Computer use · partial | 80.1% | 57% | 81.8% | – |
| ChartographyVisual chart recognition · no tools | 61.6% | 15.6% | 64.4% | 53.6% |
- Bold marks the best reported score in each row.
- Claude Opus 5.5's Terminal-Bench 4.0 score is at xhigh effort, its highest.
- OpenAI recently fixed a bug affecting GPT-6 Sol's image understanding; its GDPval-AA, AA-Briefcase and Chartography scores may predate the fix.
- A dash means the lab did not report a result for that model.
Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values
Measures how well a model completes complex, multi-step professional tasks within a command line interface.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
- GPT-6 Astra
- GPT-5.6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 38.5% | $1.29 |
| Claude Opus 5.5 | medium | 57.6% | $2.94 |
| Claude Opus 5.5 | high | 64.2% | $3.88 |
| Claude Opus 5.5 | xhigh | 66.4% | $7.35 |
| Claude Opus 5.5 | max | 64.8% | $11.2 |
| Claude Fable 5.1 | low | 40.2% | $5.70 |
| Claude Fable 5.1 | medium | 43.4% | $7.80 |
| Claude Fable 5.1 | high | 49.4% | $10.5 |
| Claude Fable 5.1 | xhigh | 51.3% | $15.8 |
| Claude Fable 5.1 | max | 55.8% | $19.5 |
| Claude Opus 5 | low | 28.5% | $4.25 |
| Claude Opus 5 | medium | 41.2% | $7.00 |
| Claude Opus 5 | high | 47% | $10.6 |
| Claude Opus 5 | xhigh | 50.6% | $13.5 |
| Claude Opus 5 | max | 52.3% | $15.8 |
| GPT-6 Astra | low | 49.7% | $4.95 |
| GPT-6 Astra | medium | 53.9% | $6.15 |
| GPT-6 Astra | high | 57.9% | $7.21 |
| GPT-6 Astra | xhigh | 57.6% | $7.48 |
| GPT-6 Astra | max | 56.7% | $10.3 |
| GPT-5.6 Sol | low | 7.9% | $1.46 |
| GPT-5.6 Sol | medium | 20.9% | $2.69 |
| GPT-5.6 Sol | high | 26.1% | $4.12 |
| GPT-5.6 Sol | xhigh | 28.5% | $5.39 |
| GPT-5.6 Sol | max | 37.3% | $7.89 |
- Exact values from the data published with Anthropic's announcement.
- GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI.
Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values
Measures whether an agent's code changes would be merged.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
- GPT-6 Astra
- GPT-5.6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 47.3% | $0.40 |
| Claude Opus 5.5 | medium | 54.64% | $0.80 |
| Claude Opus 5.5 | high | 53.99% | $1.09 |
| Claude Opus 5.5 | xhigh | 51.42% | $2.25 |
| Claude Opus 5.5 | max | 54.43% | $6.19 |
| Claude Fable 5.1 | low | 52.8% | $2.47 |
| Claude Fable 5.1 | medium | 50.91% | $3.28 |
| Claude Fable 5.1 | high | 50.34% | $5.27 |
| Claude Fable 5.1 | xhigh | 48.73% | $9.27 |
| Claude Fable 5.1 | max | 50.28% | $12.8 |
| Claude Opus 5 | low | 41.95% | $2.64 |
| Claude Opus 5 | medium | 53.38% | $4.61 |
| Claude Opus 5 | high | 47.99% | $7.62 |
| Claude Opus 5 | xhigh | 43.65% | $8.99 |
| Claude Opus 5 | max | 48.04% | $12.3 |
| GPT-6 Astra | low | 45.27% | $1.59 |
| GPT-6 Astra | medium | 48.83% | $2.28 |
| GPT-6 Astra | high | 50.94% | $2.85 |
| GPT-6 Astra | xhigh | 50.62% | $3.10 |
| GPT-6 Astra | max | 53.26% | $4.36 |
| GPT-5.6 Sol | low | 35.44% | $1.75 |
| GPT-5.6 Sol | medium | 39.93% | $2.50 |
| GPT-5.6 Sol | high | 45.06% | $3.25 |
| GPT-5.6 Sol | xhigh | 46.84% | $3.88 |
| GPT-5.6 Sol | max | 47.49% | $4.85 |
- Exact values from the data published with Anthropic's announcement.
Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values
Evaluates coding agents on ambiguous, multi-file tasks taken from real Cursor sessions.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
- GPT-5.6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 43.7% | $1.18 |
| Claude Opus 5.5 | medium | 52.5% | $2.90 |
| Claude Opus 5.5 | high | 56% | $3.97 |
| Claude Opus 5.5 | xhigh | 56% | $6.99 |
| Claude Opus 5.5 | max | 57.8% | $13.4 |
| Claude Fable 5.1 | low | 45.1% | $5.44 |
| Claude Fable 5.1 | medium | 46.8% | $7.05 |
| Claude Fable 5.1 | high | 49.2% | $9.08 |
| Claude Fable 5.1 | xhigh | 51.6% | $13 |
| Claude Fable 5.1 | max | 51.8% | $17.3 |
| Claude Opus 5 | low | 40.7% | $4.87 |
| Claude Opus 5 | medium | 43.3% | $6.94 |
| Claude Opus 5 | high | 44.7% | $9.00 |
| Claude Opus 5 | xhigh | 46.1% | $11.4 |
| Claude Opus 5 | max | 46.6% | $11.9 |
| GPT-5.6 Sol | low | 24.6% | $0.87 |
| GPT-5.6 Sol | medium | 31.1% | $1.77 |
| GPT-5.6 Sol | high | 35.7% | $2.85 |
| GPT-5.6 Sol | xhigh | 37.7% | $4.40 |
| GPT-5.6 Sol | max | 41.7% | $8.23 |
- Exact values from the data published with Anthropic's announcement.
Elo against cost · reported by Anthropic on 22 Sept 2026 · exact values
Artificial Analysis's evaluation of agents on real-world professional work across 44 occupations.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
- GPT-6 Astra
- GPT-5.6 Sol
Show data table
| Model | Effort | Elo | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 1224 Elo | $0.21 |
| Claude Opus 5.5 | medium | 1576 Elo | $0.86 |
| Claude Opus 5.5 | high | 1692 Elo | $1.54 |
| Claude Opus 5.5 | xhigh | 1820 Elo | $4.21 |
| Claude Opus 5.5 | max | 1846 Elo | $8.92 |
| Claude Fable 5.1 | low | 1450 Elo | $1.41 |
| Claude Fable 5.1 | medium | 1536 Elo | $2.17 |
| Claude Fable 5.1 | high | 1617 Elo | $3.43 |
| Claude Fable 5.1 | xhigh | 1721 Elo | $7.09 |
| Claude Fable 5.1 | max | 1735 Elo | $9.59 |
| Claude Opus 5 | low | 1294 Elo | $0.58 |
| Claude Opus 5 | medium | 1476 Elo | $1.36 |
| Claude Opus 5 | high | 1581 Elo | $3.03 |
| Claude Opus 5 | xhigh | 1676 Elo | $4.96 |
| Claude Opus 5 | max | 1708 Elo | $6.76 |
| GPT-6 Astra | low | 1366 Elo | $0.85 |
| GPT-6 Astra | medium | 1468 Elo | $1.82 |
| GPT-6 Astra | high | 1485 Elo | $2.43 |
| GPT-6 Astra | xhigh | 1516 Elo | $3.04 |
| GPT-6 Astra | max | 1542 Elo | $4.53 |
| GPT-5.6 Sol | low | 1289 Elo | $0.27 |
| GPT-5.6 Sol | medium | 1403 Elo | $0.60 |
| GPT-5.6 Sol | high | 1480 Elo | $1.11 |
| GPT-5.6 Sol | xhigh | 1548 Elo | $1.70 |
| GPT-5.6 Sol | max | 1588 Elo | $2.81 |
- Exact values from the data published with Anthropic's announcement.
Pass rate against cost · reported by Anthropic on 22 Sept 2026 · exact values
Built by Zapier, tests whether an agent can carry out real business workflows across many connected apps.
- Claude Opus 5.5
- Claude Opus 5
- GPT-6 Astra
- GPT-5.6 Sol
Show data table
| Model | Effort | Pass rate | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 23.3% | $0.50 |
| Claude Opus 5.5 | medium | 28.6% | $0.64 |
| Claude Opus 5.5 | high | 32% | $0.70 |
| Claude Opus 5.5 | xhigh | 34.4% | $0.86 |
| Claude Opus 5.5 | max | 40% | $1.37 |
| Claude Opus 5 | low | 20.4% | $0.75 |
| Claude Opus 5 | medium | 23.9% | $0.89 |
| Claude Opus 5 | high | 20.6% | $1.03 |
| Claude Opus 5 | xhigh | 25.3% | $1.15 |
| Claude Opus 5 | max | 26.9% | $1.27 |
| GPT-6 Astra | low | 30.3% | $1.08 |
| GPT-6 Astra | medium | 34.1% | $1.28 |
| GPT-6 Astra | high | 37.1% | $1.45 |
| GPT-6 Astra | xhigh | 39% | $1.53 |
| GPT-6 Astra | max | 41.4% | $1.77 |
| GPT-5.6 Sol | low | 11.7% | $0.42 |
| GPT-5.6 Sol | medium | 19.6% | $0.57 |
| GPT-5.6 Sol | high | 24.8% | $0.64 |
| GPT-5.6 Sol | xhigh | 26.3% | $0.74 |
| GPT-5.6 Sol | max | 28.8% | $0.91 |
- Exact values from the data published with Anthropic's announcement.
- Runs were performed by Zapier without fallback models, so safeguard interventions counted as failures.
Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values
Perplexity's benchmark measuring agents on large data collection tasks.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 31.2% | $1.20 |
| Claude Opus 5.5 | medium | 62.8% | $11.2 |
| Claude Opus 5.5 | high | 67.3% | $16.9 |
| Claude Opus 5.5 | xhigh | 71.3% | $29.1 |
| Claude Opus 5.5 | max | 72.3% | $37.9 |
| Claude Fable 5.1 | low | 63.3% | $23.3 |
| Claude Fable 5.1 | medium | 64.5% | $27.4 |
| Claude Fable 5.1 | high | 66.7% | $33.6 |
| Claude Fable 5.1 | xhigh | 67.7% | $42.8 |
| Claude Fable 5.1 | max | 68.7% | $49 |
| Claude Opus 5 | low | 50.5% | $10.7 |
| Claude Opus 5 | medium | 58.1% | $24 |
| Claude Opus 5 | high | 64.6% | $43.5 |
| Claude Opus 5 | xhigh | 67% | $53.8 |
| Claude Opus 5 | max | 67.2% | $61.6 |
- Exact values from the data published with Anthropic's announcement.
- Claude models ran with offline web search and fetch tools and a 980K-token task budget. This differs from Perplexity's published setup, so scores are not comparable with Perplexity's leaderboard.
Accuracy against cost · reported by Anthropic on 28 Sept 2026 · exact values
Measures how well a model completes complex, multi-step professional tasks within a command-line interface.
- Claude Sonnet 5.5
- Claude Opus 5.5
- Claude Sonnet 5
- GPT-5.6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Sonnet 5.5 | low | 20% | $0.76 |
| Claude Sonnet 5.5 | medium | 28.8% | $0.83 |
| Claude Sonnet 5.5 | high | 43% | $1.94 |
| Claude Sonnet 5.5 | xhigh | 61.5% | $5.30 |
| Claude Sonnet 5.5 | max | 70.6% | $12.5 |
| Claude Opus 5.5 | low | 38.5% | $1.29 |
| Claude Opus 5.5 | medium | 57.6% | $2.94 |
| Claude Opus 5.5 | high | 64.2% | $3.88 |
| Claude Opus 5.5 | xhigh | 66.4% | $7.35 |
| Claude Opus 5.5 | max | 64.8% | $11.2 |
| Claude Sonnet 5 | low | 3.2% | $3.78 |
| Claude Sonnet 5 | medium | 4.4% | $4.83 |
| Claude Sonnet 5 | high | 4.5% | $8.20 |
| Claude Sonnet 5 | xhigh | 7% | $9.95 |
| Claude Sonnet 5 | max | 10.3% | $11.6 |
| GPT-5.6 Sol | low | 7.9% | $1.46 |
| GPT-5.6 Sol | medium | 20.9% | $2.69 |
| GPT-5.6 Sol | high | 26.1% | $4.12 |
| GPT-5.6 Sol | xhigh | 28.5% | $5.39 |
| GPT-5.6 Sol | max | 37.3% | $7.89 |
- Exact values from the data published with Anthropic's announcement.
- OpenAI did not report GPT-6 Sol on this benchmark, so Anthropic shows GPT-5.6 Sol.
Accuracy against cost · reported by Anthropic on 28 Sept 2026 · exact values
Measures whether an agent's code changes would be merged.
- Claude Sonnet 5.5
- Claude Opus 5.5
- Claude Sonnet 5
- GPT-6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Sonnet 5.5 | low | 29.3% | $0.19 |
| Claude Sonnet 5.5 | medium | 36.5% | $0.24 |
| Claude Sonnet 5.5 | high | 49.4% | $0.42 |
| Claude Sonnet 5.5 | xhigh | 52.1% | $1.59 |
| Claude Sonnet 5.5 | max | 46.2% | $20.8 |
| Claude Opus 5.5 | low | 47.3% | $0.40 |
| Claude Opus 5.5 | medium | 54.6% | $0.80 |
| Claude Opus 5.5 | high | 54% | $1.09 |
| Claude Opus 5.5 | xhigh | 51.4% | $2.25 |
| Claude Opus 5.5 | max | 54.4% | $6.19 |
| Claude Sonnet 5 | low | 28.7% | $2.39 |
| Claude Sonnet 5 | medium | 35.2% | $3.81 |
| Claude Sonnet 5 | high | 39.4% | $6.10 |
| Claude Sonnet 5 | xhigh | 42.7% | $10.1 |
| Claude Sonnet 5 | max | 42.4% | $17.1 |
| GPT-6 Sol | low | 37.3% | $0.43 |
| GPT-6 Sol | medium | 45.9% | $0.77 |
| GPT-6 Sol | high | 47.7% | $1.04 |
| GPT-6 Sol | xhigh | 48.4% | $1.32 |
| GPT-6 Sol | max | 49.3% | $2.07 |
- Exact values from the data published with Anthropic's announcement.
Accuracy against cost · reported by Anthropic on 28 Sept 2026 · exact values
Evaluates coding agents on ambiguous, multi-file tasks taken from real Cursor sessions.
- Claude Sonnet 5.5
- Claude Opus 5.5
- Claude Sonnet 5
- GPT-5.6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Sonnet 5.5 | low | 35.8% | $0.50 |
| Claude Sonnet 5.5 | medium | 39.2% | $0.70 |
| Claude Sonnet 5.5 | high | 47.8% | $1.67 |
| Claude Sonnet 5.5 | xhigh | 53.1% | $3.88 |
| Claude Sonnet 5.5 | max | 55.5% | $9.67 |
| Claude Opus 5.5 | low | 43.7% | $1.17 |
| Claude Opus 5.5 | medium | 52.5% | $2.91 |
| Claude Opus 5.5 | high | 56% | $3.97 |
| Claude Opus 5.5 | xhigh | 56% | $6.98 |
| Claude Opus 5.5 | max | 57.8% | $13.4 |
| Claude Sonnet 5 | low | 24.1% | $1.39 |
| Claude Sonnet 5 | medium | 28% | $2.31 |
| Claude Sonnet 5 | high | 30.8% | $3.48 |
| Claude Sonnet 5 | xhigh | 32% | $4.55 |
| Claude Sonnet 5 | max | 34.1% | $7.17 |
| GPT-5.6 Sol | low | 24.6% | $0.87 |
| GPT-5.6 Sol | medium | 31.1% | $1.77 |
| GPT-5.6 Sol | high | 35.7% | $2.85 |
| GPT-5.6 Sol | xhigh | 37.7% | $4.40 |
| GPT-5.6 Sol | max | 41.7% | $8.23 |
- Exact values from the data published with Anthropic's announcement.
- CursorBench does not report GPT-6 Sol publicly, so Anthropic shows GPT-5.6 Sol.
Elo against cost · reported by Anthropic on 28 Sept 2026 · exact values
Artificial Analysis's benchmark of long-horizon knowledge work.
- Claude Sonnet 5.5
- Claude Opus 5.5
- Claude Sonnet 5
- GPT-6 Sol
Show data table
| Model | Effort | Elo | Cost per task |
|---|---|---|---|
| Claude Sonnet 5.5 | low | 1264 Elo | $0.87 |
| Claude Sonnet 5.5 | medium | 1461 Elo | $1.64 |
| Claude Sonnet 5.5 | high | 1634 Elo | $3.95 |
| Claude Sonnet 5.5 | xhigh | 1746 Elo | $9.63 |
| Claude Sonnet 5.5 | max | 1811 Elo | $29.2 |
| Claude Opus 5.5 | low | 1285 Elo | $1.15 |
| Claude Opus 5.5 | medium | 1642 Elo | $4.40 |
| Claude Opus 5.5 | high | 1705 Elo | $6.27 |
| Claude Opus 5.5 | xhigh | 1780 Elo | $12.3 |
| Claude Opus 5.5 | max | 1822 Elo | $21 |
| Claude Sonnet 5 | low | 923 Elo | $0.82 |
| Claude Sonnet 5 | medium | 1056 Elo | $1.73 |
| Claude Sonnet 5 | high | 1177 Elo | $3.76 |
| Claude Sonnet 5 | xhigh | 1274 Elo | $7.56 |
| Claude Sonnet 5 | max | 1359 Elo | $14.4 |
| GPT-6 Sol | low | 905 Elo | $0.12 |
| GPT-6 Sol | medium | 1142 Elo | $0.34 |
| GPT-6 Sol | high | 1289 Elo | $0.63 |
| GPT-6 Sol | xhigh | 1364 Elo | $1.19 |
| GPT-6 Sol | max | 1483 Elo | $2.67 |
- Exact values from the data published with Anthropic's announcement.
- Artificial Analysis ran Sonnet 5.5 on a pre-release deployment with a since-fixed structured output bug, which may slightly understate its score.