Anthropic · launched 1 Sept 2026
Claude Fable 5.1
Released alongside the access-limited Claude Mythos 5.1. Cache reads cost 75% less than Fable 5, which Anthropic estimates cuts typical costs by about 25%.
Read the official announcement- Vendor
- Anthropic
- Launch date
- 1 Sept 2026
- Input price
- $10 / 1M tokens
- Output price
- $50 / 1M tokens
- Cached input
- $0.25 / 1M tokens
- Batch discount
- 50%
- Context window
- 1M tokens
- Max output
- 128K tokens
- Input types
- Text, Image, File
- API model ID
- claude-fable-5-1
- OpenRouter ID
- anthropic/claude-fable-5.1
Input price
$10
was $10 · Same price
Output price
$50
was $50 · Same price
Blended price
$20
was $20 · Same price
- Terminal-Bench-Science 0.1max effort · $37.9 per task · reported by Anthropic52.6%
- Humanity's Last Exam (tools)xhigh effort · $2.28 per task · reported by Anthropic65.08%
- Humanity's Last Exam (no tools)max effort · $2.23 per task · reported by Anthropic60.92%
- CursorBench 3.2.0max effort · $9.64 per task · reported by Anthropic73.4%
- AutomationBench 1.0.6max effort · $2.45 per task · reported by OpenAI31.4%
- FrontierCode 1.1medium effort · $3.28 per task · reported by OpenAI50.9%
- Terminal-Bench 4.0max effort · $19.5 per attempt · reported by Anthropic55.8%
- FrontierCode v1.1low effort · $2.47 per task · reported by Anthropic52.8%
- CursorBench 4.0max effort · $17.3 per task · reported by Anthropic51.8%
- GDPval-AA v2.1max effort · $9.59 per task · reported by Anthropic1735 Elo
- WANDRmax effort · $49 per attempt · reported by Anthropic68.7%
- Terminal-Bench 4.0max effort · $19.5 per task · reported by OpenAI55.8%
Price compared with other models
| Model | Launched | Input / 1M | Output / 1M | Cached / 1M | Blended / 1M | vs Claude Fable 5.1 |
|---|---|---|---|---|---|---|
| GPT-6 LunaOpenAI | 22 Sept 2026 | $0.10 | $0.50 | $0.01 | $0.20 | 99% cheaper |
| GPT-5.6 LunaOpenAI | 9 Jul 2026 | $0.20 | $1.20 | $0.02 | $0.45 | 98% cheaper |
| Claude Sonnet 5Anthropic | 30 Jun 2026 | $2.00 | $10 | $0.20 | $4.00 | 80% cheaper |
| Claude Sonnet 5.5Anthropic | 28 Sept 2026 | $2.00 | $10 | $0.20 | $4.00 | 80% cheaper |
| GPT-6 SolOpenAI | 22 Sept 2026 | $2.00 | $10 | $0.20 | $4.00 | 80% cheaper |
| Claude Opus 5.5Anthropic | 22 Sept 2026 | $4.00 | $20 | $0.20 | $8.00 | 60% cheaper |
| GPT-5.6 SolOpenAIPrice before the GPT-6 launch, from OpenAI's announcement. OpenRouter now lists $2 / $10. | 9 Jul 2026 | $4.00 | $20 | – | $8.00 | 60% cheaper |
| Claude Opus 5Anthropic | 24 Jul 2026 | $5.00 | $25 | $0.50 | $10 | 50% cheaper |
| Claude Fable 5Anthropic · previous generation | 9 Jun 2026 | $10 | $50 | $1.00 | $20 | Same price |
| Claude Fable 5.1Anthropic | 1 Sept 2026 | $10 | $50 | $0.25 | $20 | Baseline |
| GPT-6 AstraOpenAI | 4 Sept 2026 | $10 | $50 | $1.00 | $20 | Same price |
USD per million tokens, checked 2 Oct 2026 via OpenRouter. Comparison uses the blended price (3 input tokens for every output token).
Official benchmark results
Reported by Anthropic on 1 Sept 2026 · exact values
| Benchmark | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Terminal-Bench-Science 0.1Agentic scientific research | 52.6% | 24.7% | 29% | 22.4% |
| Terminal-Bench 4.0Agentic coding | 55.8% | 42% | 52.3% | 37.3% |
| GDPval-AA v2Knowledge work | 1853 Elo | 1723 Elo | 1824 Elo | 1711 Elo |
| OSWorld 2.0Computer use · partial | 77.9% | 72.9% | 75.4% | – |
| OSWorld 2.0 (strict)Computer use · strict | 41.7% | 36.1% | 39.6% | – |
| Humanity's Last Exam (no tools)Multidisciplinary reasoning · no tools | 60.9% | 57.8% | 56.6% | – |
| Humanity's Last Exam (with tools)Multidisciplinary reasoning · with tools | 65% | 63.8% | 63.6% | – |
| AutomationBenchBusiness workflows | 31.4% | 17.1% | 26.9% | 19.6% |
| CursorBench 3.2.0Agentic coding | 73.4% | 70.5% | 70% | 67.2% |
- Bold marks the best reported score in each row.
- Fable 5.1 was evaluated with production safeguards enabled. Where they intervened, Fable 5.1 and Fable 5 scored zero on OSWorld 2.0 and Fable 5 scored zero on AutomationBench.
- Claude Mythos 5.1, the same underlying model with access limited to vetted users, scores 60.9% on Terminal-Bench 4.0. Mythos has no public API price, so it is not tracked here.
- A dash means the lab did not report a result for that model.
Reported by Anthropic on 22 Sept 2026 · exact values
| Benchmark | Claude Opus 5.5 | Claude Fable 5.1 | Claude Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0Agentic coding | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 (main)Agentic coding | 54.4% | 50.3% | 48% | 53.3% | 47.5% |
| CursorBench 4.0Agentic coding | 57.8% | 51.8% | 46.6% | – | 41.7% |
| GDPval-AA v2.1Knowledge work | 1846 Elo | 1735 Elo | 1708 Elo | 1542 Elo | 1588 Elo |
| AutomationBenchBusiness workflows | 40% | 31.4% | 26.9% | 41.4% | 28.8% |
| Humanity's Last ExamMultidisciplinary reasoning · with tools | 67.7% | 65.6% | 63.6% | 57.2% | – |
| Terminal-Bench-Science 0.1Agentic scientific research | 58.7% | 52.6% | 29% | 64.6% | 22.4% |
| OSWorld 2.1Computer use · partial | 81.8% | 80.7% | 74% | – | – |
| ChartographyVisual chart recognition · with tools | 89% | 88.4% | 83.4% | – | – |
- Bold marks the best reported score in each row.
- Claude Opus 5.5 results use adaptive thinking at max effort unless noted.
- Terminal-Bench 4.0 shows each model's highest score: Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort. GPT-6 Astra and GPT-5.6 Sol figures there are as reported by OpenAI.
- A dash means the lab did not report a result for that model.
Accuracy against cost · reported by Anthropic on 1 Sept 2026 · exact values
Terminal-based agentic scientific research tasks, in Anthropic's agentic scientific research category.
- Claude Fable 5.1
- Claude Fable 5
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Fable 5.1 | low | 26.3% | $11.1 |
| Claude Fable 5.1 | medium | 35.7% | $14.9 |
| Claude Fable 5.1 | high | 40% | $20.3 |
| Claude Fable 5.1 | xhigh | 49.5% | $31.8 |
| Claude Fable 5.1 | max | 52.6% | $37.9 |
| Claude Fable 5 | low | 12.3% | $17.1 |
| Claude Fable 5 | medium | 21.4% | $25 |
| Claude Fable 5 | high | 25% | $34.3 |
| Claude Fable 5 | xhigh | 23.4% | $36 |
| Claude Fable 5 | max | 24.7% | $44.1 |
- Exact values from the data published with Anthropic's announcement.
- Standard error is about 3.5 to 4.5 points per model.
Pass rate against cost · reported by Anthropic on 1 Sept 2026 · exact values
Expert-level questions across many disciplines, answered with access to tools.
- Claude Fable 5.1
- Claude Fable 5
Show data table
| Model | Effort | Pass rate | Cost per task |
|---|---|---|---|
| Claude Fable 5.1 | low | 60% | $0.52 |
| Claude Fable 5.1 | medium | 62.96% | $0.67 |
| Claude Fable 5.1 | high | 64.76% | $1.05 |
| Claude Fable 5.1 | xhigh | 65.08% | $2.28 |
| Claude Fable 5.1 | max | 65% | $3.20 |
| Claude Fable 5 | low | 59.64% | $0.61 |
| Claude Fable 5 | medium | 61.4% | $1.01 |
| Claude Fable 5 | high | 63.16% | $1.42 |
| Claude Fable 5 | xhigh | 63.52% | $1.92 |
| Claude Fable 5 | max | 63.8% | $3.44 |
- Exact values from the data published with Anthropic's announcement.
Pass rate against cost · reported by Anthropic on 1 Sept 2026 · exact values
Expert-level questions across many disciplines, answered without tools.
- Claude Fable 5.1
- Claude Fable 5
Show data table
| Model | Effort | Pass rate | Cost per task |
|---|---|---|---|
| Claude Fable 5.1 | low | 53.16% | $0.30 |
| Claude Fable 5.1 | medium | 55.92% | $0.46 |
| Claude Fable 5.1 | high | 57.96% | $0.75 |
| Claude Fable 5.1 | xhigh | 60.36% | $1.53 |
| Claude Fable 5.1 | max | 60.92% | $2.23 |
| Claude Fable 5 | low | 50.6% | $0.17 |
| Claude Fable 5 | medium | 55.88% | $0.40 |
| Claude Fable 5 | high | 56.88% | $0.62 |
| Claude Fable 5 | xhigh | 57.44% | $0.91 |
| Claude Fable 5 | max | 57.76% | $1.70 |
- Exact values from the data published with Anthropic's announcement.
Accuracy against cost · reported by Anthropic on 1 Sept 2026 · exact values
Evaluates coding agents on tasks taken from real Cursor sessions (earlier release than CursorBench 4.0).
- Claude Fable 5.1
- Claude Fable 5
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Fable 5.1 | low | 66.2% | $2.90 |
| Claude Fable 5.1 | medium | 68% | $3.53 |
| Claude Fable 5.1 | high | 69.4% | $4.80 |
| Claude Fable 5.1 | xhigh | 72.8% | $6.96 |
| Claude Fable 5.1 | max | 73.4% | $9.64 |
| Claude Fable 5 | low | 62.1% | $4.46 |
| Claude Fable 5 | medium | 65.2% | $6.80 |
| Claude Fable 5 | high | 66.5% | $8.77 |
| Claude Fable 5 | xhigh | 68.4% | $11.7 |
| Claude Fable 5 | max | 70.5% | $17.3 |
- Exact values from the data published with Anthropic's announcement.
Score against cost · reported by OpenAI on 22 Sept 2026 · exact values
AI agents are tested on end-to-end workflows using 47 tools across sales, marketing, operations, support, finance and HR.
- GPT-6 Luna
- GPT-5.6 Luna
- GPT-6 Sol
- GPT-5.6 Sol
- GPT-6 Astra
- Claude Opus 5
- Claude Fable 5.1
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| GPT-6 Luna | low | 1.2% | $0.006 |
| GPT-6 Luna | medium | 9.4% | $0.016 |
| GPT-6 Luna | high | 14.5% | $0.021 |
| GPT-6 Luna | xhigh | 12.6% | $0.025 |
| GPT-6 Luna | max | 20.7% | $0.037 |
| GPT-5.6 Luna | low | 1.8% | $0.01 |
| GPT-5.6 Luna | medium | 4.3% | $0.02 |
| GPT-5.6 Luna | high | 9.1% | $0.05 |
| GPT-5.6 Luna | xhigh | 12.9% | $0.06 |
| GPT-5.6 Luna | max | 17% | $0.07 |
| GPT-6 Sol | low | 21.2% | $0.19 |
| GPT-6 Sol | medium | 26.9% | $0.21 |
| GPT-6 Sol | high | 31.2% | $0.24 |
| GPT-6 Sol | xhigh | 33.2% | $0.27 |
| GPT-6 Sol | max | 32% | $0.34 |
| GPT-5.6 Sol | low | 11.7% | $0.31 |
| GPT-5.6 Sol | medium | 19.6% | $0.42 |
| GPT-5.6 Sol | high | 24.8% | $0.47 |
| GPT-5.6 Sol | xhigh | 26.3% | $0.54 |
| GPT-5.6 Sol | max | 28.8% | $0.67 |
| GPT-6 Astra | low | 30.3% | $1.08 |
| GPT-6 Astra | medium | 34.1% | $1.27 |
| GPT-6 Astra | high | 37.1% | $1.44 |
| GPT-6 Astra | xhigh | 39% | $1.50 |
| GPT-6 Astra | max | 41.4% | $1.73 |
| Claude Opus 5 | low | 20.4% | $1.64 |
| Claude Opus 5 | medium | 23.9% | $2.22 |
| Claude Opus 5 | high | 20.5% | $2.27 |
| Claude Opus 5 | xhigh | 25.3% | $2.71 |
| Claude Opus 5 | max | 26.9% | $3.05 |
| Claude Fable 5.1 | max | 31.4% | $2.45 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
- Claude Fable 5.1 ran with Claude Opus 5 as a fallback. OpenAI notes its cost omits the fallback runs (about 40% of tasks), so its real cost is higher.
Score against cost · reported by OpenAI on 22 Sept 2026 · exact values
AI agents write code graded on correctness and mergeability: test quality, scope discipline, code style and codebase standards.
- GPT-6 Luna
- GPT-5.6 Luna
- GPT-6 Sol
- GPT-5.6 Sol
- GPT-6 Astra
- Claude Opus 5
- Claude Fable 5.1
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| GPT-6 Luna | low | 25.7% | $0.021 |
| GPT-6 Luna | medium | 35.5% | $0.053 |
| GPT-6 Luna | high | 37.3% | $0.067 |
| GPT-6 Luna | xhigh | 37.1% | $0.073 |
| GPT-6 Luna | max | 42.4% | $0.11 |
| GPT-5.6 Luna | low | 15.4% | $0.06 |
| GPT-5.6 Luna | medium | 25.7% | $0.13 |
| GPT-5.6 Luna | high | 35.9% | $0.23 |
| GPT-5.6 Luna | xhigh | 38.9% | $0.31 |
| GPT-5.6 Luna | max | 39.8% | $0.37 |
| GPT-6 Sol | low | 37.3% | $0.45 |
| GPT-6 Sol | medium | 45.9% | $0.80 |
| GPT-6 Sol | high | 47.7% | $1.08 |
| GPT-6 Sol | xhigh | 48.4% | $1.37 |
| GPT-6 Sol | max | 49.3% | $2.14 |
| GPT-5.6 Sol | low | 35.4% | $1.89 |
| GPT-5.6 Sol | medium | 39.9% | $2.69 |
| GPT-5.6 Sol | high | 45.1% | $3.48 |
| GPT-5.6 Sol | xhigh | 46.8% | $4.15 |
| GPT-5.6 Sol | max | 47.5% | $5.19 |
| GPT-6 Astra | low | 45.3% | $1.70 |
| GPT-6 Astra | medium | 48.8% | $2.43 |
| GPT-6 Astra | high | 50.9% | $3.01 |
| GPT-6 Astra | xhigh | 50.6% | $3.28 |
| GPT-6 Astra | max | 53.3% | $4.59 |
| Claude Opus 5 | low | 41.9% | $2.68 |
| Claude Opus 5 | medium | 53.4% | $4.31 |
| Claude Opus 5 | high | 48% | $7.24 |
| Claude Opus 5 | xhigh | 43.6% | $9.14 |
| Claude Opus 5 | max | 48% | $11.4 |
| Claude Fable 5.1 | low | 49.8% | $2.38 |
| Claude Fable 5.1 | medium | 50.9% | $3.28 |
| Claude Fable 5.1 | high | 50.3% | $5.27 |
| Claude Fable 5.1 | xhigh | 48.7% | $9.27 |
| Claude Fable 5.1 | max | 50.3% | $12.8 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values
Measures how well a model completes complex, multi-step professional tasks within a command line interface.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
- GPT-6 Astra
- GPT-5.6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 38.5% | $1.29 |
| Claude Opus 5.5 | medium | 57.6% | $2.94 |
| Claude Opus 5.5 | high | 64.2% | $3.88 |
| Claude Opus 5.5 | xhigh | 66.4% | $7.35 |
| Claude Opus 5.5 | max | 64.8% | $11.2 |
| Claude Fable 5.1 | low | 40.2% | $5.70 |
| Claude Fable 5.1 | medium | 43.4% | $7.80 |
| Claude Fable 5.1 | high | 49.4% | $10.5 |
| Claude Fable 5.1 | xhigh | 51.3% | $15.8 |
| Claude Fable 5.1 | max | 55.8% | $19.5 |
| Claude Opus 5 | low | 28.5% | $4.25 |
| Claude Opus 5 | medium | 41.2% | $7.00 |
| Claude Opus 5 | high | 47% | $10.6 |
| Claude Opus 5 | xhigh | 50.6% | $13.5 |
| Claude Opus 5 | max | 52.3% | $15.8 |
| GPT-6 Astra | low | 49.7% | $4.95 |
| GPT-6 Astra | medium | 53.9% | $6.15 |
| GPT-6 Astra | high | 57.9% | $7.21 |
| GPT-6 Astra | xhigh | 57.6% | $7.48 |
| GPT-6 Astra | max | 56.7% | $10.3 |
| GPT-5.6 Sol | low | 7.9% | $1.46 |
| GPT-5.6 Sol | medium | 20.9% | $2.69 |
| GPT-5.6 Sol | high | 26.1% | $4.12 |
| GPT-5.6 Sol | xhigh | 28.5% | $5.39 |
| GPT-5.6 Sol | max | 37.3% | $7.89 |
- Exact values from the data published with Anthropic's announcement.
- GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI.
Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values
Measures whether an agent's code changes would be merged.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
- GPT-6 Astra
- GPT-5.6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 47.3% | $0.40 |
| Claude Opus 5.5 | medium | 54.64% | $0.80 |
| Claude Opus 5.5 | high | 53.99% | $1.09 |
| Claude Opus 5.5 | xhigh | 51.42% | $2.25 |
| Claude Opus 5.5 | max | 54.43% | $6.19 |
| Claude Fable 5.1 | low | 52.8% | $2.47 |
| Claude Fable 5.1 | medium | 50.91% | $3.28 |
| Claude Fable 5.1 | high | 50.34% | $5.27 |
| Claude Fable 5.1 | xhigh | 48.73% | $9.27 |
| Claude Fable 5.1 | max | 50.28% | $12.8 |
| Claude Opus 5 | low | 41.95% | $2.64 |
| Claude Opus 5 | medium | 53.38% | $4.61 |
| Claude Opus 5 | high | 47.99% | $7.62 |
| Claude Opus 5 | xhigh | 43.65% | $8.99 |
| Claude Opus 5 | max | 48.04% | $12.3 |
| GPT-6 Astra | low | 45.27% | $1.59 |
| GPT-6 Astra | medium | 48.83% | $2.28 |
| GPT-6 Astra | high | 50.94% | $2.85 |
| GPT-6 Astra | xhigh | 50.62% | $3.10 |
| GPT-6 Astra | max | 53.26% | $4.36 |
| GPT-5.6 Sol | low | 35.44% | $1.75 |
| GPT-5.6 Sol | medium | 39.93% | $2.50 |
| GPT-5.6 Sol | high | 45.06% | $3.25 |
| GPT-5.6 Sol | xhigh | 46.84% | $3.88 |
| GPT-5.6 Sol | max | 47.49% | $4.85 |
- Exact values from the data published with Anthropic's announcement.
Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values
Evaluates coding agents on ambiguous, multi-file tasks taken from real Cursor sessions.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
- GPT-5.6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 43.7% | $1.18 |
| Claude Opus 5.5 | medium | 52.5% | $2.90 |
| Claude Opus 5.5 | high | 56% | $3.97 |
| Claude Opus 5.5 | xhigh | 56% | $6.99 |
| Claude Opus 5.5 | max | 57.8% | $13.4 |
| Claude Fable 5.1 | low | 45.1% | $5.44 |
| Claude Fable 5.1 | medium | 46.8% | $7.05 |
| Claude Fable 5.1 | high | 49.2% | $9.08 |
| Claude Fable 5.1 | xhigh | 51.6% | $13 |
| Claude Fable 5.1 | max | 51.8% | $17.3 |
| Claude Opus 5 | low | 40.7% | $4.87 |
| Claude Opus 5 | medium | 43.3% | $6.94 |
| Claude Opus 5 | high | 44.7% | $9.00 |
| Claude Opus 5 | xhigh | 46.1% | $11.4 |
| Claude Opus 5 | max | 46.6% | $11.9 |
| GPT-5.6 Sol | low | 24.6% | $0.87 |
| GPT-5.6 Sol | medium | 31.1% | $1.77 |
| GPT-5.6 Sol | high | 35.7% | $2.85 |
| GPT-5.6 Sol | xhigh | 37.7% | $4.40 |
| GPT-5.6 Sol | max | 41.7% | $8.23 |
- Exact values from the data published with Anthropic's announcement.
Elo against cost · reported by Anthropic on 22 Sept 2026 · exact values
Artificial Analysis's evaluation of agents on real-world professional work across 44 occupations.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
- GPT-6 Astra
- GPT-5.6 Sol
Show data table
| Model | Effort | Elo | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 1224 Elo | $0.21 |
| Claude Opus 5.5 | medium | 1576 Elo | $0.86 |
| Claude Opus 5.5 | high | 1692 Elo | $1.54 |
| Claude Opus 5.5 | xhigh | 1820 Elo | $4.21 |
| Claude Opus 5.5 | max | 1846 Elo | $8.92 |
| Claude Fable 5.1 | low | 1450 Elo | $1.41 |
| Claude Fable 5.1 | medium | 1536 Elo | $2.17 |
| Claude Fable 5.1 | high | 1617 Elo | $3.43 |
| Claude Fable 5.1 | xhigh | 1721 Elo | $7.09 |
| Claude Fable 5.1 | max | 1735 Elo | $9.59 |
| Claude Opus 5 | low | 1294 Elo | $0.58 |
| Claude Opus 5 | medium | 1476 Elo | $1.36 |
| Claude Opus 5 | high | 1581 Elo | $3.03 |
| Claude Opus 5 | xhigh | 1676 Elo | $4.96 |
| Claude Opus 5 | max | 1708 Elo | $6.76 |
| GPT-6 Astra | low | 1366 Elo | $0.85 |
| GPT-6 Astra | medium | 1468 Elo | $1.82 |
| GPT-6 Astra | high | 1485 Elo | $2.43 |
| GPT-6 Astra | xhigh | 1516 Elo | $3.04 |
| GPT-6 Astra | max | 1542 Elo | $4.53 |
| GPT-5.6 Sol | low | 1289 Elo | $0.27 |
| GPT-5.6 Sol | medium | 1403 Elo | $0.60 |
| GPT-5.6 Sol | high | 1480 Elo | $1.11 |
| GPT-5.6 Sol | xhigh | 1548 Elo | $1.70 |
| GPT-5.6 Sol | max | 1588 Elo | $2.81 |
- Exact values from the data published with Anthropic's announcement.
Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values
Perplexity's benchmark measuring agents on large data collection tasks.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 31.2% | $1.20 |
| Claude Opus 5.5 | medium | 62.8% | $11.2 |
| Claude Opus 5.5 | high | 67.3% | $16.9 |
| Claude Opus 5.5 | xhigh | 71.3% | $29.1 |
| Claude Opus 5.5 | max | 72.3% | $37.9 |
| Claude Fable 5.1 | low | 63.3% | $23.3 |
| Claude Fable 5.1 | medium | 64.5% | $27.4 |
| Claude Fable 5.1 | high | 66.7% | $33.6 |
| Claude Fable 5.1 | xhigh | 67.7% | $42.8 |
| Claude Fable 5.1 | max | 68.7% | $49 |
| Claude Opus 5 | low | 50.5% | $10.7 |
| Claude Opus 5 | medium | 58.1% | $24 |
| Claude Opus 5 | high | 64.6% | $43.5 |
| Claude Opus 5 | xhigh | 67% | $53.8 |
| Claude Opus 5 | max | 67.2% | $61.6 |
- Exact values from the data published with Anthropic's announcement.
- Claude models ran with offline web search and fetch tools and a 980K-token task budget. This differs from Perplexity's published setup, so scores are not comparable with Perplexity's leaderboard.
Accuracy against API cost · reported by OpenAI on 9 Sept 2026 · exact values
Tests agents on complex terminal-based tasks, including software engineering, system configuration and data analysis.
- GPT-6 Astra
- GPT-5.6 Sol
- Claude Fable 5.1
- Claude Fable 5
- Claude Opus 5
Show data table
| Model | Effort | Accuracy | Cost per task |
|---|---|---|---|
| GPT-6 Astra | low | 49.7% | $4.95 |
| GPT-6 Astra | medium | 53.9% | $6.15 |
| GPT-6 Astra | high | 57.9% | $7.21 |
| GPT-6 Astra | xhigh | 57.6% | $7.48 |
| GPT-6 Astra | max | 56.7% | $10.3 |
| GPT-5.6 Sol | low | 7.9% | $1.46 |
| GPT-5.6 Sol | medium | 20.9% | $2.69 |
| GPT-5.6 Sol | high | 26.1% | $4.12 |
| GPT-5.6 Sol | xhigh | 28.5% | $5.39 |
| GPT-5.6 Sol | max | 37.3% | $7.89 |
| Claude Fable 5.1 | low | 40.2% | $5.70 |
| Claude Fable 5.1 | medium | 43.4% | $7.80 |
| Claude Fable 5.1 | high | 49.4% | $10.5 |
| Claude Fable 5.1 | xhigh | 51.3% | $15.8 |
| Claude Fable 5.1 | max | 55.8% | $19.5 |
| Claude Fable 5 | max | 44.5% | $22.2 |
| Claude Opus 5 | max | 52.6% | $18.4 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
- Claude Fable 5 and Claude Opus 5 are shown at max effort only, as published.