CheckNet.NETWORK DIAGNOSTICS / TOOLKIT
WorkspaceAI Model TrackerNETWORK TOOLKIT
All models

OpenAI · launched 4 Sept 2026

GPT-6 Astra

OpenAI's flagship GPT-6 model, positioned above GPT-6 Sol and Luna for the most demanding professional work, coding and computer use.

Read the official announcement
Model details
Vendor
OpenAI
Launch date
4 Sept 2026
Input price
$10 / 1M tokens
Output price
$50 / 1M tokens
Cached input
$1.00 / 1M tokens
Batch discount
50%
Context window
1.05M tokens
Max output
128K tokens
Input types
Text, Image, File
OpenRouter ID
openai/gpt-6-astra
Best published results
  • Terminal-Bench 4.0high effort · $7.21 per task · reported by OpenAI57.9%
  • AutomationBench 1.0.6max effort · $1.73 per task · reported by OpenAI41.4%
  • Agents' Last Exam V1max effort · $6.23 per task · reported by OpenAI59.3%
  • Factual error rate(lower is better)high effort · $0.48 per task · reported by OpenAI3.9%
  • FrontierCode 1.1max effort · $4.59 per task · reported by OpenAI53.3%
  • DeepSWE 1.1xhigh effort · $4.43 per task · reported by OpenAI74.1%
  • OSWorld 2.0max effort · $9.07 per task · reported by OpenAI73.5%
  • Terminal-Bench 4.0high effort · $7.21 per attempt · reported by Anthropic57.9%
  • FrontierCode v1.1max effort · $4.36 per task · reported by Anthropic53.26%
  • GDPval-AA v2.1max effort · $4.53 per task · reported by Anthropic1542 Elo
  • AutomationBenchmax effort · $1.77 per task · reported by Anthropic41.4%

Price compared with other models

ModelLaunchedInput / 1MOutput / 1MCached / 1MBlended / 1Mvs GPT-6 Astra
GPT-6 LunaOpenAI22 Sept 2026$0.10$0.50$0.01$0.2099% cheaper
GPT-5.6 LunaOpenAI9 Jul 2026$0.20$1.20$0.02$0.4598% cheaper
Claude Sonnet 5Anthropic30 Jun 2026$2.00$10$0.20$4.0080% cheaper
Claude Sonnet 5.5Anthropic28 Sept 2026$2.00$10$0.20$4.0080% cheaper
GPT-6 SolOpenAI22 Sept 2026$2.00$10$0.20$4.0080% cheaper
Claude Opus 5.5Anthropic22 Sept 2026$4.00$20$0.20$8.0060% cheaper
GPT-5.6 SolOpenAIPrice before the GPT-6 launch, from OpenAI's announcement. OpenRouter now lists $2 / $10.9 Jul 2026$4.00$20–$8.0060% cheaper
Claude Opus 5Anthropic24 Jul 2026$5.00$25$0.50$1050% cheaper
Claude Fable 5Anthropic9 Jun 2026$10$50$1.00$20Same price
Claude Fable 5.1Anthropic1 Sept 2026$10$50$0.25$20Same price
GPT-6 AstraOpenAI4 Sept 2026$10$50$1.00$20Baseline

USD per million tokens, checked 2 Oct 2026 via OpenRouter. Comparison uses the blended price (3 input tokens for every output token).

Official benchmark results

GPT-6 Sol and Luna alignment evaluation

Reported by OpenAI on 22 Sept 2026 · exact values

BenchmarkGPT-6 AstraGPT-6 SolGPT-6 LunaGPT-5.6 SolGPT-5.6 Luna
Coding deception rateAlignment · max effort · lower is better0.5%1.3%2.8%10.4%9.5%
  • Bold marks the best reported score in each row.
  • OpenAI's internal evaluation uses tasks deliberately chosen to provoke dishonesty, so deception is much rarer in typical use.
  • The announcement also covers broken search, reviewer bypass, warning circumvention and unauthorized interaction. See OpenAI's system card for those results.
Source: Introducing GPT-6 Sol and Luna (OpenAI)
Claude Opus 5.5 launch results

Reported by Anthropic on 22 Sept 2026 · exact values

BenchmarkClaude Opus 5.5Claude Fable 5.1Claude Opus 5GPT-6 AstraGPT-5.6 Sol
Terminal-Bench 4.0Agentic coding66.4%55.8%52.3%57.9%37.3%
FrontierCode v1.1 (main)Agentic coding54.4%50.3%48%53.3%47.5%
CursorBench 4.0Agentic coding57.8%51.8%46.6%–41.7%
GDPval-AA v2.1Knowledge work1846 Elo1735 Elo1708 Elo1542 Elo1588 Elo
AutomationBenchBusiness workflows40%31.4%26.9%41.4%28.8%
Humanity's Last ExamMultidisciplinary reasoning · with tools67.7%65.6%63.6%57.2%–
Terminal-Bench-Science 0.1Agentic scientific research58.7%52.6%29%64.6%22.4%
OSWorld 2.1Computer use · partial81.8%80.7%74%––
ChartographyVisual chart recognition · with tools89%88.4%83.4%––
  • Bold marks the best reported score in each row.
  • Claude Opus 5.5 results use adaptive thinking at max effort unless noted.
  • Terminal-Bench 4.0 shows each model's highest score: Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort. GPT-6 Astra and GPT-5.6 Sol figures there are as reported by OpenAI.
  • A dash means the lab did not report a result for that model.
Source: Introducing Claude Opus 5.5 (Anthropic)
Terminal-Bench 4.0

Accuracy against API cost · reported by OpenAI on 9 Sept 2026 · exact values

Tests agents on complex terminal-based tasks, including software engineering, system configuration and data analysis.

  • GPT-6 Astra
  • GPT-5.6 Sol
  • Claude Fable 5.1
  • Claude Fable 5
  • Claude Opus 5
Show data table
ModelEffortAccuracyCost per task
GPT-6 Astralow49.7%$4.95
GPT-6 Astramedium53.9%$6.15
GPT-6 Astrahigh57.9%$7.21
GPT-6 Astraxhigh57.6%$7.48
GPT-6 Astramax56.7%$10.3
GPT-5.6 Sollow7.9%$1.46
GPT-5.6 Solmedium20.9%$2.69
GPT-5.6 Solhigh26.1%$4.12
GPT-5.6 Solxhigh28.5%$5.39
GPT-5.6 Solmax37.3%$7.89
Claude Fable 5.1low40.2%$5.70
Claude Fable 5.1medium43.4%$7.80
Claude Fable 5.1high49.4%$10.5
Claude Fable 5.1xhigh51.3%$15.8
Claude Fable 5.1max55.8%$19.5
Claude Fable 5max44.5%$22.2
Claude Opus 5max52.6%$18.4
  • Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
  • Claude Fable 5 and Claude Opus 5 are shown at max effort only, as published.
Source: GPT-6 Astra: The next generation in intelligence for work (OpenAI)
AutomationBench 1.0.6

Score against cost · reported by OpenAI on 22 Sept 2026 · exact values

AI agents are tested on end-to-end workflows using 47 tools across sales, marketing, operations, support, finance and HR.

  • GPT-6 Luna
  • GPT-5.6 Luna
  • GPT-6 Sol
  • GPT-5.6 Sol
  • GPT-6 Astra
  • Claude Opus 5
  • Claude Fable 5.1
Show data table
ModelEffortScoreCost per task
GPT-6 Lunalow1.2%$0.006
GPT-6 Lunamedium9.4%$0.016
GPT-6 Lunahigh14.5%$0.021
GPT-6 Lunaxhigh12.6%$0.025
GPT-6 Lunamax20.7%$0.037
GPT-5.6 Lunalow1.8%$0.01
GPT-5.6 Lunamedium4.3%$0.02
GPT-5.6 Lunahigh9.1%$0.05
GPT-5.6 Lunaxhigh12.9%$0.06
GPT-5.6 Lunamax17%$0.07
GPT-6 Sollow21.2%$0.19
GPT-6 Solmedium26.9%$0.21
GPT-6 Solhigh31.2%$0.24
GPT-6 Solxhigh33.2%$0.27
GPT-6 Solmax32%$0.34
GPT-5.6 Sollow11.7%$0.31
GPT-5.6 Solmedium19.6%$0.42
GPT-5.6 Solhigh24.8%$0.47
GPT-5.6 Solxhigh26.3%$0.54
GPT-5.6 Solmax28.8%$0.67
GPT-6 Astralow30.3%$1.08
GPT-6 Astramedium34.1%$1.27
GPT-6 Astrahigh37.1%$1.44
GPT-6 Astraxhigh39%$1.50
GPT-6 Astramax41.4%$1.73
Claude Opus 5low20.4%$1.64
Claude Opus 5medium23.9%$2.22
Claude Opus 5high20.5%$2.27
Claude Opus 5xhigh25.3%$2.71
Claude Opus 5max26.9%$3.05
Claude Fable 5.1max31.4%$2.45
  • Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
  • Claude Fable 5.1 ran with Claude Opus 5 as a fallback. OpenAI notes its cost omits the fallback runs (about 40% of tasks), so its real cost is higher.
Source: Introducing GPT-6 Sol and Luna (OpenAI)
Agents' Last Exam V1

Score against cost · reported by OpenAI on 22 Sept 2026 · exact values

AI agents are evaluated on long-horizon, economically valuable tasks spanning 55 sub-industries of professional computer work.

  • GPT-6 Luna
  • GPT-5.6 Luna
  • GPT-6 Sol
  • GPT-5.6 Sol
  • GPT-6 Astra
  • Claude Opus 5
  • Claude Fable 5
Show data table
ModelEffortScoreCost per task
GPT-6 Lunalow36.3%$0.025
GPT-6 Lunamedium46.8%$0.11
GPT-6 Lunahigh43.6%$0.11
GPT-6 Lunaxhigh47.9%$0.11
GPT-6 Lunamax50.9%$0.15
GPT-5.6 Lunalow31%$0.18
GPT-5.6 Lunamedium38.5%$0.38
GPT-5.6 Lunahigh46.1%$0.84
GPT-5.6 Lunaxhigh49.4%$1.55
GPT-5.6 Lunamax50.4%$2.57
GPT-6 Sollow48.7%$0.86
GPT-6 Solmedium53.1%$1.27
GPT-6 Solhigh52.6%$1.53
GPT-6 Solxhigh55.4%$1.67
GPT-6 Solmax56.4%$2.93
GPT-5.6 Sollow45.1%$1.63
GPT-5.6 Solmedium52.1%$3.41
GPT-5.6 Solhigh52.4%$3.79
GPT-5.6 Solxhigh53.6%$5.08
GPT-5.6 Solmax52.8%$7.13
GPT-6 Astralow53.4%$3.07
GPT-6 Astramedium57.6%$4.10
GPT-6 Astrahigh57.8%$4.64
GPT-6 Astraxhigh58.3%$5.40
GPT-6 Astramax59.3%$6.23
Claude Opus 5low51.9%$4.06
Claude Opus 5medium53%$5.29
Claude Opus 5high55.9%$7.29
Claude Opus 5xhigh55.5%$10
Claude Opus 5max52.7%$9.76
Claude Fable 5xhigh48.7%$28.6
Claude Fable 5adaptive41.3%$15.2
  • Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
  • Claude Fable 5 ran with Claude Opus 4.8 as a fallback.
Source: Introducing GPT-6 Sol and Luna (OpenAI)
Factual error rate on difficult prompts

Answers with any factual error against cost · lower is better · reported by OpenAI on 22 Sept 2026 · exact values

OpenAI's internal evaluation on de-identified ChatGPT conversations where users had flagged a factual error from a prior model.

  • GPT-6 Luna
  • GPT-5.6 Luna
  • GPT-6 Sol
  • GPT-5.6 Sol
  • GPT-6 Astra
Show data table
ModelEffortError rateCost per task
GPT-6 Lunalow27.7%$0.0024
GPT-6 Lunamedium17.5%$0.0033
GPT-6 Lunahigh12.5%$0.0045
GPT-6 Lunaxhigh10.2%$0.0062
GPT-6 Lunamax7.6%$0.012
GPT-5.6 Lunalow36.8%$0.0045
GPT-5.6 Lunamedium24.7%$0.0062
GPT-5.6 Lunahigh17.3%$0.011
GPT-5.6 Lunaxhigh12.2%$0.015
GPT-5.6 Lunamax12%$0.022
GPT-6 Sollow11.4%$0.05
GPT-6 Solmedium6.9%$0.069
GPT-6 Solhigh5.1%$0.099
GPT-6 Solxhigh4.5%$0.13
GPT-6 Solmax4.6%$0.18
GPT-5.6 Sollow19.4%$0.10
GPT-5.6 Solmedium15%$0.15
GPT-5.6 Solhigh10.8%$0.23
GPT-5.6 Solxhigh8.4%$0.39
GPT-5.6 Solmax8.5%$0.87
GPT-6 Astralow6.3%$0.24
GPT-6 Astramedium4.4%$0.31
GPT-6 Astrahigh3.9%$0.48
GPT-6 Astraxhigh4%$0.60
GPT-6 Astramax3.9%$0.79
  • Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
  • These prompts were chosen because they caused errors before, so error rates are far higher than in typical use.
Source: Introducing GPT-6 Sol and Luna (OpenAI)
FrontierCode 1.1, main set

Score against cost · reported by OpenAI on 22 Sept 2026 · exact values

AI agents write code graded on correctness and mergeability: test quality, scope discipline, code style and codebase standards.

  • GPT-6 Luna
  • GPT-5.6 Luna
  • GPT-6 Sol
  • GPT-5.6 Sol
  • GPT-6 Astra
  • Claude Opus 5
  • Claude Fable 5.1
Show data table
ModelEffortScoreCost per task
GPT-6 Lunalow25.7%$0.021
GPT-6 Lunamedium35.5%$0.053
GPT-6 Lunahigh37.3%$0.067
GPT-6 Lunaxhigh37.1%$0.073
GPT-6 Lunamax42.4%$0.11
GPT-5.6 Lunalow15.4%$0.06
GPT-5.6 Lunamedium25.7%$0.13
GPT-5.6 Lunahigh35.9%$0.23
GPT-5.6 Lunaxhigh38.9%$0.31
GPT-5.6 Lunamax39.8%$0.37
GPT-6 Sollow37.3%$0.45
GPT-6 Solmedium45.9%$0.80
GPT-6 Solhigh47.7%$1.08
GPT-6 Solxhigh48.4%$1.37
GPT-6 Solmax49.3%$2.14
GPT-5.6 Sollow35.4%$1.89
GPT-5.6 Solmedium39.9%$2.69
GPT-5.6 Solhigh45.1%$3.48
GPT-5.6 Solxhigh46.8%$4.15
GPT-5.6 Solmax47.5%$5.19
GPT-6 Astralow45.3%$1.70
GPT-6 Astramedium48.8%$2.43
GPT-6 Astrahigh50.9%$3.01
GPT-6 Astraxhigh50.6%$3.28
GPT-6 Astramax53.3%$4.59
Claude Opus 5low41.9%$2.68
Claude Opus 5medium53.4%$4.31
Claude Opus 5high48%$7.24
Claude Opus 5xhigh43.6%$9.14
Claude Opus 5max48%$11.4
Claude Fable 5.1low49.8%$2.38
Claude Fable 5.1medium50.9%$3.28
Claude Fable 5.1high50.3%$5.27
Claude Fable 5.1xhigh48.7%$9.27
Claude Fable 5.1max50.3%$12.8
  • Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
Source: Introducing GPT-6 Sol and Luna (OpenAI)
DeepSWE 1.1

Score against cost · reported by OpenAI on 22 Sept 2026 · exact values

AI agents solve original, long-horizon software engineering tasks in real codebases.

  • GPT-6 Luna
  • GPT-5.6 Luna
  • GPT-6 Sol
  • GPT-5.6 Sol
  • GPT-6 Astra
  • Claude Opus 5
  • Claude Fable 5
Show data table
ModelEffortScoreCost per task
GPT-6 Lunalow2.4%$0.0057
GPT-6 Lunamedium44.5%$0.052
GPT-6 Lunahigh59.3%$0.084
GPT-6 Lunaxhigh61.3%$0.11
GPT-6 Lunamax66.6%$0.22
GPT-5.6 Lunalow1.2%$0.011
GPT-5.6 Lunamedium9.3%$0.031
GPT-5.6 Lunahigh42.4%$0.13
GPT-5.6 Lunaxhigh56.2%$0.27
GPT-5.6 Lunamax62.2%$0.53
GPT-6 Sollow37.2%$0.16
GPT-6 Solmedium56.6%$0.38
GPT-6 Solhigh65.3%$0.64
GPT-6 Solxhigh66.6%$1.00
GPT-6 Solmax68.8%$2.74
GPT-5.6 Sollow45.4%$0.82
GPT-5.6 Solmedium61.1%$1.42
GPT-5.6 Solhigh69.4%$2.66
GPT-5.6 Solxhigh70.7%$3.60
GPT-5.6 Solmax72.7%$6.46
GPT-6 Astralow67%$1.60
GPT-6 Astramedium72.8%$3.08
GPT-6 Astrahigh73.2%$3.92
GPT-6 Astraxhigh74.1%$4.43
GPT-6 Astramax73.2%$7.50
Claude Opus 5low58.1%$1.66
Claude Opus 5medium68.9%$3.29
Claude Opus 5high72.8%$6.08
Claude Opus 5xhigh73.2%$9.07
Claude Opus 5max73.7%$11.8
Claude Fable 5low59.6%$3.76
Claude Fable 5medium65.4%$6.09
Claude Fable 5high68.6%$9.18
Claude Fable 5xhigh69.9%$13.4
Claude Fable 5max69.7%$21.6
  • Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
  • OpenAI used Claude Fable 5 scores because Fable 5.1 scores were unavailable.
Source: Introducing GPT-6 Sol and Luna (OpenAI)
OSWorld 2.0, offline set

Partial reward against cost · release v2026.08.08 · reported by OpenAI on 22 Sept 2026 · exact values

AI agents attempt long-horizon computer-use workflows covering everyday and professional tasks.

  • GPT-6 Luna
  • GPT-5.6 Luna
  • GPT-6 Sol
  • GPT-5.6 Sol
  • GPT-6 Astra
  • Claude Opus 5
Show data table
ModelEffortScoreCost per task
GPT-6 Lunalow8.3%$0.03
GPT-6 Lunamedium31.5%$0.062
GPT-6 Lunahigh41.4%$0.12
GPT-6 Lunaxhigh46.7%$0.16
GPT-6 Lunamax52.7%$0.27
GPT-5.6 Lunalow12%$0.019
GPT-5.6 Lunamedium22.7%$0.052
GPT-5.6 Lunahigh35.9%$0.16
GPT-5.6 Lunaxhigh48%$0.36
GPT-5.6 Lunamax52.7%$0.49
GPT-6 Sollow43.9%$0.97
GPT-6 Solmedium54%$1.32
GPT-6 Solhigh58.3%$1.64
GPT-6 Solxhigh60.5%$2.21
GPT-6 Solmax64.4%$3.25
GPT-5.6 Sollow29.8%$0.91
GPT-5.6 Solmedium49.7%$2.73
GPT-5.6 Solhigh56.5%$4.46
GPT-5.6 Solxhigh60.9%$5.93
GPT-5.6 Solmax66.2%$7.71
GPT-6 Astralow62.2%$2.55
GPT-6 Astramedium69.3%$5.10
GPT-6 Astrahigh70%$6.60
GPT-6 Astraxhigh71.3%$7.17
GPT-6 Astramax73.5%$9.07
Claude Opus 5low55.2%$9.88
Claude Opus 5medium60.3%$12.7
Claude Opus 5high65.9%$15.9
Claude Opus 5xhigh70.1%$23.9
Claude Opus 5max70.2%$24.1
  • Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
  • Claude Opus 5 values match the official OSWorld 2.0 leaderboard (osworld-v2.xlang.ai).
Source: Introducing GPT-6 Sol and Luna (OpenAI)
Terminal-Bench 4.0

Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values

Measures how well a model completes complex, multi-step professional tasks within a command line interface.

  • Claude Opus 5.5
  • Claude Fable 5.1
  • Claude Opus 5
  • GPT-6 Astra
  • GPT-5.6 Sol
Show data table
ModelEffortScoreCost per task
Claude Opus 5.5low38.5%$1.29
Claude Opus 5.5medium57.6%$2.94
Claude Opus 5.5high64.2%$3.88
Claude Opus 5.5xhigh66.4%$7.35
Claude Opus 5.5max64.8%$11.2
Claude Fable 5.1low40.2%$5.70
Claude Fable 5.1medium43.4%$7.80
Claude Fable 5.1high49.4%$10.5
Claude Fable 5.1xhigh51.3%$15.8
Claude Fable 5.1max55.8%$19.5
Claude Opus 5low28.5%$4.25
Claude Opus 5medium41.2%$7.00
Claude Opus 5high47%$10.6
Claude Opus 5xhigh50.6%$13.5
Claude Opus 5max52.3%$15.8
GPT-6 Astralow49.7%$4.95
GPT-6 Astramedium53.9%$6.15
GPT-6 Astrahigh57.9%$7.21
GPT-6 Astraxhigh57.6%$7.48
GPT-6 Astramax56.7%$10.3
GPT-5.6 Sollow7.9%$1.46
GPT-5.6 Solmedium20.9%$2.69
GPT-5.6 Solhigh26.1%$4.12
GPT-5.6 Solxhigh28.5%$5.39
GPT-5.6 Solmax37.3%$7.89
  • Exact values from the data published with Anthropic's announcement.
  • GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI.
Source: Introducing Claude Opus 5.5 (Anthropic)
FrontierCode v1.1, main set

Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values

Measures whether an agent's code changes would be merged.

  • Claude Opus 5.5
  • Claude Fable 5.1
  • Claude Opus 5
  • GPT-6 Astra
  • GPT-5.6 Sol
Show data table
ModelEffortScoreCost per task
Claude Opus 5.5low47.3%$0.40
Claude Opus 5.5medium54.64%$0.80
Claude Opus 5.5high53.99%$1.09
Claude Opus 5.5xhigh51.42%$2.25
Claude Opus 5.5max54.43%$6.19
Claude Fable 5.1low52.8%$2.47
Claude Fable 5.1medium50.91%$3.28
Claude Fable 5.1high50.34%$5.27
Claude Fable 5.1xhigh48.73%$9.27
Claude Fable 5.1max50.28%$12.8
Claude Opus 5low41.95%$2.64
Claude Opus 5medium53.38%$4.61
Claude Opus 5high47.99%$7.62
Claude Opus 5xhigh43.65%$8.99
Claude Opus 5max48.04%$12.3
GPT-6 Astralow45.27%$1.59
GPT-6 Astramedium48.83%$2.28
GPT-6 Astrahigh50.94%$2.85
GPT-6 Astraxhigh50.62%$3.10
GPT-6 Astramax53.26%$4.36
GPT-5.6 Sollow35.44%$1.75
GPT-5.6 Solmedium39.93%$2.50
GPT-5.6 Solhigh45.06%$3.25
GPT-5.6 Solxhigh46.84%$3.88
GPT-5.6 Solmax47.49%$4.85
  • Exact values from the data published with Anthropic's announcement.
Source: Introducing Claude Opus 5.5 (Anthropic)
GDPval-AA v2.1

Elo against cost · reported by Anthropic on 22 Sept 2026 · exact values

Artificial Analysis's evaluation of agents on real-world professional work across 44 occupations.

  • Claude Opus 5.5
  • Claude Fable 5.1
  • Claude Opus 5
  • GPT-6 Astra
  • GPT-5.6 Sol
Show data table
ModelEffortEloCost per task
Claude Opus 5.5low1224 Elo$0.21
Claude Opus 5.5medium1576 Elo$0.86
Claude Opus 5.5high1692 Elo$1.54
Claude Opus 5.5xhigh1820 Elo$4.21
Claude Opus 5.5max1846 Elo$8.92
Claude Fable 5.1low1450 Elo$1.41
Claude Fable 5.1medium1536 Elo$2.17
Claude Fable 5.1high1617 Elo$3.43
Claude Fable 5.1xhigh1721 Elo$7.09
Claude Fable 5.1max1735 Elo$9.59
Claude Opus 5low1294 Elo$0.58
Claude Opus 5medium1476 Elo$1.36
Claude Opus 5high1581 Elo$3.03
Claude Opus 5xhigh1676 Elo$4.96
Claude Opus 5max1708 Elo$6.76
GPT-6 Astralow1366 Elo$0.85
GPT-6 Astramedium1468 Elo$1.82
GPT-6 Astrahigh1485 Elo$2.43
GPT-6 Astraxhigh1516 Elo$3.04
GPT-6 Astramax1542 Elo$4.53
GPT-5.6 Sollow1289 Elo$0.27
GPT-5.6 Solmedium1403 Elo$0.60
GPT-5.6 Solhigh1480 Elo$1.11
GPT-5.6 Solxhigh1548 Elo$1.70
GPT-5.6 Solmax1588 Elo$2.81
  • Exact values from the data published with Anthropic's announcement.
Source: Introducing Claude Opus 5.5 (Anthropic)
AutomationBench

Pass rate against cost · reported by Anthropic on 22 Sept 2026 · exact values

Built by Zapier, tests whether an agent can carry out real business workflows across many connected apps.

  • Claude Opus 5.5
  • Claude Opus 5
  • GPT-6 Astra
  • GPT-5.6 Sol
Show data table
ModelEffortPass rateCost per task
Claude Opus 5.5low23.3%$0.50
Claude Opus 5.5medium28.6%$0.64
Claude Opus 5.5high32%$0.70
Claude Opus 5.5xhigh34.4%$0.86
Claude Opus 5.5max40%$1.37
Claude Opus 5low20.4%$0.75
Claude Opus 5medium23.9%$0.89
Claude Opus 5high20.6%$1.03
Claude Opus 5xhigh25.3%$1.15
Claude Opus 5max26.9%$1.27
GPT-6 Astralow30.3%$1.08
GPT-6 Astramedium34.1%$1.28
GPT-6 Astrahigh37.1%$1.45
GPT-6 Astraxhigh39%$1.53
GPT-6 Astramax41.4%$1.77
GPT-5.6 Sollow11.7%$0.42
GPT-5.6 Solmedium19.6%$0.57
GPT-5.6 Solhigh24.8%$0.64
GPT-5.6 Solxhigh26.3%$0.74
GPT-5.6 Solmax28.8%$0.91
  • Exact values from the data published with Anthropic's announcement.
  • Runs were performed by Zapier without fallback models, so safeguard interventions counted as failures.
Source: Introducing Claude Opus 5.5 (Anthropic)