CheckNet.NETWORK DIAGNOSTICS / TOOLKIT
WorkspaceAI Model TrackerNETWORK TOOLKIT

AI Model Tracker

Recent model launches with their API prices and the benchmark results published by each lab, linked back to the official announcement.

Latest releases

Anthropic · launched 28 Sept 2026

Claude Sonnet 5.5

The second model in the Claude 5.5 family. Priced the same as Sonnet 5, it runs more than 30% faster and needs fewer tokens per task.

Input / 1M
$2.00
Output / 1M
$10
Context
1M

Anthropic · launched 22 Sept 2026

Claude Opus 5.5

Anthropic's newest Opus model. Tokens cost 20% less than Opus 5 and cache reads 60% less, with output generated more than 30% faster.

Input / 1M
$4.00
Output / 1M
$20
Context
1M

OpenAI · launched 22 Sept 2026

GPT-6 Luna

The lowest-cost model in the GPT-6 family, positioned below GPT-6 Sol and the flagship GPT-6 Astra.

Input / 1M
$0.10
Output / 1M
$0.50
Context
1.05M

OpenAI · launched 22 Sept 2026

GPT-6 Sol

The mid-tier GPT-6 model, sitting between GPT-6 Luna and the flagship GPT-6 Astra.

Input / 1M
$2.00
Output / 1M
$10
Context
1.05M

OpenAI · launched 4 Sept 2026

GPT-6 Astra

OpenAI's flagship GPT-6 model, positioned above GPT-6 Sol and Luna for the most demanding professional work, coding and computer use.

Input / 1M
$10
Output / 1M
$50
Context
1.05M

Anthropic · launched 1 Sept 2026

Claude Fable 5.1

Released alongside the access-limited Claude Mythos 5.1. Cache reads cost 75% less than Fable 5, which Anthropic estimates cuts typical costs by about 25%.

Input / 1M
$10
Output / 1M
$50
Context
1M

API pricing

ModelLaunchedInput / 1MOutput / 1MCached / 1MBlended / 1M
Claude Sonnet 5.5Anthropic28 Sept 2026$2.00$10$0.20$4.00
Claude Opus 5.5Anthropic22 Sept 2026$4.00$20$0.20$8.00
GPT-6 LunaOpenAI22 Sept 2026$0.10$0.50$0.01$0.20
GPT-6 SolOpenAI22 Sept 2026$2.00$10$0.20$4.00
GPT-6 AstraOpenAI4 Sept 2026$10$50$1.00$20
Claude Fable 5.1Anthropic1 Sept 2026$10$50$0.25$20
Claude Opus 5Anthropic24 Jul 2026$5.00$25$0.50$10
GPT-5.6 LunaOpenAI9 Jul 2026$0.20$1.20$0.02$0.45
GPT-5.6 SolOpenAIPrice before the GPT-6 launch, from OpenAI's announcement. OpenRouter now lists $2 / $10.9 Jul 2026$4.00$20–$8.00
Claude Sonnet 5Anthropic30 Jun 2026$2.00$10$0.20$4.00
Claude Fable 5Anthropic9 Jun 2026$10$50$1.00$20

Standard API prices in USD per million tokens, checked 2 Oct 2026 via OpenRouter. Blended assumes 3 input tokens for every output token. Batch requests are 50% cheaper for every model listed.

Official benchmark results

OpenAI: Introducing GPT-6 Sol and LunaRead announcement

GPT-6 Sol and Luna alignment evaluation

Reported by OpenAI on 22 Sept 2026 · exact values

BenchmarkGPT-6 AstraGPT-6 SolGPT-6 LunaGPT-5.6 SolGPT-5.6 Luna
Coding deception rateAlignment · max effort · lower is better0.5%1.3%2.8%10.4%9.5%
  • Bold marks the best reported score in each row.
  • OpenAI's internal evaluation uses tasks deliberately chosen to provoke dishonesty, so deception is much rarer in typical use.
  • The announcement also covers broken search, reviewer bypass, warning circumvention and unauthorized interaction. See OpenAI's system card for those results.
Source: Introducing GPT-6 Sol and Luna (OpenAI)
AutomationBench 1.0.6

Score against cost · reported by OpenAI on 22 Sept 2026 · exact values

AI agents are tested on end-to-end workflows using 47 tools across sales, marketing, operations, support, finance and HR.

  • GPT-6 Luna
  • GPT-5.6 Luna
  • GPT-6 Sol
  • GPT-5.6 Sol
  • GPT-6 Astra
  • Claude Opus 5
  • Claude Fable 5.1
Show data table
ModelEffortScoreCost per task
GPT-6 Lunalow1.2%$0.006
GPT-6 Lunamedium9.4%$0.016
GPT-6 Lunahigh14.5%$0.021
GPT-6 Lunaxhigh12.6%$0.025
GPT-6 Lunamax20.7%$0.037
GPT-5.6 Lunalow1.8%$0.01
GPT-5.6 Lunamedium4.3%$0.02
GPT-5.6 Lunahigh9.1%$0.05
GPT-5.6 Lunaxhigh12.9%$0.06
GPT-5.6 Lunamax17%$0.07
GPT-6 Sollow21.2%$0.19
GPT-6 Solmedium26.9%$0.21
GPT-6 Solhigh31.2%$0.24
GPT-6 Solxhigh33.2%$0.27
GPT-6 Solmax32%$0.34
GPT-5.6 Sollow11.7%$0.31
GPT-5.6 Solmedium19.6%$0.42
GPT-5.6 Solhigh24.8%$0.47
GPT-5.6 Solxhigh26.3%$0.54
GPT-5.6 Solmax28.8%$0.67
GPT-6 Astralow30.3%$1.08
GPT-6 Astramedium34.1%$1.27
GPT-6 Astrahigh37.1%$1.44
GPT-6 Astraxhigh39%$1.50
GPT-6 Astramax41.4%$1.73
Claude Opus 5low20.4%$1.64
Claude Opus 5medium23.9%$2.22
Claude Opus 5high20.5%$2.27
Claude Opus 5xhigh25.3%$2.71
Claude Opus 5max26.9%$3.05
Claude Fable 5.1max31.4%$2.45
  • Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
  • Claude Fable 5.1 ran with Claude Opus 5 as a fallback. OpenAI notes its cost omits the fallback runs (about 40% of tasks), so its real cost is higher.
Source: Introducing GPT-6 Sol and Luna (OpenAI)
Agents' Last Exam V1

Score against cost · reported by OpenAI on 22 Sept 2026 · exact values

AI agents are evaluated on long-horizon, economically valuable tasks spanning 55 sub-industries of professional computer work.

  • GPT-6 Luna
  • GPT-5.6 Luna
  • GPT-6 Sol
  • GPT-5.6 Sol
  • GPT-6 Astra
  • Claude Opus 5
  • Claude Fable 5
Show data table
ModelEffortScoreCost per task
GPT-6 Lunalow36.3%$0.025
GPT-6 Lunamedium46.8%$0.11
GPT-6 Lunahigh43.6%$0.11
GPT-6 Lunaxhigh47.9%$0.11
GPT-6 Lunamax50.9%$0.15
GPT-5.6 Lunalow31%$0.18
GPT-5.6 Lunamedium38.5%$0.38
GPT-5.6 Lunahigh46.1%$0.84
GPT-5.6 Lunaxhigh49.4%$1.55
GPT-5.6 Lunamax50.4%$2.57
GPT-6 Sollow48.7%$0.86
GPT-6 Solmedium53.1%$1.27
GPT-6 Solhigh52.6%$1.53
GPT-6 Solxhigh55.4%$1.67
GPT-6 Solmax56.4%$2.93
GPT-5.6 Sollow45.1%$1.63
GPT-5.6 Solmedium52.1%$3.41
GPT-5.6 Solhigh52.4%$3.79
GPT-5.6 Solxhigh53.6%$5.08
GPT-5.6 Solmax52.8%$7.13
GPT-6 Astralow53.4%$3.07
GPT-6 Astramedium57.6%$4.10
GPT-6 Astrahigh57.8%$4.64
GPT-6 Astraxhigh58.3%$5.40
GPT-6 Astramax59.3%$6.23
Claude Opus 5low51.9%$4.06
Claude Opus 5medium53%$5.29
Claude Opus 5high55.9%$7.29
Claude Opus 5xhigh55.5%$10
Claude Opus 5max52.7%$9.76
Claude Fable 5xhigh48.7%$28.6
Claude Fable 5adaptive41.3%$15.2
  • Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
  • Claude Fable 5 ran with Claude Opus 4.8 as a fallback.
Source: Introducing GPT-6 Sol and Luna (OpenAI)
Factual error rate on difficult prompts

Answers with any factual error against cost · lower is better · reported by OpenAI on 22 Sept 2026 · exact values

OpenAI's internal evaluation on de-identified ChatGPT conversations where users had flagged a factual error from a prior model.

  • GPT-6 Luna
  • GPT-5.6 Luna
  • GPT-6 Sol
  • GPT-5.6 Sol
  • GPT-6 Astra
Show data table
ModelEffortError rateCost per task
GPT-6 Lunalow27.7%$0.0024
GPT-6 Lunamedium17.5%$0.0033
GPT-6 Lunahigh12.5%$0.0045
GPT-6 Lunaxhigh10.2%$0.0062
GPT-6 Lunamax7.6%$0.012
GPT-5.6 Lunalow36.8%$0.0045
GPT-5.6 Lunamedium24.7%$0.0062
GPT-5.6 Lunahigh17.3%$0.011
GPT-5.6 Lunaxhigh12.2%$0.015
GPT-5.6 Lunamax12%$0.022
GPT-6 Sollow11.4%$0.05
GPT-6 Solmedium6.9%$0.069
GPT-6 Solhigh5.1%$0.099
GPT-6 Solxhigh4.5%$0.13
GPT-6 Solmax4.6%$0.18
GPT-5.6 Sollow19.4%$0.10
GPT-5.6 Solmedium15%$0.15
GPT-5.6 Solhigh10.8%$0.23
GPT-5.6 Solxhigh8.4%$0.39
GPT-5.6 Solmax8.5%$0.87
GPT-6 Astralow6.3%$0.24
GPT-6 Astramedium4.4%$0.31
GPT-6 Astrahigh3.9%$0.48
GPT-6 Astraxhigh4%$0.60
GPT-6 Astramax3.9%$0.79
  • Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
  • These prompts were chosen because they caused errors before, so error rates are far higher than in typical use.
Source: Introducing GPT-6 Sol and Luna (OpenAI)
FrontierCode 1.1, main set

Score against cost · reported by OpenAI on 22 Sept 2026 · exact values

AI agents write code graded on correctness and mergeability: test quality, scope discipline, code style and codebase standards.

  • GPT-6 Luna
  • GPT-5.6 Luna
  • GPT-6 Sol
  • GPT-5.6 Sol
  • GPT-6 Astra
  • Claude Opus 5
  • Claude Fable 5.1
Show data table
ModelEffortScoreCost per task
GPT-6 Lunalow25.7%$0.021
GPT-6 Lunamedium35.5%$0.053
GPT-6 Lunahigh37.3%$0.067
GPT-6 Lunaxhigh37.1%$0.073
GPT-6 Lunamax42.4%$0.11
GPT-5.6 Lunalow15.4%$0.06
GPT-5.6 Lunamedium25.7%$0.13
GPT-5.6 Lunahigh35.9%$0.23
GPT-5.6 Lunaxhigh38.9%$0.31
GPT-5.6 Lunamax39.8%$0.37
GPT-6 Sollow37.3%$0.45
GPT-6 Solmedium45.9%$0.80
GPT-6 Solhigh47.7%$1.08
GPT-6 Solxhigh48.4%$1.37
GPT-6 Solmax49.3%$2.14
GPT-5.6 Sollow35.4%$1.89
GPT-5.6 Solmedium39.9%$2.69
GPT-5.6 Solhigh45.1%$3.48
GPT-5.6 Solxhigh46.8%$4.15
GPT-5.6 Solmax47.5%$5.19
GPT-6 Astralow45.3%$1.70
GPT-6 Astramedium48.8%$2.43
GPT-6 Astrahigh50.9%$3.01
GPT-6 Astraxhigh50.6%$3.28
GPT-6 Astramax53.3%$4.59
Claude Opus 5low41.9%$2.68
Claude Opus 5medium53.4%$4.31
Claude Opus 5high48%$7.24
Claude Opus 5xhigh43.6%$9.14
Claude Opus 5max48%$11.4
Claude Fable 5.1low49.8%$2.38
Claude Fable 5.1medium50.9%$3.28
Claude Fable 5.1high50.3%$5.27
Claude Fable 5.1xhigh48.7%$9.27
Claude Fable 5.1max50.3%$12.8
  • Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
Source: Introducing GPT-6 Sol and Luna (OpenAI)
DeepSWE 1.1

Score against cost · reported by OpenAI on 22 Sept 2026 · exact values

AI agents solve original, long-horizon software engineering tasks in real codebases.

  • GPT-6 Luna
  • GPT-5.6 Luna
  • GPT-6 Sol
  • GPT-5.6 Sol
  • GPT-6 Astra
  • Claude Opus 5
  • Claude Fable 5
Show data table
ModelEffortScoreCost per task
GPT-6 Lunalow2.4%$0.0057
GPT-6 Lunamedium44.5%$0.052
GPT-6 Lunahigh59.3%$0.084
GPT-6 Lunaxhigh61.3%$0.11
GPT-6 Lunamax66.6%$0.22
GPT-5.6 Lunalow1.2%$0.011
GPT-5.6 Lunamedium9.3%$0.031
GPT-5.6 Lunahigh42.4%$0.13
GPT-5.6 Lunaxhigh56.2%$0.27
GPT-5.6 Lunamax62.2%$0.53
GPT-6 Sollow37.2%$0.16
GPT-6 Solmedium56.6%$0.38
GPT-6 Solhigh65.3%$0.64
GPT-6 Solxhigh66.6%$1.00
GPT-6 Solmax68.8%$2.74
GPT-5.6 Sollow45.4%$0.82
GPT-5.6 Solmedium61.1%$1.42
GPT-5.6 Solhigh69.4%$2.66
GPT-5.6 Solxhigh70.7%$3.60
GPT-5.6 Solmax72.7%$6.46
GPT-6 Astralow67%$1.60
GPT-6 Astramedium72.8%$3.08
GPT-6 Astrahigh73.2%$3.92
GPT-6 Astraxhigh74.1%$4.43
GPT-6 Astramax73.2%$7.50
Claude Opus 5low58.1%$1.66
Claude Opus 5medium68.9%$3.29
Claude Opus 5high72.8%$6.08
Claude Opus 5xhigh73.2%$9.07
Claude Opus 5max73.7%$11.8
Claude Fable 5low59.6%$3.76
Claude Fable 5medium65.4%$6.09
Claude Fable 5high68.6%$9.18
Claude Fable 5xhigh69.9%$13.4
Claude Fable 5max69.7%$21.6
  • Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
  • OpenAI used Claude Fable 5 scores because Fable 5.1 scores were unavailable.
Source: Introducing GPT-6 Sol and Luna (OpenAI)
OSWorld 2.0, offline set

Partial reward against cost · release v2026.08.08 · reported by OpenAI on 22 Sept 2026 · exact values

AI agents attempt long-horizon computer-use workflows covering everyday and professional tasks.

  • GPT-6 Luna
  • GPT-5.6 Luna
  • GPT-6 Sol
  • GPT-5.6 Sol
  • GPT-6 Astra
  • Claude Opus 5
Show data table
ModelEffortScoreCost per task
GPT-6 Lunalow8.3%$0.03
GPT-6 Lunamedium31.5%$0.062
GPT-6 Lunahigh41.4%$0.12
GPT-6 Lunaxhigh46.7%$0.16
GPT-6 Lunamax52.7%$0.27
GPT-5.6 Lunalow12%$0.019
GPT-5.6 Lunamedium22.7%$0.052
GPT-5.6 Lunahigh35.9%$0.16
GPT-5.6 Lunaxhigh48%$0.36
GPT-5.6 Lunamax52.7%$0.49
GPT-6 Sollow43.9%$0.97
GPT-6 Solmedium54%$1.32
GPT-6 Solhigh58.3%$1.64
GPT-6 Solxhigh60.5%$2.21
GPT-6 Solmax64.4%$3.25
GPT-5.6 Sollow29.8%$0.91
GPT-5.6 Solmedium49.7%$2.73
GPT-5.6 Solhigh56.5%$4.46
GPT-5.6 Solxhigh60.9%$5.93
GPT-5.6 Solmax66.2%$7.71
GPT-6 Astralow62.2%$2.55
GPT-6 Astramedium69.3%$5.10
GPT-6 Astrahigh70%$6.60
GPT-6 Astraxhigh71.3%$7.17
GPT-6 Astramax73.5%$9.07
Claude Opus 5low55.2%$9.88
Claude Opus 5medium60.3%$12.7
Claude Opus 5high65.9%$15.9
Claude Opus 5xhigh70.1%$23.9
Claude Opus 5max70.2%$24.1
  • Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
  • Claude Opus 5 values match the official OSWorld 2.0 leaderboard (osworld-v2.xlang.ai).
Source: Introducing GPT-6 Sol and Luna (OpenAI)

Anthropic: Introducing Claude Opus 5.5Read announcement

Claude Opus 5.5 launch results

Reported by Anthropic on 22 Sept 2026 · exact values

BenchmarkClaude Opus 5.5Claude Fable 5.1Claude Opus 5GPT-6 AstraGPT-5.6 Sol
Terminal-Bench 4.0Agentic coding66.4%55.8%52.3%57.9%37.3%
FrontierCode v1.1 (main)Agentic coding54.4%50.3%48%53.3%47.5%
CursorBench 4.0Agentic coding57.8%51.8%46.6%–41.7%
GDPval-AA v2.1Knowledge work1846 Elo1735 Elo1708 Elo1542 Elo1588 Elo
AutomationBenchBusiness workflows40%31.4%26.9%41.4%28.8%
Humanity's Last ExamMultidisciplinary reasoning · with tools67.7%65.6%63.6%57.2%–
Terminal-Bench-Science 0.1Agentic scientific research58.7%52.6%29%64.6%22.4%
OSWorld 2.1Computer use · partial81.8%80.7%74%––
ChartographyVisual chart recognition · with tools89%88.4%83.4%––
  • Bold marks the best reported score in each row.
  • Claude Opus 5.5 results use adaptive thinking at max effort unless noted.
  • Terminal-Bench 4.0 shows each model's highest score: Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort. GPT-6 Astra and GPT-5.6 Sol figures there are as reported by OpenAI.
  • A dash means the lab did not report a result for that model.
Source: Introducing Claude Opus 5.5 (Anthropic)
Terminal-Bench 4.0

Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values

Measures how well a model completes complex, multi-step professional tasks within a command line interface.

  • Claude Opus 5.5
  • Claude Fable 5.1
  • Claude Opus 5
  • GPT-6 Astra
  • GPT-5.6 Sol
Show data table
ModelEffortScoreCost per task
Claude Opus 5.5low38.5%$1.29
Claude Opus 5.5medium57.6%$2.94
Claude Opus 5.5high64.2%$3.88
Claude Opus 5.5xhigh66.4%$7.35
Claude Opus 5.5max64.8%$11.2
Claude Fable 5.1low40.2%$5.70
Claude Fable 5.1medium43.4%$7.80
Claude Fable 5.1high49.4%$10.5
Claude Fable 5.1xhigh51.3%$15.8
Claude Fable 5.1max55.8%$19.5
Claude Opus 5low28.5%$4.25
Claude Opus 5medium41.2%$7.00
Claude Opus 5high47%$10.6
Claude Opus 5xhigh50.6%$13.5
Claude Opus 5max52.3%$15.8
GPT-6 Astralow49.7%$4.95
GPT-6 Astramedium53.9%$6.15
GPT-6 Astrahigh57.9%$7.21
GPT-6 Astraxhigh57.6%$7.48
GPT-6 Astramax56.7%$10.3
GPT-5.6 Sollow7.9%$1.46
GPT-5.6 Solmedium20.9%$2.69
GPT-5.6 Solhigh26.1%$4.12
GPT-5.6 Solxhigh28.5%$5.39
GPT-5.6 Solmax37.3%$7.89
  • Exact values from the data published with Anthropic's announcement.
  • GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI.
Source: Introducing Claude Opus 5.5 (Anthropic)
FrontierCode v1.1, main set

Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values

Measures whether an agent's code changes would be merged.

  • Claude Opus 5.5
  • Claude Fable 5.1
  • Claude Opus 5
  • GPT-6 Astra
  • GPT-5.6 Sol
Show data table
ModelEffortScoreCost per task
Claude Opus 5.5low47.3%$0.40
Claude Opus 5.5medium54.64%$0.80
Claude Opus 5.5high53.99%$1.09
Claude Opus 5.5xhigh51.42%$2.25
Claude Opus 5.5max54.43%$6.19
Claude Fable 5.1low52.8%$2.47
Claude Fable 5.1medium50.91%$3.28
Claude Fable 5.1high50.34%$5.27
Claude Fable 5.1xhigh48.73%$9.27
Claude Fable 5.1max50.28%$12.8
Claude Opus 5low41.95%$2.64
Claude Opus 5medium53.38%$4.61
Claude Opus 5high47.99%$7.62
Claude Opus 5xhigh43.65%$8.99
Claude Opus 5max48.04%$12.3
GPT-6 Astralow45.27%$1.59
GPT-6 Astramedium48.83%$2.28
GPT-6 Astrahigh50.94%$2.85
GPT-6 Astraxhigh50.62%$3.10
GPT-6 Astramax53.26%$4.36
GPT-5.6 Sollow35.44%$1.75
GPT-5.6 Solmedium39.93%$2.50
GPT-5.6 Solhigh45.06%$3.25
GPT-5.6 Solxhigh46.84%$3.88
GPT-5.6 Solmax47.49%$4.85
  • Exact values from the data published with Anthropic's announcement.
Source: Introducing Claude Opus 5.5 (Anthropic)
CursorBench 4.0

Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values

Evaluates coding agents on ambiguous, multi-file tasks taken from real Cursor sessions.

  • Claude Opus 5.5
  • Claude Fable 5.1
  • Claude Opus 5
  • GPT-5.6 Sol
Show data table
ModelEffortScoreCost per task
Claude Opus 5.5low43.7%$1.18
Claude Opus 5.5medium52.5%$2.90
Claude Opus 5.5high56%$3.97
Claude Opus 5.5xhigh56%$6.99
Claude Opus 5.5max57.8%$13.4
Claude Fable 5.1low45.1%$5.44
Claude Fable 5.1medium46.8%$7.05
Claude Fable 5.1high49.2%$9.08
Claude Fable 5.1xhigh51.6%$13
Claude Fable 5.1max51.8%$17.3
Claude Opus 5low40.7%$4.87
Claude Opus 5medium43.3%$6.94
Claude Opus 5high44.7%$9.00
Claude Opus 5xhigh46.1%$11.4
Claude Opus 5max46.6%$11.9
GPT-5.6 Sollow24.6%$0.87
GPT-5.6 Solmedium31.1%$1.77
GPT-5.6 Solhigh35.7%$2.85
GPT-5.6 Solxhigh37.7%$4.40
GPT-5.6 Solmax41.7%$8.23
  • Exact values from the data published with Anthropic's announcement.
Source: Introducing Claude Opus 5.5 (Anthropic)
GDPval-AA v2.1

Elo against cost · reported by Anthropic on 22 Sept 2026 · exact values

Artificial Analysis's evaluation of agents on real-world professional work across 44 occupations.

  • Claude Opus 5.5
  • Claude Fable 5.1
  • Claude Opus 5
  • GPT-6 Astra
  • GPT-5.6 Sol
Show data table
ModelEffortEloCost per task
Claude Opus 5.5low1224 Elo$0.21
Claude Opus 5.5medium1576 Elo$0.86
Claude Opus 5.5high1692 Elo$1.54
Claude Opus 5.5xhigh1820 Elo$4.21
Claude Opus 5.5max1846 Elo$8.92
Claude Fable 5.1low1450 Elo$1.41
Claude Fable 5.1medium1536 Elo$2.17
Claude Fable 5.1high1617 Elo$3.43
Claude Fable 5.1xhigh1721 Elo$7.09
Claude Fable 5.1max1735 Elo$9.59
Claude Opus 5low1294 Elo$0.58
Claude Opus 5medium1476 Elo$1.36
Claude Opus 5high1581 Elo$3.03
Claude Opus 5xhigh1676 Elo$4.96
Claude Opus 5max1708 Elo$6.76
GPT-6 Astralow1366 Elo$0.85
GPT-6 Astramedium1468 Elo$1.82
GPT-6 Astrahigh1485 Elo$2.43
GPT-6 Astraxhigh1516 Elo$3.04
GPT-6 Astramax1542 Elo$4.53
GPT-5.6 Sollow1289 Elo$0.27
GPT-5.6 Solmedium1403 Elo$0.60
GPT-5.6 Solhigh1480 Elo$1.11
GPT-5.6 Solxhigh1548 Elo$1.70
GPT-5.6 Solmax1588 Elo$2.81
  • Exact values from the data published with Anthropic's announcement.
Source: Introducing Claude Opus 5.5 (Anthropic)
AutomationBench

Pass rate against cost · reported by Anthropic on 22 Sept 2026 · exact values

Built by Zapier, tests whether an agent can carry out real business workflows across many connected apps.

  • Claude Opus 5.5
  • Claude Opus 5
  • GPT-6 Astra
  • GPT-5.6 Sol
Show data table
ModelEffortPass rateCost per task
Claude Opus 5.5low23.3%$0.50
Claude Opus 5.5medium28.6%$0.64
Claude Opus 5.5high32%$0.70
Claude Opus 5.5xhigh34.4%$0.86
Claude Opus 5.5max40%$1.37
Claude Opus 5low20.4%$0.75
Claude Opus 5medium23.9%$0.89
Claude Opus 5high20.6%$1.03
Claude Opus 5xhigh25.3%$1.15
Claude Opus 5max26.9%$1.27
GPT-6 Astralow30.3%$1.08
GPT-6 Astramedium34.1%$1.28
GPT-6 Astrahigh37.1%$1.45
GPT-6 Astraxhigh39%$1.53
GPT-6 Astramax41.4%$1.77
GPT-5.6 Sollow11.7%$0.42
GPT-5.6 Solmedium19.6%$0.57
GPT-5.6 Solhigh24.8%$0.64
GPT-5.6 Solxhigh26.3%$0.74
GPT-5.6 Solmax28.8%$0.91
  • Exact values from the data published with Anthropic's announcement.
  • Runs were performed by Zapier without fallback models, so safeguard interventions counted as failures.
Source: Introducing Claude Opus 5.5 (Anthropic)
WANDR

Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values

Perplexity's benchmark measuring agents on large data collection tasks.

  • Claude Opus 5.5
  • Claude Fable 5.1
  • Claude Opus 5
Show data table
ModelEffortScoreCost per task
Claude Opus 5.5low31.2%$1.20
Claude Opus 5.5medium62.8%$11.2
Claude Opus 5.5high67.3%$16.9
Claude Opus 5.5xhigh71.3%$29.1
Claude Opus 5.5max72.3%$37.9
Claude Fable 5.1low63.3%$23.3
Claude Fable 5.1medium64.5%$27.4
Claude Fable 5.1high66.7%$33.6
Claude Fable 5.1xhigh67.7%$42.8
Claude Fable 5.1max68.7%$49
Claude Opus 5low50.5%$10.7
Claude Opus 5medium58.1%$24
Claude Opus 5high64.6%$43.5
Claude Opus 5xhigh67%$53.8
Claude Opus 5max67.2%$61.6
  • Exact values from the data published with Anthropic's announcement.
  • Claude models ran with offline web search and fetch tools and a 980K-token task budget. This differs from Perplexity's published setup, so scores are not comparable with Perplexity's leaderboard.
Source: Introducing Claude Opus 5.5 (Anthropic)

OpenAI: GPT-6 Astra: The next generation in intelligence for workRead announcement

Terminal-Bench 4.0

Accuracy against API cost · reported by OpenAI on 9 Sept 2026 · exact values

Tests agents on complex terminal-based tasks, including software engineering, system configuration and data analysis.

  • GPT-6 Astra
  • GPT-5.6 Sol
  • Claude Fable 5.1
  • Claude Fable 5
  • Claude Opus 5
Show data table
ModelEffortAccuracyCost per task
GPT-6 Astralow49.7%$4.95
GPT-6 Astramedium53.9%$6.15
GPT-6 Astrahigh57.9%$7.21
GPT-6 Astraxhigh57.6%$7.48
GPT-6 Astramax56.7%$10.3
GPT-5.6 Sollow7.9%$1.46
GPT-5.6 Solmedium20.9%$2.69
GPT-5.6 Solhigh26.1%$4.12
GPT-5.6 Solxhigh28.5%$5.39
GPT-5.6 Solmax37.3%$7.89
Claude Fable 5.1low40.2%$5.70
Claude Fable 5.1medium43.4%$7.80
Claude Fable 5.1high49.4%$10.5
Claude Fable 5.1xhigh51.3%$15.8
Claude Fable 5.1max55.8%$19.5
Claude Fable 5max44.5%$22.2
Claude Opus 5max52.6%$18.4
  • Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
  • Claude Fable 5 and Claude Opus 5 are shown at max effort only, as published.
Source: GPT-6 Astra: The next generation in intelligence for work (OpenAI)

Anthropic: Introducing Claude Sonnet 5.5Read announcement

Claude Sonnet 5.5 launch results

Reported by Anthropic on 28 Sept 2026 · exact values

BenchmarkClaude Sonnet 5.5Claude Sonnet 5Claude Opus 5.5GPT-6 Sol
Terminal-Bench 4.0Agentic coding70.6%10.3%66.4%–
FrontierCode 1.1 (main)Agentic coding · Sonnet 5.5 at max effort, 52.1% at xhigh46.2%42.4%54.4%49.3%
CursorBench 4.0Agentic coding55.5%34.1%57.8%–
GDPval-AA v2.1Knowledge work1844 Elo1449 Elo1846 Elo1487 Elo
AA-Briefcase v1.1Knowledge work1811 Elo1359 Elo1822 Elo1483 Elo
Humanity's Last ExamMultidisciplinary reasoning · with tools64.5%54.9%67.7%–
OSWorld 2.1Computer use · partial80.1%57%81.8%–
ChartographyVisual chart recognition · no tools61.6%15.6%64.4%53.6%
  • Bold marks the best reported score in each row.
  • Claude Opus 5.5's Terminal-Bench 4.0 score is at xhigh effort, its highest.
  • OpenAI recently fixed a bug affecting GPT-6 Sol's image understanding; its GDPval-AA, AA-Briefcase and Chartography scores may predate the fix.
  • A dash means the lab did not report a result for that model.
Source: Introducing Claude Sonnet 5.5 (Anthropic)
Terminal-Bench 4.0

Accuracy against cost · reported by Anthropic on 28 Sept 2026 · exact values

Measures how well a model completes complex, multi-step professional tasks within a command-line interface.

  • Claude Sonnet 5.5
  • Claude Opus 5.5
  • Claude Sonnet 5
  • GPT-5.6 Sol
Show data table
ModelEffortScoreCost per task
Claude Sonnet 5.5low20%$0.76
Claude Sonnet 5.5medium28.8%$0.83
Claude Sonnet 5.5high43%$1.94
Claude Sonnet 5.5xhigh61.5%$5.30
Claude Sonnet 5.5max70.6%$12.5
Claude Opus 5.5low38.5%$1.29
Claude Opus 5.5medium57.6%$2.94
Claude Opus 5.5high64.2%$3.88
Claude Opus 5.5xhigh66.4%$7.35
Claude Opus 5.5max64.8%$11.2
Claude Sonnet 5low3.2%$3.78
Claude Sonnet 5medium4.4%$4.83
Claude Sonnet 5high4.5%$8.20
Claude Sonnet 5xhigh7%$9.95
Claude Sonnet 5max10.3%$11.6
GPT-5.6 Sollow7.9%$1.46
GPT-5.6 Solmedium20.9%$2.69
GPT-5.6 Solhigh26.1%$4.12
GPT-5.6 Solxhigh28.5%$5.39
GPT-5.6 Solmax37.3%$7.89
  • Exact values from the data published with Anthropic's announcement.
  • OpenAI did not report GPT-6 Sol on this benchmark, so Anthropic shows GPT-5.6 Sol.
Source: Introducing Claude Sonnet 5.5 (Anthropic)
FrontierCode v1.1, main set

Accuracy against cost · reported by Anthropic on 28 Sept 2026 · exact values

Measures whether an agent's code changes would be merged.

  • Claude Sonnet 5.5
  • Claude Opus 5.5
  • Claude Sonnet 5
  • GPT-6 Sol
Show data table
ModelEffortScoreCost per task
Claude Sonnet 5.5low29.3%$0.19
Claude Sonnet 5.5medium36.5%$0.24
Claude Sonnet 5.5high49.4%$0.42
Claude Sonnet 5.5xhigh52.1%$1.59
Claude Sonnet 5.5max46.2%$20.8
Claude Opus 5.5low47.3%$0.40
Claude Opus 5.5medium54.6%$0.80
Claude Opus 5.5high54%$1.09
Claude Opus 5.5xhigh51.4%$2.25
Claude Opus 5.5max54.4%$6.19
Claude Sonnet 5low28.7%$2.39
Claude Sonnet 5medium35.2%$3.81
Claude Sonnet 5high39.4%$6.10
Claude Sonnet 5xhigh42.7%$10.1
Claude Sonnet 5max42.4%$17.1
GPT-6 Sollow37.3%$0.43
GPT-6 Solmedium45.9%$0.77
GPT-6 Solhigh47.7%$1.04
GPT-6 Solxhigh48.4%$1.32
GPT-6 Solmax49.3%$2.07
  • Exact values from the data published with Anthropic's announcement.
Source: Introducing Claude Sonnet 5.5 (Anthropic)
CursorBench 4.0

Accuracy against cost · reported by Anthropic on 28 Sept 2026 · exact values

Evaluates coding agents on ambiguous, multi-file tasks taken from real Cursor sessions.

  • Claude Sonnet 5.5
  • Claude Opus 5.5
  • Claude Sonnet 5
  • GPT-5.6 Sol
Show data table
ModelEffortScoreCost per task
Claude Sonnet 5.5low35.8%$0.50
Claude Sonnet 5.5medium39.2%$0.70
Claude Sonnet 5.5high47.8%$1.67
Claude Sonnet 5.5xhigh53.1%$3.88
Claude Sonnet 5.5max55.5%$9.67
Claude Opus 5.5low43.7%$1.17
Claude Opus 5.5medium52.5%$2.91
Claude Opus 5.5high56%$3.97
Claude Opus 5.5xhigh56%$6.98
Claude Opus 5.5max57.8%$13.4
Claude Sonnet 5low24.1%$1.39
Claude Sonnet 5medium28%$2.31
Claude Sonnet 5high30.8%$3.48
Claude Sonnet 5xhigh32%$4.55
Claude Sonnet 5max34.1%$7.17
GPT-5.6 Sollow24.6%$0.87
GPT-5.6 Solmedium31.1%$1.77
GPT-5.6 Solhigh35.7%$2.85
GPT-5.6 Solxhigh37.7%$4.40
GPT-5.6 Solmax41.7%$8.23
  • Exact values from the data published with Anthropic's announcement.
  • CursorBench does not report GPT-6 Sol publicly, so Anthropic shows GPT-5.6 Sol.
Source: Introducing Claude Sonnet 5.5 (Anthropic)
AA-Briefcase v1.1

Elo against cost · reported by Anthropic on 28 Sept 2026 · exact values

Artificial Analysis's benchmark of long-horizon knowledge work.

  • Claude Sonnet 5.5
  • Claude Opus 5.5
  • Claude Sonnet 5
  • GPT-6 Sol
Show data table
ModelEffortEloCost per task
Claude Sonnet 5.5low1264 Elo$0.87
Claude Sonnet 5.5medium1461 Elo$1.64
Claude Sonnet 5.5high1634 Elo$3.95
Claude Sonnet 5.5xhigh1746 Elo$9.63
Claude Sonnet 5.5max1811 Elo$29.2
Claude Opus 5.5low1285 Elo$1.15
Claude Opus 5.5medium1642 Elo$4.40
Claude Opus 5.5high1705 Elo$6.27
Claude Opus 5.5xhigh1780 Elo$12.3
Claude Opus 5.5max1822 Elo$21
Claude Sonnet 5low923 Elo$0.82
Claude Sonnet 5medium1056 Elo$1.73
Claude Sonnet 5high1177 Elo$3.76
Claude Sonnet 5xhigh1274 Elo$7.56
Claude Sonnet 5max1359 Elo$14.4
GPT-6 Sollow905 Elo$0.12
GPT-6 Solmedium1142 Elo$0.34
GPT-6 Solhigh1289 Elo$0.63
GPT-6 Solxhigh1364 Elo$1.19
GPT-6 Solmax1483 Elo$2.67
  • Exact values from the data published with Anthropic's announcement.
  • Artificial Analysis ran Sonnet 5.5 on a pre-release deployment with a since-fixed structured output bug, which may slightly understate its score.
Source: Introducing Claude Sonnet 5.5 (Anthropic)

Anthropic: Introducing Claude Fable 5.1 and Claude Mythos 5.1Read announcement

Claude Fable 5.1 launch results

Reported by Anthropic on 1 Sept 2026 · exact values

BenchmarkClaude Fable 5.1Claude Fable 5Claude Opus 5GPT-5.6 Sol
Terminal-Bench-Science 0.1Agentic scientific research52.6%24.7%29%22.4%
Terminal-Bench 4.0Agentic coding55.8%42%52.3%37.3%
GDPval-AA v2Knowledge work1853 Elo1723 Elo1824 Elo1711 Elo
OSWorld 2.0Computer use · partial77.9%72.9%75.4%–
OSWorld 2.0 (strict)Computer use · strict41.7%36.1%39.6%–
Humanity's Last Exam (no tools)Multidisciplinary reasoning · no tools60.9%57.8%56.6%–
Humanity's Last Exam (with tools)Multidisciplinary reasoning · with tools65%63.8%63.6%–
AutomationBenchBusiness workflows31.4%17.1%26.9%19.6%
CursorBench 3.2.0Agentic coding73.4%70.5%70%67.2%
  • Bold marks the best reported score in each row.
  • Fable 5.1 was evaluated with production safeguards enabled. Where they intervened, Fable 5.1 and Fable 5 scored zero on OSWorld 2.0 and Fable 5 scored zero on AutomationBench.
  • Claude Mythos 5.1, the same underlying model with access limited to vetted users, scores 60.9% on Terminal-Bench 4.0. Mythos has no public API price, so it is not tracked here.
  • A dash means the lab did not report a result for that model.
Source: Introducing Claude Fable 5.1 and Claude Mythos 5.1 (Anthropic)
Terminal-Bench-Science 0.1

Accuracy against cost · reported by Anthropic on 1 Sept 2026 · exact values

Terminal-based agentic scientific research tasks, in Anthropic's agentic scientific research category.

  • Claude Fable 5.1
  • Claude Fable 5
Show data table
ModelEffortScoreCost per task
Claude Fable 5.1low26.3%$11.1
Claude Fable 5.1medium35.7%$14.9
Claude Fable 5.1high40%$20.3
Claude Fable 5.1xhigh49.5%$31.8
Claude Fable 5.1max52.6%$37.9
Claude Fable 5low12.3%$17.1
Claude Fable 5medium21.4%$25
Claude Fable 5high25%$34.3
Claude Fable 5xhigh23.4%$36
Claude Fable 5max24.7%$44.1
  • Exact values from the data published with Anthropic's announcement.
  • Standard error is about 3.5 to 4.5 points per model.
Source: Introducing Claude Fable 5.1 and Claude Mythos 5.1 (Anthropic)
Humanity's Last Exam, with tools

Pass rate against cost · reported by Anthropic on 1 Sept 2026 · exact values

Expert-level questions across many disciplines, answered with access to tools.

  • Claude Fable 5.1
  • Claude Fable 5
Show data table
ModelEffortPass rateCost per task
Claude Fable 5.1low60%$0.52
Claude Fable 5.1medium62.96%$0.67
Claude Fable 5.1high64.76%$1.05
Claude Fable 5.1xhigh65.08%$2.28
Claude Fable 5.1max65%$3.20
Claude Fable 5low59.64%$0.61
Claude Fable 5medium61.4%$1.01
Claude Fable 5high63.16%$1.42
Claude Fable 5xhigh63.52%$1.92
Claude Fable 5max63.8%$3.44
  • Exact values from the data published with Anthropic's announcement.
Source: Introducing Claude Fable 5.1 and Claude Mythos 5.1 (Anthropic)
Humanity's Last Exam, no tools

Pass rate against cost · reported by Anthropic on 1 Sept 2026 · exact values

Expert-level questions across many disciplines, answered without tools.

  • Claude Fable 5.1
  • Claude Fable 5
Show data table
ModelEffortPass rateCost per task
Claude Fable 5.1low53.16%$0.30
Claude Fable 5.1medium55.92%$0.46
Claude Fable 5.1high57.96%$0.75
Claude Fable 5.1xhigh60.36%$1.53
Claude Fable 5.1max60.92%$2.23
Claude Fable 5low50.6%$0.17
Claude Fable 5medium55.88%$0.40
Claude Fable 5high56.88%$0.62
Claude Fable 5xhigh57.44%$0.91
Claude Fable 5max57.76%$1.70
  • Exact values from the data published with Anthropic's announcement.
Source: Introducing Claude Fable 5.1 and Claude Mythos 5.1 (Anthropic)
CursorBench 3.2.0

Accuracy against cost · reported by Anthropic on 1 Sept 2026 · exact values

Evaluates coding agents on tasks taken from real Cursor sessions (earlier release than CursorBench 4.0).

  • Claude Fable 5.1
  • Claude Fable 5
Show data table
ModelEffortScoreCost per task
Claude Fable 5.1low66.2%$2.90
Claude Fable 5.1medium68%$3.53
Claude Fable 5.1high69.4%$4.80
Claude Fable 5.1xhigh72.8%$6.96
Claude Fable 5.1max73.4%$9.64
Claude Fable 5low62.1%$4.46
Claude Fable 5medium65.2%$6.80
Claude Fable 5high66.5%$8.77
Claude Fable 5xhigh68.4%$11.7
Claude Fable 5max70.5%$17.3
  • Exact values from the data published with Anthropic's announcement.
Source: Introducing Claude Fable 5.1 and Claude Mythos 5.1 (Anthropic)