fiveagents logofiveagents.io
Released 01–02 Sep 2026 · Anthropic

Claude Fable 5.1
benchmark by benchmark

Every number below is quoted directly from Anthropic's own release page, browser-checked 02sep26 — not a third-party recap. AutomationBench is the one that matters if you run a business, not a lab.

Every Fable 5.1 benchmark against Fable 5 and Opus 5, quoted straight from Anthropic’s own release page — pricing and safeguard changes included. (8:54).youtube.com/watch?v=efhxZPpBXeI

The number that matters most
17.1% → 31.4%

AutomationBench (Zapier) — a real, multi-step business workflow, scored on whether the model finishes it without a human stepping in. Fable 5.1 nearly doubled Fable 5's clear rate, and beat Opus 5's 26.9%.

All nine benchmarks

BenchmarkWhat it measuresFable 5Opus 5Fable 5.1
AutomationBenchBusiness workflows, end to end17.1%26.9%31.4%
Terminal-Bench-Science 0.1Agentic scientific research24.7%29.0%52.6%
Terminal-Bench 4.0Agentic coding42.0%52.3%55.8%
CursorBench 3.2.0Agentic coding70.5%70.0%73.4%
GDPval-AA v2Knowledge work / economic value172318241853
OSWorld 2.0 (loose)Computer use — task basically done72.9%75.4%77.9%
OSWorld 2.0 (strict)Computer use — done cleanly36.1%39.6%41.7%
Humanity's Last Exam (no tools)Multidisciplinary reasoning57.8%56.6%60.9%
Humanity's Last Exam (with tools)Multidisciplinary reasoning63.8%63.6%65.0%

GPT-5.6 Sol scores, where Anthropic published them: Terminal-Bench-Science 22.4% · Terminal-Bench 4.0 37.3% · CursorBench 67.2% · GDPval-AA 1711 · AutomationBench 19.6%.

x.com/claudeai
Anthropic's published benchmark table for Claude Fable 5.1, comparing Fable 5.1, Fable 5, Opus 5 and GPT-5.6 Sol across agentic scientific research, agentic coding, knowledge work, computer use, multidisciplinary reasoning and business workflows.

Anthropic's own table, posted to @claudeai. The table above restates it with GPT-5.6 Sol split out.

What it costs

Token pricing — unchanged

$10 / million input tokens
$50 / million output tokens
Same as Fable 5.

Cache reads −75%

Down to $0.25 / million tokens. Anthropic states this cuts typical workload cost ~25%, and up to ~45% for heavily agentic use.

x.com/claudeai
Indexed cost of Fable usage: a typical workload falls from 100 to 75 (about 25% less) and a highly agentic workload from 100 to 55 (about 45% less), with cache reads making up most of the saving.

Where the ~25% and ~45% come from — cache reads are most of the bar on agentic work.

What a point of accuracy costs

The table above is every model at one setting. These are Fable 5.1 against Fable 5 at each effort level, plotted against what the task actually cost to run. The point is the shape: 5.1's curve sits above and to the left, so the same score comes cheaper, and the extra effort levels keep paying rather than flattening out.

Terminal-Bench-Science 0.1 accuracy versus cost. Fable 5.1 rises from about 26% at low effort to 52.6% at max, while Fable 5 stays between 12% and 25% at a higher cost per task.

Terminal-Bench-Science 0.1 — agentic scientific research

Terminal-Bench 4.0 accuracy versus cost, showing Mythos 5.1, Fable 5.1 and Mythos 5 curves at each effort level.

Terminal-Bench 4.0 — agentic terminal coding. Anthropic notes Fable 5.1 and Mythos 5.1 are the same underlying model; the gap between them is where the older cyber safeguards intervened.

Humanity's Last Exam pass rate versus cost, with and without tools, for Fable 5.1 and Fable 5.

Humanity's Last Exam — multidisciplinary reasoning, with and without tools

CursorBench 3.2.0 score versus cost per task, Fable 5.1 above Fable 5 at every effort level.

CursorBench 3.2.0 — agentic coding

All four charts posted to x.com/claudeai, 02sep26.

Safeguards

−60% unnecessary interventions−85% false-positive refusals

~60% fewer safety interventions per Claude Code session (cybersecurity). ~85% fewer false-positive refusals on benign biology questions.

Source: anthropic.com/claude-fable-and-mythos-5-1 · browser-verified 02sep26.

Petric Manurung · youtube.com/@petricbuilds