The real cost of
getting an answer.
Seven frontier models. Five reasoning efforts. A practical comparison of token usage, benchmark performance and API spending.
More reasoning is not always better value.
Compare performance against dollars or generated tokens. Higher and farther left is better on these two dimensions. Tap a point to inspect a configuration.
Performance vs resources
Logarithmic horizontal axis. Solid paths connect published efforts; the dashed Fable 5 path uses modelled lower efforts. Points are rounded composite scores, not success probabilities. Sources and limitations ↗
| Model / effort | Generated tokens | Reasoning included | Index points | Output charge only | Full API $/task | 1,000 attempts | Evidence |
|---|
Swipe the table horizontally. k = 1,000 tokens. Amber rows are estimates, not measurements.
Output charges are calculated using rounded output counts and base API prices. Full task charges are published separately and include input and caching. Rounding matters most for Luna. Fable 5.1 uses default fallback; Fable 5 Max uses Opus 4.8 fallback. Table filters are independent of chart legend toggles.
Where does extra effort pay off?
Effort labels are vendor-specific. They are not identical compute budgets across model families.
Follow the effort curve
Changes refer to rounded mixed-benchmark averages. An unchanged integer score does not establish equal underlying capability or statistical equivalence. Fable 5 lower-effort steps are estimated.
Success depends on the test.
Specific evaluation percentages are more informative than treating an aggregate index as a universal success rate. All results in this section use Max effort unless labelled otherwise.
Benchmark performance
| Model · Max | HLE | Terminal pass@1 | SciCode | Automation | GDP.pdf all-pass | AA-LCR |
|---|
Bold cells mark the highest displayed value in each column, including ties. Each evaluation has its own scoring rules. Sources: Astra / Fable 5.1 · Opus / Sol · Luna / Terra · Fable 5.
Translate task costs into a budget.
A planning calculator for attempts distributed like the benchmark mix. It is not a quote for your application or a cost-per-success calculator.
Attempt-based budget
Output charge = generated tokens × output price / 1,000,000
Reasoning is a subset of output. Never charge for it twice. External tool fees, taxes, negotiated discounts and different workloads can change the bill.
Token price is only half the story.
Standard first-party API dollars per million tokens, using base context tiers. Subscription allowances and vendors’ internal inference costs are different quantities.
| Model | Fresh input | Cached input read | Cache write | Output incl. reasoning |
|---|
Claude cache-write prices shown are the five-minute tier. Higher context tiers, Batch, Flex, Fast mode, regional uplifts and tool fees are excluded from these base-price examples. Sol pricing is promotional; the official page says it is available at least through November 21, 2026. OpenAI pricing ↗ · Anthropic pricing ↗
Know what is measured. Know what is guessed.
The published benchmark numbers are the backbone. Lower-effort Fable 5 and raw input-traffic counts are explicitly modelled rather than observed.
Generated tokens
All output across a task’s turns: reasoning plus answers and tool-call output. The reasoning column is already included in the generated total.
Processed tokens
Generated output plus cumulative input traffic, including context reread on later turns. This is not the same as unique context length.
Full task charge
Benchmark-weighted token spending covering input, cache reads and writes, reasoning and answers. An average attempt, not a guaranteed completion.
Fable 5 lower-effort estimates: assumptions and margins
Only Max has a comparable published token measurement in this analysis: 67k generated tokens, 50 index points and $8.75 per attempt, with Opus 4.8 fallback. The four lower-effort rows are analyst estimates.
For each lower effort, normalize Fable 5.1 and Opus 5 against their own Max values. Average those two ratios and multiply by Fable 5’s Max value. For the index, average the two absolute score drops from Max, then subtract that drop from 50.
Fable 5 cost(e) = $8.75 × average[Fable 5.1 cost(e) / $7.63, Opus 5 cost(e) / $5.86]
Subjective planning margins: ±50% around estimated average tokens and spending; approximately ±4 index points. These are not statistical confidence intervals or limits on individual tasks. Reasoning-token splits for these four rows are left unknown.
Anchors: Fable 5 Max · Fable 5.1 effort profile · Opus 5 effort profile.
Total processed tokens: an editable inverse-pricing scenario
The scenario holds the published Max-effort bill fixed, subtracts the calculated output charge, and assumes every remaining dollar pays for cache reads or base-tier cache writes. No fallback-price correction, long-context premium or external fee is applied.
| Model · Max | Published output | Full bill | Estimated total input + output |
|---|
At the default 90% assumption, Astra Max implies about 0.92 million processed tokens; Fable 5.1 Max implies about 2.61 million. These are repeated traffic counts, not unique context sizes.
Moving the slider toward more caching implies that the same bill buys more discounted tokens. It does not mean caching increases cost. Reading a 50k-token context 20 times already creates one million cumulative input tokens.
Benchmark methodology and limits on interpretation
A consistent comparison, not a new benchmark run
This site republishes the preceding research comparison, with source checks on September 14, 2026. It does not claim independently executed benchmark experiments. Artificial Analysis supplies the published configurations and weighted averages; the site’s calculations and estimates are identified separately.
What is inside the index?
AA Intelligence Index v4.3 combines ten evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. Index points are not probabilities.
Why the averages cannot promise an outcome
Individual tasks can consume very different token counts. Means are not caps. Prompting, tools, context, model versions, fallback routing and scoring rules affect the results. Same-number effort settings are not equal-compute interventions across providers.
Rounding and fallback
The displayed token counts, costs and scores are rounded. Output-only charges and residual input/cache charges therefore are approximate. Fable 5.1 uses default fallback, and Fable 5 Max uses Opus 4.8 fallback: neither should be described as an isolated Fable-only measurement.
Sources, not mystery numbers.
Static snapshot: September 14, 2026. Linked source pages may subsequently change. Row-level source URLs are included in the CSV export.
Published output, reasoning, composite scores and full cost per attempt.
Artificial Analysis release comparison ↗Five reasoning-effort configurations for each model.
Artificial Analysis release comparison ↗Separate comparison pages for each matched effort.
Max-effort results with Opus 4.8 fallback; lower efforts on this site are estimated.
Fable 5 Max comparison ↗Max-effort evaluation percentages, rather than inferred universal success rates.
Astra / Fable 5.1 ↗ · Opus / Sol ↗Index v4.3 construction, evaluation setup and benchmark weighting.
Artificial Analysis methodology ↗Standard base-tier input, cache-read, cache-write and output rates.
Official OpenAI pricing ↗Official unit rates and prompt-caching tiers.
Official Claude pricing ↗