AIFrontier Economics
Research snapshot · 14 September 2026

The real cost of
getting an answer.

Seven frontier models. Five reasoning efforts. A practical comparison of token usage, benchmark performance and API spending.

31 published configurations4 explicit estimatesAA Intelligence Index v4.3
Explore the comparison
17kAstra Xhigh generated tokens
$2.31Astra Xhigh full API cost/task
53Composite points, not 53% success
$2,3101,000 similar Astra Xhigh attempts
A task attempt is not a guaranteed success. Costs and token counts are weighted averages across a benchmark mix, not fixed budgets or the length of a typical chat reply. Generated tokens already include reasoning. Read the definitions →
01 / The trade-off

More reasoning is not always better value.

Compare performance against dollars or generated tokens. Higher and farther left is better on these two dimensions. Tap a point to inspect a configuration.

Performance vs resources

Logarithmic horizontal axis. Solid paths connect published efforts; the dashed Fable 5 path uses modelled lower efforts. Points are rounded composite scores, not success probabilities. Sources and limitations ↗

Every model. Every effort.

31 published + 4 modelled configurations
Model / effortGenerated
tokens
Reasoning
included
Index
points
Output
charge only
Full API
$/task
1,000
attempts
Evidence

Swipe the table horizontally. k = 1,000 tokens. Amber rows are estimates, not measurements.

Output charges are calculated using rounded output counts and base API prices. Full task charges are published separately and include input and caching. Rounding matters most for Luna. Fable 5.1 uses default fallback; Fable 5 Max uses Opus 4.8 fallback. Table filters are independent of chart legend toggles.

02 / Diminishing returns

Where does extra effort pay off?

Effort labels are vendor-specific. They are not identical compute budgets across model families.

Follow the effort curve

Changes refer to rounded mixed-benchmark averages. An unchanged integer score does not establish equal underlying capability or statistical equivalence. Fable 5 lower-effort steps are estimated.

03 / Actual benchmark results

Success depends on the test.

Specific evaluation percentages are more informative than treating an aggregate index as a universal success rate. All results in this section use Max effort unless labelled otherwise.

Benchmark performance

Model · MaxHLETerminal
pass@1
SciCodeAutomationGDP.pdf
all-pass
AA-LCR

Bold cells mark the highest displayed value in each column, including ties. Each evaluation has its own scoring rules. Sources: Astra / Fable 5.1 · Opus / Sol · Luna / Terra · Fable 5.

Max is not a guaranteed improvement. On Terminal-Bench, Astra reports 60% at Xhigh versus 59% at Max; Fable 5.1 reports 55% versus 52%. These small reversals can reflect evaluation variation and should not be interpreted as proof that lower effort is inherently smarter. Xhigh source ↗
04 / Put a number on it

Translate task costs into a budget.

A planning calculator for attempts distributed like the benchmark mix. It is not a quote for your application or a cost-per-success calculator.

Attempt-based budget

Total budget = attempts × full API cost per attempt
Output charge = generated tokens × output price / 1,000,000

Reasoning is a subset of output. Never charge for it twice. External tool fees, taxes, negotiated discounts and different workloads can change the bill.

Estimated total API spending
$2,310
Do not divide mixed-benchmark cost by an unrelated success percentage. Valid cost-per-success analysis needs spending and successes from the same workload. Retry projections also need assumptions about correlated failures.
05 / Unit economics

Token price is only half the story.

Standard first-party API dollars per million tokens, using base context tiers. Subscription allowances and vendors’ internal inference costs are different quantities.

ModelFresh
input
Cached
input read
Cache
write
Output incl.
reasoning

Claude cache-write prices shown are the five-minute tier. Higher context tiers, Batch, Flex, Fast mode, regional uplifts and tool fees are excluded from these base-price examples. Sol pricing is promotional; the official page says it is available at least through November 21, 2026. OpenAI pricing ↗ · Anthropic pricing ↗

API token bill = (fresh input × input rate + cache reads × read rate + cache writes × write rate + output × output rate) / 1,000,000
Cheaper tokens do not always produce cheaper tasks. Fable 5.1 High and Opus 5 Max both display 51 points, but their full task costs are $3.91 and $5.86. Fable’s higher output-token price is outweighed by its shorter run in this comparison. Sources →
06 / Evidence and uncertainty

Know what is measured. Know what is guessed.

The published benchmark numbers are the backbone. Lower-effort Fable 5 and raw input-traffic counts are explicitly modelled rather than observed.

Generated tokens

All output across a task’s turns: reasoning plus answers and tool-call output. The reasoning column is already included in the generated total.

Processed tokens

Generated output plus cumulative input traffic, including context reread on later turns. This is not the same as unique context length.

Full task charge

Benchmark-weighted token spending covering input, cache reads and writes, reasoning and answers. An average attempt, not a guaranteed completion.

Fable 5 lower-effort estimates: assumptions and margins

Only Max has a comparable published token measurement in this analysis: 67k generated tokens, 50 index points and $8.75 per attempt, with Opus 4.8 fallback. The four lower-effort rows are analyst estimates.

For each lower effort, normalize Fable 5.1 and Opus 5 against their own Max values. Average those two ratios and multiply by Fable 5’s Max value. For the index, average the two absolute score drops from Max, then subtract that drop from 50.

Fable 5 tokens(e) = 67,000 × average[Fable 5.1 tokens(e) / 78,000, Opus 5 tokens(e) / 73,000]
Fable 5 cost(e) = $8.75 × average[Fable 5.1 cost(e) / $7.63, Opus 5 cost(e) / $5.86]

Subjective planning margins: ±50% around estimated average tokens and spending; approximately ±4 index points. These are not statistical confidence intervals or limits on individual tasks. Reasoning-token splits for these four rows are left unknown.

Anchors: Fable 5 Max · Fable 5.1 effort profile · Opus 5 effort profile.

Total processed tokens: an editable inverse-pricing scenario
Not measured input counts. This back-calculation is considerably less reliable than the published dollar costs. Its purpose is to show the scale implied by explicit cache assumptions.

The scenario holds the published Max-effort bill fixed, subtracts the calculated output charge, and assumes every remaining dollar pays for cache reads or base-tier cache writes. No fallback-price correction, long-context premium or external fee is applied.

Model · MaxPublished
output
Full billEstimated total
input + output
Total tokens = output tokens + (full bill − output charge) × 1,000,000 / [(1 − cache-read fraction) × write rate + cache-read fraction × read rate]

At the default 90% assumption, Astra Max implies about 0.92 million processed tokens; Fable 5.1 Max implies about 2.61 million. These are repeated traffic counts, not unique context sizes.

Moving the slider toward more caching implies that the same bill buys more discounted tokens. It does not mean caching increases cost. Reading a 50k-token context 20 times already creates one million cumulative input tokens.

Benchmark methodology and limits on interpretation

A consistent comparison, not a new benchmark run

This site republishes the preceding research comparison, with source checks on September 14, 2026. It does not claim independently executed benchmark experiments. Artificial Analysis supplies the published configurations and weighted averages; the site’s calculations and estimates are identified separately.

What is inside the index?

AA Intelligence Index v4.3 combines ten evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. Index points are not probabilities.

Why the averages cannot promise an outcome

Individual tasks can consume very different token counts. Means are not caps. Prompting, tools, context, model versions, fallback routing and scoring rules affect the results. Same-number effort settings are not equal-compute interventions across providers.

Rounding and fallback

The displayed token counts, costs and scores are rounded. Output-only charges and residual input/cache charges therefore are approximate. Fable 5.1 uses default fallback, and Fable 5 Max uses Opus 4.8 fallback: neither should be described as an isolated Fable-only measurement.

Read Artificial Analysis’s evaluation methodology ↗

07 / Follow the evidence

Sources, not mystery numbers.

Static snapshot: September 14, 2026. Linked source pages may subsequently change. Row-level source URLs are included in the CSV export.

01 · Astra & Fable 5.1 effort data

Published output, reasoning, composite scores and full cost per attempt.

Artificial Analysis release comparison ↗
02 · Opus 5 & Sol effort data

Five reasoning-effort configurations for each model.

Artificial Analysis release comparison ↗
03 · Luna & Terra effort data

Separate comparison pages for each matched effort.

04 · Fable 5 published anchor

Max-effort results with Opus 4.8 fallback; lower efforts on this site are estimated.

Fable 5 Max comparison ↗
05 · Benchmark-specific results

Max-effort evaluation percentages, rather than inferred universal success rates.

Astra / Fable 5.1 ↗ · Opus / Sol ↗
06 · Evaluation definitions

Index v4.3 construction, evaluation setup and benchmark weighting.

Artificial Analysis methodology ↗
07 · OpenAI API pricing

Standard base-tier input, cache-read, cache-write and output rates.

Official OpenAI pricing ↗
08 · Anthropic API pricing

Official unit rates and prompt-caching tiers.

Official Claude pricing ↗