30% Cheaper Per Task Costs 7x the Tokens: Inside Sonnet 5.5's Efficiency Trap
Key takeaways
- Sonnet 5.5's "up to 30% less per task" headline holds at low and medium effort but inverts at max: Anthropic charged the same $2/$10 rate card while the model spent ~193,000 output tokens per index task — the heaviest token use Artificial Analysis has measured — making its measured cost $7.60 per task, about 50% higher than Sonnet 5's and ~27% above Opus 5.5's $5.98 (Anthropic, Sep 28; Artificial Analysis, Sep 28).
- The same release window produced two opposite strategies: OpenAI cut GPT-6 Sol/Luna prices 50% to $2/$10 and $0.10/$0.50 per million tokens on September 22 — GPT-6 Sol max at $1.06 per Intelligence Index task, roughly 50% below its predecessor — while Anthropic kept Sonnet's rate card flat and let the token bill rise (OpenAI; Artificial Analysis).
- Opus 5.5 (September 22) took a third path: 20% off the sticker price and 60% off cache reads ($0.50 → $0.20 per million), targeting the line item that dominates agentic workloads — which is why its measured $5.98 cost per task undercuts its own junior sibling at max effort (Anthropic; Finout).
- At high-to-max effort, Sonnet 5.5 sits off the cost-efficiency frontier: GPT-6 Sol delivers near-equal intelligence at "effectively the same cost per task," and GPT-6 Luna max does equivalent-intelligence work for $0.07 per task — roughly one-hundredth of Sonnet 5.5 max's bill (Artificial Analysis, Sep 28).
- Roughly 180 of 672 tracked models had a September price change — the rate card is now a monthly hazard, and an agent whose contract names only per-token pricing has no defense against a swap it never approved (pricepertoken, Sep 28; the DeepSeek auto-routing reversal is the worked example).
Anthropic shipped Claude Sonnet 5.5 on September 28 with the cleanest cost claim in the market: "It's a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work." Every word is defensible — and the mechanism is not what the headline implies. The rate card did not move: $2 per million input tokens, $10 output, $0.20 cache reads, identical to Sonnet 5. What changed is the token bill, and it moved in the opposite direction from the claim at the effort level agentic work actually uses. Artificial Analysis ran the model the day it shipped: at max effort, Sonnet 5.5 spent ~193,000 output tokens per Intelligence Index task — the highest their testing has ever recorded, about 60% above Opus 5.5 max and roughly seven times GPT-6 Astra — producing a measured cost of $7.60 per task (Artificial Analysis, Sep 28; first recorded by the team at kingy.ai: "Sonnet 5.5 is cheap per token but can be expensive per task").
By itself that is one model's quirk. Read against the same two-week window it is the industry's fork in the road. On September 22, OpenAI cut GPT-6 Sol and Luna prices by 50% (OpenAI) and Anthropic cut Opus 5.5's cache reads by 60% (Anthropic). Six days later, Anthropic's "30% cheaper" entry prices the same token volume that competitors sell at half the rate — and buys #2-on-the-index capability by spending tokens instead of buying them down. Three release strategies now compete on the same metric, cost-per-task, and they diverge most exactly where inference procurement guidance lives. This article maps the three strategies against Artificial Analysis's measured numbers, works the token-efficiency arithmetic a per-token rate card hides, and closes with the four contract clauses that keep the next model swap from silently repricing your agent's bill.
The claim, the mechanism, and the measurement
Anthropic's claim is about efficiency, and the launch page is explicit about the mechanism: the model "typically needs far fewer tokens to do the same work. In our testing, it costs up to 30% less per task than its predecessor." That is true on Anthropic's own test distribution at Anthropic's chosen settings — the company reports Sonnet 5.5 beating Sonnet 5's best benchmark scores "for about a tenth of the cost per task" at low and medium effort (as summarized by kingy.ai). If your workload runs lightweight and Anthropic-shaped — short drafts, classification, bounded lookups — the claim will probably hold on your bill.
Agentic workloads are not that distribution. The Artificial Analysis Intelligence Index runs multi-step tasks under five effort settings, and the model's adaptive reasoning escalates effort on hard tasks. At the top of that range the token bill explodes: 193k output tokens per task at max, versus ~120k for Sonnet 5 max and Opus 5.5 max (+60%) and 27k for GPT-6 Astra max (7x). The token volume is the point — Sonnet 5.5 (max) scores 56 on the index, 2 points behind Opus 5.5's 58, with Terminal-Bench 4.0 at 64%, slightly above Opus 5.5 and GPT-6 Astra. It reaches that capability by thinking longer per step, and output tokens bill at $10 per million regardless of how defensible each one was. A fallback footnote completes the mechanics: Anthropic's default adaptive routing "falls back in ~0.1% of tasks... falling back to Sonnet 5 in all cases" — another rate-card-neutral substitution an operator only sees on the invoice (Artificial Analysis).
Cost per task is the only metric that survives contact with this structure, because it composes the rate card with the token bill. The same-day measurements, all cost-per-task on the Intelligence Index at max effort:
| Model (max effort) | Sticker ($/M in/out) | Output tokens/task | Measured cost/task | Release strategy |
|---|---|---|---|---|
| Claude Sonnet 5.5 | $2 / $10 | ~193k | $7.60 | Capability-led: flat rate card, token bill up |
| Claude Opus 5.5 | $4 / $20 (cache $0.20) | ~120k | $5.98 | Cache-led: sticker −20%, agentic line item −60% |
| Claude Sonnet 5 | $2 / $10 | ~120k | ~$5.00 | Baseline |
| GPT-6 Sol | $2 / $10 | ~31k | $1.06 | Price-led: sticker −50% |
| GPT-6 Luna | $0.10 / $0.50 | ~51k | $0.07 | Price-led: floor rate |
(AA Intelligence Index v4.x scores: Sonnet 5.5 56 at #2; Opus 5.5 58 at #1; GPT-6 Sol ~56-class equivalent on coding-agent work. Scores carry a lineage caveat — AA has rebased its index twice this year, so cite scores only against same-sourced publications.)
The trap is the first and third rows: identical sticker prices, 7.6x apart on measured cost. A procurement decision made on the rate card — "Sonnet 5.5 matches Sonnet 5's pricing, so it is the economical upgrade" — produces the exact inverse of the bill. And the measured ranking inverts the intuition a second way: the more expensive sticker on the table (Opus 5.5 at $4/$20) is the cheaper task, because its cache-read cut attacks the line item agentic workloads actually dominate — re-reading long context thousands of times per task (Anthropic: "Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens").
Three release strategies, one battleground
The September window formalized what the price war had been implying since June: per-token price is no longer where vendors compete, because per-token price has become nearly free. The battleground moved to cost-per-task, and each lab chose a different weapon.
Capability-led (Anthropic, Sonnet line). Ship a near-frontier model on the old rate card and win the benchmark table. Sonnet 5.5 at #2 with a 60%-heavier token bill is the purest form: the intelligence gain is real (+18 index points over Sonnet 5 at max), and the cost of it lands on the invoice's token line, where "30% cheaper" and "$7.60" both stay technically true. The strategy bets that buyers read headlines and benchmarks and treat token volume as a free good — and per Gartner's September 16 forecast that worldwide AI spending grows 49.5% in 2026 to $2.7T, with the agents-and-assistants segment on a $16.5B → $29.2B → $65.5B path, the bet has a real audience.
Price-led (OpenAI, Sol/Luna line). Halve the sticker, publish the per-task math yourself. OpenAI's launch table is procurement-ready in a way no model launch has been: GPT-6 Sol (xhigh) scores 33.2% at $0.27 per task versus Claude Opus 5 (max) at 26.9% for 11.1x the cost. On the Artificial Analysis side, GPT-6 Sol max halved its own predecessor's $1.99 cost to $1.06 while gaining 2 points on the Coding Agent Index (Artificial Analysis). The strategy bets that measured, published cost-per-task is the future's buying metric — and that a vendor confident in its efficiency wins procurement by making cost comparisons effortless.
Cache-led (Anthropic, Opus line). Keep the frontier sticker, gut the agentic cost driver. Opus 5.5's cache-read cut from $0.50 to $0.20 targets exactly the workload shape where agentic tokens live — repeated reads of large, stable context. Anthropic estimates ~40% lower total cost on typical work (Fello AI summarizing the launch claims), and the measured $5.98/task versus $7.60 confirms the direction independently.
The diagram's point is the arithmetic box, not the bars: a rate card is a price per million; a cost-per-task contract is a price per outcome. On a 200-step workload run 1,200 times a week, Sonnet 5.5 max spends 193k output tokens per step and bills about $2,316 weekly where GPT-6 Sol bills $121 for near-equal measured intelligence (both score in the mid-50s on the index; Sol's coding-agent economics at $2.99/task are on the Pareto frontier per AA's coding index analysis). The weekly delta between the "30% cheaper" model and its own predecessor is roughly $1,100 — the "up to 30% less" claim and a 2x cost increase are simultaneously true, distinguished only by effort distribution. That is why the contract matters more than the launch post.
The procurement problem September just made harder
Two structural facts turn this from commentary into contract language. First, the rate-card hazard is now monthly: pricepertoken reads roughly 180 of 672 tracked models with a September price change (the counter's drift across the month: 178/672 → 181/671), and LLM Gateway's timeline lists 23 new models from 15 providers this month — the bill your agent ran in August is not the bill it runs in October. Second, the substitution layer is already shipping as product: LLM Gateway's Smart Route released September 25 makes vendor-side routing a purchasable feature, and Anthropic's adaptive fallback ("falls back in ~0.1% of tasks, primarily in Terminal-Bench, falling back to Sonnet 5 in all cases" per AA's testing) is the same mechanism inside the model call itself. Neither is misbehavior — the DeepSeek episode showed vendors can even announce a fleet-wide model substitution and reverse it 45 hours later under user pressure. The routing decision is real and it belongs to whoever holds the runtime.
Against that backdrop, buying on the rate card has three specific failure modes:
- The benchmark-headline failure. Sonnet 5.5 tops the capability story (index #2, Terminal-Bench 4.0 at 64%) and the cost story (30% less per task) in the same launch — both true, composed against different baselines. A procurement pass that reads the launch post and the benchmark table but not the token column approves the most expensive cost-per-task configuration on the table.
- The effort-blind forecast. Most published cost math assumes published settings. AA's sweep shows Sonnet 5.5's bill swings by effort (max $7.60 vs Sonnet 5 ~$5.00; low/medium settings where the "30% cheaper" claim lives); a forecast built on the API default inherits whichever setting the vendor's SDK escalation logic picks per task.
- The cache-blind unit price. Two Anthropic models 30 days apart chose opposite axes — Sonnet 5.5 left tokens expensive, Opus 5.5 made cached reading cheap. An agent workload's actual cache-hit ratio (long, stable system prompts; repeated tool schemas; big retrieved context) is now the single biggest determinant of which of the two is cheaper — a number most RFQ processes never ask for.
None of these are vendor deception. All three are the normal output of a market whose unit economics shifted underneath its own pricing language — the same conclusion AA reached in one line: "Sonnet 5.5 is cheap per token but can be expensive per task."
The contract: four clauses that survive the next release
The fix is not picking the winning lab; on cost-per-task, the winner changes weekly (pricepertoken's counter moves monthly). It is instrumenting the decision so the next swap is a measured configuration change instead of an invoice surprise. The four clauses map to the same runtime this site builds into every agent:
1. Pin model IDs per step and log the served model, not the requested one. The routing layer is operator infrastructure. model fields live in your agent configuration — in our reference runtime, agents reference registered model resources by provider + name, and changing the model is a data operation, not a redeploy. The audit trail records the actual served model per tool call (the DeepSeek article covers the full incident pattern). A vendor that swaps your model silently now produces a diff you can see.
2. Contract on cost-per-task with a weekly ceiling, not per-token with a monthly invoice. Per-task pricing is becoming a published first-party metric (OpenAI publishes it in launch tables; AA measures everyone against the same tasks). A scoped contract names the workload (steps, context shape, expected cache-hit ratio), the model and effort tier per step class, and a measured ceiling per task — with the vendor's published per-task price as the benchmark row. The six-cost-vectors framework provides the audit checklist; the fifth vector (vendor-substitution risk) is the clause this section operationalizes.
3. Budget tokens with the effort dial as a procurement surface, not a developer setting. Sonnet 5.5's 30x bill swing across effort levels is the concrete case. The routing layer should treat reasoning effort as a typed configuration field per step class — escalation thresholds explicit, ceilings hard-stopped at the runtime (the long-running-agent cost-ceiling pattern), and the escalation policy visible in the audit trail. A "max effort" default across 1,200 daily steps is a procurement decision that was never made.
4. Test the claimed efficiency on your workload before you commit a fleet to it. Anthropic's "up to 30% less" is honest, scoped, and workload-conditional — and AA's counter-measurement of $7.60 is also honest, scoped, and workload-conditional. Neither tells you your bill. The 30-minute pilot: take one week of real agent runs, replay them against the candidate model at each effort tier, measure tokens and cache-hit ratio per step, and put the three numbers (rate card, token bill, measured per-task) in the procurement one-pager. The TypeSafe Jev three-layer stack makes this continuous — a calibrated decision-layer classifier watches the bill per class of step, because the watcher is nearly free per call.
One honesty marker applies to everything above, including the arithmetic box: all measured per-task numbers come from two independent same-day test sources (AA, kingy.ai) plus vendor-published tables, on benchmark tasks, not your workload. AA's Intelligence Index has also been rebased twice this year; the scores and costs cited here are same-publication-date comparisons and can shift at the next rebase. Treat all of it as the starting baseline a pilot replaces — the direction (rate card ≠ bill; tokens are the variable) is stable, the exact constants are not.
The outcome, on a real profile
Run the numbers on the mid-market build profile this site writes around: a distributor's quoting agent, 200 RFQs per week, roughly 1,200 model steps per week across catalog lookup, tier pricing, knowledge-graph substitutes, and draft generation. At Sonnet 5.5 max, that profile bills around $2,300/week. Routed deliberately — frontier effort only on negotiation-relevant steps, a mid-tier model on lookups, a near-floor model on classification, cache-heavy context shared across steps — the same workload bills under $300/week with the quality concentrated where the business outcome actually lives: the quote the customer sees. The 87% difference is not model selection skill; it is the runtime enforcing a cost decision the launch post never made for you. At these volumes the integration layer — MCP modules, the routing policy, the audit trail — remains the 90% of the build that a per-token rate card cannot price; the model bill is the noise, and the pattern above is how you keep it noise instead of the signal.
Related reading
- Inference Economics: Why Always-On Production Agents Are Now Affordable — the parent article: the 1,000x per-token cost collapse, the Jevons-paradox framing, and the hyperscaler capex context behind this window's price cuts
- Beyond Per-Token: Six Cost Vectors Reshaping Inference Procurement — the procurement framework this article's four clauses slot into; token efficiency is the model-level companion to its vendor-continuity vector
- DeepSeek V4.1-Flash and the Silent Model Swap — the substitution incident that makes the served-model log and the pinned-ID contract necessary rather than prudent
A mid-market distributor running NetSuite and BigCommerce quotes 200 RFQs a week through a governed agent stack: MCP modules for the ERP and three supplier catalogs, a knowledge graph for part substitutes, an RFQ engine handling availability holds — and a routing layer that pins model IDs per step class, caps tokens per task, and logs the served model on every call. When the next launch post claims "30% cheaper," the layer runs a one-week pilot replay and the procurement one-pager gets three numbers instead of one adjective. First agent live in 5–8 weeks.
Request a scoped build. One-week discovery. You get a system inventory, workflow map, and fixed scope — whether or not you build with us.
Want this built for your systems?
Every document here comes from real production work. If you have a target system and a workflow in mind, we can scope a build in one week.
Request a scoped buildOne-week discovery. You get a system inventory, workflow map, and fixed scope — whether or not you build with us.