The $0.10 Subagent: Claude Haiku 5.5 Collapses the Frontier-vs-Open-Weight Cost Gap
Key takeaways
- Claude Haiku 5.5 is priced at $0.10/$0.50 per million tokens for prompts under 100K — 90% below Haiku 4.5 and identical to GPT-6 Luna's pricing — making the subagent tier effectively free for 90% of requests across three providers (Anthropic, Oct 7).
- OSWorld 2.1 jumped from 15.7% (Haiku 4.5) to 72.4% (Haiku 5.5) — a 4.6× improvement that puts the cheap model within 12 points of Sonnet 5.5's 83.9% (Anthropic, Oct 7).
- Sonnet 5.5 cache reads were halved from $0.20 to $0.10 per million tokens, cutting Sonnet 5.5's agentic cost by ~20% — the mid-tier model got cheaper the same day the subagent tier launched (Anthropic, Oct 7).
- GLM-5.3-Flash at $0.15/$0.50 and GPT-6 Luna at $0.10/$0.50 mean the cost gap between frontier-closed and open-weight in the small-model tier is now $0.05/M input — a number that rounds to zero for most workloads (Z.AI; Anthropic).
- The competitive axis has shifted from price to capability, safety posture, and ecosystem integration — the subagent tier is no longer a cost decision, it is an architecture decision (Yahoo Finance, Oct 7).
This builds on 30% Cheaper Per Task Costs 7x the Tokens: Inside Sonnet 5.5's Efficiency Trap, which mapped the three release strategies competing on cost-per-task at the flagship tier. Here the focus is one layer down: the subagent tier, where the model that handles compaction, classification, database queries, and coding sidekick work just hit a price floor that was unthinkable six months ago. The Sonnet 5.5 cache-read halving — shipped in the same announcement — is the direct companion datapoint: the article's token-efficiency thesis gains a pricing change that makes the mid-tier model 20% cheaper on the agentic workloads where cache reads dominate.
The subagent tier is now a three-vendor market at one price
The pricing table for the small-model tier as of October 8, 2026:
| Model | Input $/M | Output $/M | Type | Provider |
|---|---|---|---|---|
| GPT-6 Luna | $0.10 | $0.50 | Frontier closed | OpenAI |
| Claude Haiku 5.5 | $0.10 | $0.50 | Frontier closed | Anthropic |
| GLM-5.3-Flash | $0.15 | $0.50 | Open-weight | Z.ai |
The output price is identical across all three. The input gap between the frontier-closed pair and the open-weight model is $0.05 per million tokens — for a subagent that compacts a 50,000-token conversation context, that is $0.0025. The cost of choosing the wrong subagent model is now measured in fractions of a cent per task, not dollars per day.
The subagent tier pricing landscape as of October 8, 2026:
Yahoo Finance framed the Haiku 5.5 launch as "AI pricing war intensifies" (Yahoo Finance, Oct 7). The framing is accurate but understates what happened. A pricing war implies margins compressing on the same product. What actually happened is that the subagent tier — the layer where most agentic token consumption occurs — collapsed to a price that was previously only available from the cheapest open-weight model. The frontier closed model and the open-weight model now sit at the same price point for output tokens. The question "should I use a frontier or an open-weight subagent" is no longer a cost question.
What changed: the subagent got capable
The price is only half the story. Haiku 5.5's benchmark jump is the part that changes architecture decisions:
- OSWorld 2.1 (offline subset): 72.4% — up from Haiku 4.5's 15.7%, a 4.6× improvement. GPT-6 Luna sits at 48.9%. The subagent can now operate a computer to finish multi-step tasks at a level that was flagship-tier twelve months ago (Anthropic).
- Terminal-Bench 4.0: 39.2% — up from Haiku 4.5's 0.0%. GPT-6 Luna sits at 16.4%. The subagent can complete complex command-line tasks that the previous generation could not attempt (Anthropic).
- FrontierCode 1.1 (Main): 46.4% — above GPT-6 Luna's 42.4%, below Sonnet 5.5's 52.1% at xhigh. The subagent writes production code well enough to serve as a coding sidekick — Cognition ships it as the sidekick in Devin Fusion, holding a top-tier FrontierCode score of 66.2 at reduced cost and latency (Anthropic).
- HubSpot CRM task suite: 92.8% — the best score HubSpot has measured on its simulated-portal CRM evaluation. The subagent identifies stale but ambiguous records with the highest hit rate and lowest false positive rate of any model tested (Anthropic).
The HubSpot datapoint matters beyond the benchmark. HubSpot is a major CRM platform evaluating frontier small models for CRM-automation tasks — the exact use case that a custom MCP module addresses. When the subagent tier can identify stale CRM records at 92.8% accuracy for $0.10/M input, the cost of running an agent that monitors a deal pipeline 24/7 drops from a budget line item to a rounding error.
Adjustable effort: the same model, five price points
Haiku 5.5 is the first Haiku-class model with adjustable effort settings: low, medium, high, xhigh, and max. This is the same configurable-cost pattern the DeepSeek auto-routing analysis flagged as a procurement risk when the routing is invisible — but here the control is in the operator's hands. A subagent that compacts conversation context can run at low effort for $0.002 per task; the same model can run at max effort for a harder database query. The cost-per-task curve is now a dial, not a fixed rate.
The implication for inference procurement is direct: a contract that names only the model and the rate card has no control over the effort setting. An agent running at max effort burns more tokens than the same agent at medium — and the operator sees only the invoice. The four contract clauses the Sonnet 5.5 article recommended (effort-level pinning, substitution disclosure, cache-read pricing floors, token-volume ceilings) apply equally to the subagent tier now that the effort dial exists there too.
The Sonnet 5.5 cache-read halving: the mid-tier got cheaper too
The same announcement halved Sonnet 5.5's cache reads from $0.20 to $0.10 per million tokens. Cache reads make up a large share of agentic token consumption — an agent that re-reads the same system prompt, tool definitions, and conversation history on every turn hits the cache-read line item more than the input or output line items. A 50% cut on the line item that dominates agentic spending is a ~20% reduction in Sonnet 5.5's total agentic cost (Anthropic).
This matters for the agent architecture decision. The subagent tier at $0.10/$0.50 handles compaction, classification, and sidekick work. The mid-tier at $2/$10 — now with halved cache reads — handles the complex agentic coding and multi-step reasoning. The flagship tier at $4/$20 handles the hardest tasks. The three-tier stack is now cleanly priced: cheap subagent, capable mid-tier with cheap cache, powerful flagship. The cost of running the full stack for an always-on production agent dropped meaningfully on October 7.
What this means for the build
The inference economics analysis established the 10:1 ratio: $10 of process, governance, and integration work for every $1 of model spend. The Haiku 5.5 release widens that ratio further. The model got 90% cheaper; the integration layer did not. An MCP module that connects an agent to NetSuite, a knowledge graph that resolves substitute parts, an audit trail that logs every tool call — these costs are unchanged. The model is now a smaller fraction of the total build cost than it was a week ago.
This is the procurement insight. The subagent tier pricing collapse does not make agent systems cheaper to build. It makes them cheaper to run — which means the teams that have already built the integration layer get a larger cost reduction than teams that have not. A team running a production RFQ agent with typed MCP modules, a knowledge graph, and human-in-the-loop checkpoints sees its monthly inference bill drop by 75% overnight. A team that has not built the integration layer sees no change — because the integration layer is where the cost lives.
Related reading
- 30% Cheaper Per Task Costs 7x the Tokens: Inside Sonnet 5.5's Efficiency Trap — the parent article mapping the three release strategies competing on cost-per-task; the cache-read halving is the direct companion datapoint
- Inference Economics: Why Always-On Production Agents Are Now Affordable — the 10:1 ratio of integration to model cost; the subagent-tier collapse widens the gap
- GLM-5.3-Flash: A 320B Open-Weight Model Approaches Claude Opus 4.8 at One-Tenth the Flagship Price — the open-weight half of the cost-collapse thesis; GLM-5.3-Flash at $0.15/$0.50 is the third vendor in the three-provider subagent tier
A regional distributor running NetSuite, BigCommerce, and three supplier catalogs processes 200 RFQs a week. The quoting agent runs Haiku 5.5 as the subagent for catalog lookup and price-tier classification, Sonnet 5.5 with halved cache reads for the quote-drafting step, and a knowledge graph for substitute-part resolution. The monthly inference cost for the full stack dropped 75% on October 7 — from roughly $4,000 to under $1,000 — without changing a single line of integration code. The MCP modules, the knowledge graph, and the audit trail are what make the cost reduction real. The model price is the easy part.
Request a scoped build
One-week discovery. You get a system inventory, workflow map, and fixed scope — whether or not you build with us.
Want this built for your systems?
Every document here comes from real production work. If you have a target system and a workflow in mind, we can scope a build in one week.
Request a scoped buildOne-week discovery. You get a system inventory, workflow map, and fixed scope — whether or not you build with us.