AGENTCONN

Field report · · AgentConn Team

Opus Plans, Haiku Executes: The Agent Cost Stack

Anthropic's advisor tool pairs Opus as planner with Haiku as executor. Build the orchestrator-worker cost stack that cuts agent spend without losing quality.

AI AgentsCost OptimizationClaudeOrchestrator PatternModel Routing2026
Orchestrator-worker cost stack: Opus brain directing Haiku execution nodes in a dark minimalist tech illustration

A developer ran an overnight agent orchestrator for a week. The bill: $31,000. The final diff on most tasks: a handful of lines. The root cause was not a bug — it was an architecture decision. Every subagent inherited the parent session’s model. Every model was Opus. Every task, from scanning a directory to formatting a string, burned frontier-model tokens at $5/$25 per million.

This is the agent cost crisis hiding in plain sight. When Anthropic formalized the advisor strategy in April 2026, they were not inventing a new pattern. They were standardizing what production teams had already converged on: the most expensive model should not do the cheapest work.

The orchestrator-worker cost stack is the answer. Opus plans. Haiku executes. And the gap between a $31,000 bill and a $4,200 one is which model you assign to which job.

Anthropic blog — The Advisor Strategy: Give Agents an Intelligence Boost

Read the full post on Anthropic’s blog →

The Advisor Strategy: What Anthropic Actually Ships

Anthropic’s advisor tool is a single API addition — one more entry in the tools array alongside web search and code execution. The setup is three lines:

response = client.messages.create(
    model="claude-sonnet-4-6",  # executor
    tools=[{
        "type": "advisor_20260301",
        "name": "advisor",
        "model": "claude-opus-4-6",
        "max_uses": 3,
    }],
    messages=[...]
)

The executor (Sonnet or Haiku) runs the task end-to-end — calling tools, reading results, iterating toward a solution. When it hits a decision it cannot reasonably solve, it consults Opus. The advisor returns a plan, a correction, or a stop signal. Then the executor resumes. Opus never calls tools. Opus never produces user-facing output. Opus only advises.

This is the inverse of the standard orchestrator-worker pattern. Instead of a smart model decomposing work and delegating to dumb workers, a cheap model drives and escalates only when stuck.

The Numbers That Matter

Anthropic’s own benchmarks tell a specific story:

ConfigurationSWE-bench MultilingualCost per TaskChange
Sonnet 4.6 solo72.1%baseline—
Sonnet 4.6 + Opus advisor74.8%-11.9%+2.7 pts, cheaper
Haiku 4.5 solo (BrowseComp)19.7%very low—
Haiku 4.5 + Opus advisor (BrowseComp)41.2%85% less than Sonnet+21.5 pts

The Haiku result is the headline. Haiku with an Opus advisor more than doubled its standalone BrowseComp score while costing 85% less per task than Sonnet alone. The advisor typically generates only 400-700 text tokens per consultation — a short plan, not a long monologue.

Customer reports back this up. Eve Legal’s case study reports that Haiku 4.5 with Opus consultation matched frontier-model quality on structured document extraction at a fraction of the cost. Bolt saw better architectural decisions on complex tasks with no overhead on simple ones.

The Cost Multiplier Problem Nobody Talks About

The advisor tool is only half the story. The other half is what happens when you do not have a cost stack at all.

Dev.to postmortem — My Agent Orchestrator Burned 1-2M Opus Tokens Per Task

Read the full postmortem on Dev.to →

A detailed postmortem on Dev.to by developer Avraham K. dissected an orchestration skill that burned 1-2 million Opus tokens per task. The diagnosis was not a single bug. It was three modest multipliers stacking:

MultiplierFactorWhat Went Wrong
Model tier~1.7xSubagent model was optional — inherited Opus by default
Cache misses~12xFresh subagents paid cold cache writes instead of cache reads
Fan-out~5x5+ agents per task, each reloading full repo context
Review loops~1-3x”Loop until clean” with no termination bound
Combined100-300xvs. one well-cached Sonnet agent

The fix was not switching models. The fix was treating model assignment as infrastructure, not a default. Every dispatch got an explicit model field. Curated briefs replaced full context dumps. A PreToolUse hook enforced budgets outside the model’s control — because, as the author put it, “a budget rule in a prompt is a preference. A budget rule in a PreToolUse hook is an invariant.”

HackerNoon — You Don't Need Claude Opus for Every Step: Cut Agent Costs by 90%

Read the full article on HackerNoon →

A separate HackerNoon analysis puts the same insight in starker terms: 70-90% of token volume can move to models priced at $0.10-$0.80 per million tokens, with frontier models reserved for the 10-30% of steps that genuinely need deep reasoning. The author’s own pipeline dropped from roughly $31,000 to $4,200 per month — an 86% reduction.

The Three-Tier Cost Stack

The pattern that works in production is not two tiers. It is three.

Tier 1: The Planner (Opus)

Opus decomposes the task, makes architecture decisions, and reviews final output. It touches 10-30% of the steps. Its job is to be right, not fast. At $5/$25 per million tokens (input/output), every Opus call should count.

Use Opus for:

  • Task decomposition and planning
  • Architecture and design decisions
  • Final review and quality gates
  • Ambiguous decisions where the cost of being wrong is high

Tier 2: The Implementer (Sonnet)

Sonnet handles code generation, complex tool orchestration, and multi-step reasoning that Haiku would struggle with. At $3/$15 per million tokens, it is the workhorse for tasks that need judgment but not frontier-level reasoning.

Use Sonnet for:

  • Code generation and refactoring
  • Complex tool orchestration chains
  • Code review where context matters
  • Multi-step reasoning with moderate ambiguity

Tier 3: The Executor (Haiku)

Haiku runs the high-volume, low-complexity work that constitutes the majority of agent token spend. At $1/$5 per million tokens, it is 5x cheaper than Opus on output tokens. Vercel’s model page describes Haiku 4.5 as built for “specialized sub-agent roles inside agentic pipelines.”

Use Haiku for:

  • File search and directory scanning
  • Simple code formatting and linting
  • Classification and extraction
  • Status checks and validation
  • Data retrieval and summarization

The Math

Consider a typical agent sprint with 100 tasks. Without model routing:

ModelTasksInput (MTok)Output (MTok)Cost
Opus (all)1002.00.5$22.50

With three-tier routing:

ModelTasksInput (MTok)Output (MTok)Cost
Opus (plan/review)150.30.08$3.50
Sonnet (implement)250.50.13$3.45
Haiku (execute)601.20.30$2.70
Total1002.00.51$9.65

That is a 57% cost reduction at the same task volume. CloudZero’s analysis models a similar split — one Opus orchestrator with four Sonnet workers at roughly 40% less than five Opus agents.

Configuring the Stack in Practice

Claude Code: Set the Model, Don’t Inherit It

Claude Code resolves a subagent’s model in a fixed precedence order:

  1. CLAUDE_CODE_SUBAGENT_MODEL environment variable (overrides everything)
  2. model parameter at invocation time
  3. model field in the subagent’s frontmatter (.md definition)
  4. Main conversation’s model (the fallback you want to avoid)

The critical mistake is relying on fallback 4. If your session runs Opus, every subagent silently inherits Opus — and you get the $31,000 bill.

# .claude/agents/explorer.md
---
name: explorer
model: haiku
description: Fast file search and context gathering
---
Search the codebase for relevant files and report back.
# .claude/agents/implementer.md
---
name: implementer
model: sonnet
description: Code generation and refactoring
---
Implement the changes described in the plan.

The orchestrator itself runs on Opus (the session model). Workers are explicitly routed.

The Advisor Tool: API-Level Routing

For API users, the advisor tool moves the routing inside the request. No separate orchestration layer. No subagent pool. The executor model handles the task and calls Opus only when it needs guidance:

import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-haiku-4-5",
    max_tokens=8096,
    tools=[
        {
            "type": "advisor_20260301",
            "name": "advisor",
            "model": "claude-opus-4-6",
            "max_uses": 3,
        },
        # your regular tools here
    ],
    system="You are a document extraction agent. "
           "Consult the advisor for ambiguous formatting decisions.",
    messages=[{
        "role": "user",
        "content": "Extract all contract terms from this PDF."
    }],
    betas=["advisor-tool-2026-03-01"],
)

What the Community Is Building

The community is not waiting for the advisor tool. Practitioners are building their own routing layers.

Hacker News — Show HN: Smart model routing directly in Claude, Codex and Cursor (216 points)

View discussion on Hacker News →

Hacker News — Show HN: Rayline routes Claude Code subagents to on-device and cheaper models

View discussion on Hacker News →

LiteLLM’s subtask routing is the most data-rich example. Their routed setup used a cheap exploration model for most turns, Haiku for verification, and Opus only for edits. The result: cost per solved task fell from $0.31 to $0.17 — a 46% saving at the same quality. In the full run, 86% of all turns cost a combined $0.14.

LiteLLM blog — Subtask Type Routing for agent cost optimization

Read the full analysis on LiteLLM →

On Hacker News, Rayline routes Claude Code subagents to on-device and cheaper models. The Claude Code Router (CCR) intercepts API calls and forwards them to any model — cheap models run 20-50x below Claude Sonnet per token. Vercel’s orchestration guide warns that multi-agent systems can use up to 15x the token volume of standard chats — and that under equal compute budgets, single-agent loops often match multi-agent accuracy.

An ICML 2026 paper formalizes the problem as a constrained optimization problem, comparing sequential delegation, parallel fan-out, and hierarchical decomposition on a combined cost-latency objective. The takeaway: coordination overhead can dominate execution cost when orchestrator context grows faster than worker output.

Hacker News — Show HN: Open-source model routing for coding agents at Astra-level performance (123 points)

View discussion on Hacker News →

The lesson: routing by task type works better than routing by prompt difficulty. Your code already knows whether it is planning, extracting, validating, or drafting. Use that signal.

Contrarian Corner: Do You Even Need Multi-Model?

A credible counter-argument exists. The advisor tool saved 11.9% on SWE-bench Multilingual — not 85%. The 85% figure compares Haiku+advisor against Sonnet, not against Opus alone. For workflows where every step genuinely needs frontier reasoning, the overhead of routing — classifier maintenance, integration testing, debugging model-specific failures — may exceed the token savings. One ICML 2026 paper on the "orchestrator bottleneck" found that coordination cost can dominate execution cost when the task is uniformly hard. The honest question: does your workload have a genuine split between easy and hard steps, or are you just hoping it does?

Hacker News — Launch HN: Bullet (YC S26) – A Faster Coding Agent (121 points)

View discussion on Hacker News →

The Five Rules

After reviewing the postmortems, case studies, and production patterns, five rules emerge for building an effective agent cost stack:

1. Set the model explicitly on every dispatch. Never rely on inheritance. The dev who burned $31K did not choose Opus for directory scans. Opus chose itself by default.

2. Route by task type, not prompt difficulty. Classify steps as plan/implement/execute at the architecture level. A lightweight classifier guessing “how hard is this prompt?” is brittle. Your task graph already encodes the answer.

3. Cap your review loops. “Loop until clean” with no termination bound is how review rounds multiply cost 1-3x. Set a maximum — two rounds is enough for most tasks. After two failures, escalate to a human, not a third retry.

4. Enforce budgets outside the model. A cost guard in the system prompt is a suggestion the model can ignore. A PreToolUse hook that blocks dispatches over a threshold is an invariant. AWS recommends tagging orchestrator and worker calls separately and computing the orchestration overhead ratio.

5. Measure the orchestration overhead ratio. Divide orchestrator cost by total workflow cost. If more than 30% of your spend is coordination, your hierarchy is too deep. Flatten it.

Quick start: Set CLAUDE_CODE_SUBAGENT_MODEL=haiku as your environment variable. This single line routes all subagent work to the cheapest model. Most read-only and mechanical tasks produce equivalent results on Haiku.

What This Means for You

If you are building agents today, the cost stack is the highest-leverage optimization available. Not prompt engineering. Not model selection. The architecture of which model does which job.

Start here:

  • Audit your current spend. Tag orchestrator and worker calls. Compute the split. Most teams discover 60-70% of tokens go to execution tasks that do not need a frontier model.
  • Set CLAUDE_CODE_SUBAGENT_MODEL=haiku as your environment variable. This is the simplest possible change — a single line that routes all subagent work to the cheapest model. Test it. Most read-only and mechanical tasks produce equivalent results.
  • Evaluate the advisor tool for API-driven workflows. The max_uses cap gives you direct control over how often the executor escalates. Start with max_uses: 3 and tune from there.
  • Instrument your cost. If you cannot see the per-model, per-task cost breakdown, you cannot optimize it. Build the dashboard before you build the router.

The orchestrator-worker cost stack is not about making agents cheaper. It is about making cheap agents that are still smart enough. Opus for the decisions that matter. Haiku for everything else.


Read more on AgentConn:

The AgentConn Weekly

Weekly digest of new AI agent releases, framework comparisons, and deployment guides. Built for builders.

Weekly. Unsubscribe anytime.

Explore AI Agents

Discover the best AI agents for your workflow in our directory.

Browse Directory