/00 — boot sequence

Hello.

Article

Claude Code Token Overhead: 4.7x More vs OpenCode

July 13, 2026•7 min read
claude-code opencode ai-coding token-efficiency developer-tools llm

If you use Claude Code for AI-assisted development, here is a number that might surprise you. Before you even type a prompt, Claude Code sends roughly 33,000 tokens to the model. That is 4.7 times more than OpenCode, an open-source alternative that sends about 7,000 tokens for the same setup.

Systima AI ran a controlled benchmark comparing Claude Code 2.1.207 and OpenCode 1.17.18 on identical hardware, using the same Sonnet 4.5 model. The results reveal where those extra tokens come from, how they affect your costs, and what you can do about it.

Background

AI coding agents come with a hidden cost. Every time you open a session, the agent loads its system prompt, tool definitions, configuration files, and environment context before your actual work begins. This scaffolding ensures the agent knows what tools it has available and how to behave, but it also consumes your context window and your token budget.

Until now, developers had limited visibility into what their coding agent actually sent to the model. Systima's logging proxy changed that by capturing exact request payloads and API usage blocks at the boundary between each harness and the model endpoint.

The Floor: Fixed Overhead Per Session

The simplest measurement is the floor. What happens when you ask each harness to "Reply with exactly: OK"?

Claude Code sent 33,000 tokens on its first request. OpenCode sent about 6,900. The dominant factor is tool schemas. Claude Code ships 27 tools comprising 99,778 characters of schema definitions roughly 24,000 tokens. OpenCode ships 10 tools consuming about 4,800 tokens.

Claude Code carries an entire orchestration suite behind its tools, including worktree management, push notifications, subagent families, and an internal agent catalogue. OpenCode keeps its tool set lean with only the essentials.

Even with tools stripped entirely, Claude Code's system prompt alone is 6,500 tokens versus OpenCode's 2,000. The residual is behavioural doctrine, safety guidance, and environment description.

The Multipliers: Where Costs Really Add Up

The floor tells only part of the story. Real development sessions layer configuration on top of the baseline, and each layer multiplies the cost.

Instruction Files

A 72KB instruction file such as CLAUDE.md or AGENTS.md adds roughly 20,000 tokens to every single request in both harnesses. One critical caveat emerged during testing: Claude Code 2.1.207 ignored a file named AGENTS.md entirely and only read it when renamed to CLAUDE.md. An ignored instruction file is silent, so verify which filename your harness honours.

MCP Servers

Each MCP server adds 1,000 to 1,400 tokens per request depending on schema size. Five modest servers grow the tool count from 27 to 69 in Claude Code and from 10 to 52 in OpenCode. Production servers with rich APIs ship schemas several times larger.

Subagents

Subagent delegation is the single largest token multiplier. A small task that cost 121,000 tokens when done directly cost 513,000 tokens when fanned out to two subagents, a 4.2x multiplier. Every subagent pays its own bootstrap cost, and the parent must then ingest the subagent's entire transcript.

The Everything Number

With 11 MCP servers plus a 72KB instruction file, OpenCode's first cold-cache request was 90,817 tokens carrying 179 tools. Claude Code with 4 MCP servers plus plugins hit roughly 85,000 tokens. Your configuration sets the bill.

Comparison Table

MetricClaude CodeOpenCode
System prompt (no tools)6,500 tokens2,000 tokens
First request baseline~33,000 tokens~7,000 tokens
Tools shipped27 (99,778 chars)10 (20,856 chars)
Cache prefix stabilityUnstable, mid-session rewritesByte-identical, stable
Cache write volume (same task)5.9x to 54x more than OpenCodeBaseline 1,003 tokens
Multi-step task efficiencyBatches aggressively, fewer requestsSerial calls, more requests
Multi-step total inputConverges with OpenCodeConverges with Claude Code
MCP server tax~1,000-1,400 tokens per server~1,000-1,400 tokens per server
Subagent multiplier4.2x (two subagents)Leaner per-agent design

Cache Stability: The Decisive Difference

Prompt caching should offset these costs, but only if the request prefix stays stable. OpenCode emitted byte-identical prefixes across every request and every run. Three separate sessions produced identical tool bytes, system bytes, and message bytes.

Claude Code emitted three distinct request classes per session: a warmup probe, the main conversation, and subagent calls. Each has its own prefix and its own cache entry. On the identical file-summarise task, Claude Code wrote 53,839 cache tokens across five requests including one complete mid-task re-write of its full 43,000-token prefix. OpenCode wrote 1,003.

The behaviour reproduced across models. On Claude Fable 5, the same pattern appeared: another complete mid-session re-write of 50,053 tokens with zero cache read. Cache writes bill at a premium of 1.25x base rate for the 5-minute tier.

Developer Impact

A 33,000-token baseline means every turn starts a sixth of the way into a 200k context window before any code enters the conversation. With a heavy instruction file and MCP servers, that grows to 85,000 tokens occupying 42 per cent of the window.

For teams running Claude Code in production, particularly under regulations like the EU AI Act's Article 12 which expects you to log and understand your system behaviour, knowing what your agent actually sends is essential data rather than folklore.

The good news is that on multi-step tasks the gap closes significantly. Claude Code's aggressive parallel tool batching compresses what would be nine separate round trips into three, so the per-turn baseline penalty is paid fewer times. On a write-run-test-fix loop, the cumulative totals converged between both harnesses.

Best Practices for Minimizing Token Overhead

Keep your instruction file lean. Every kilobyte in CLAUDE.md or AGENTS.md is paid on every turn. Audit which filename your harness actually reads, because an ignored file gives you neither guidance nor overhead.

Review your MCP server list. Five small servers add 5,000 to 7,000 tokens per request. Production servers with rich schemas cost more. Remove any server your workflow does not need for the current task.

Be deliberate about subagent delegation. A 4.2x cost multiplier for two subagents means subagent fan-out is powerful but expensive. Use it only when parallel work truly adds value.

Check your cache hit rate. If your tool schemas or system prompt are changing between requests, you are paying write rates for every session rather than read rates.

Frequently Asked Questions

Does the gap matter if I use prompt caching? Yes. Three costs survive caching: the initial cache write (repaid after any 5-minute pause), the per-turn cache read (multiplied by request count), and context-window consumption (completely immune to caching).

Is Claude Code worse on newer models? The gap narrows. On Claude Fable 5, Claude Code's floor dropped to about 3.3x instead of 4.7x because Anthropic sends newer models a much smaller system prompt.

Is OpenCode always cheaper then? Not always. On multi-step tasks with many tool calls, Claude Code's aggressive parallel batching means it makes fewer requests. The smaller baseline Advantage of OpenCode is multiplied by more requests, and the totals can converge.

Can I reduce Claude Code's overhead today? Yes. Trim your CLAUDE.md file, audit your MCP server list, avoid unnecessary subagent fan-out, and check whether your cache prefix stays stable across requests within a session.

Key Takeaways

  • Claude Code sends 33,000 tokens before your prompt versus OpenCode's 7,000, a 4.7x gap on Sonnet 4.5
  • Tool schemas are the dominant term: 24,000 tokens in Claude Code versus 4,800 in OpenCode
  • Subagent delegation is the largest token multiplier at 4.2x for a two-subagent fan-out
  • Claude Code's cache prefixes are unstable, leading to up to 54x more cache-write tokens than OpenCode
  • On multi-step tasks with batching, the cost gap converges significantly
  • The gap narrows on newer models (3.3x on Fable 5) but remains substantial
  • Keep your instruction file lean and audit which filename your harness actually reads

Conclusion

Token overhead in AI coding agents is not just a cost issue it is a context-window budget issue. Every token spent on scaffolding is a token you cannot spend on understanding your codebase or generating solutions. The 4.7x gap between Claude Code and OpenCode means Claude Code sessions consume significantly more context before any productive work begins.

The surprise finding is that Claude Code's aggressive batching means the gap can close on complex multi-step tasks, and the cash difference for teams using prompt caching is smaller than the raw token counts suggest. But the context-window cost remains, and cache instability means Claude Code users are more likely to pay write rates than read rates.

Knowing what your agent actually sends is the first step to optimizing it. Measure your own setup, trim where you can, and choose the harness that matches your workflow.


Sources: Systima AI - Claude Code vs OpenCode Token Overhead Benchmark

Automated Transmission

This entry was synthesized and populated dynamically using native API integrations.

Resources & Links