Token benchmarks
Historical first-call input measurements, their source data, and limits of the comparison.
Measured 2026-09-02. These are historical results, not measurements of the current releases.
A model request includes more than your message. It can contain system instructions, tool definitions, project context, and conversation history. This chart records input tokens for the first model call using the prompt "Reply with exactly OK and nothing else." and model {benchmarkRun.model} through OpenRouter.
The OpenWaggle row comes from a separate bundled-Pi-SDK probe, not a fully configured GUI session. The other seven rows come from container runs through a logging proxy. Different tool profiles and provider routes limit direct comparisons.
The largest value is 17,608 tokens. Short bars have a minimum visible width of 6% and enough room for their label, so use the printed counts for exact comparisons.
Details worth knowing
- The container runs used empty projects. Real repositories can add instructions, repository maps, skills, extensions, and MCP tools.
- The DeepSeek Harness row used its
sdk-minimalprofile, not its full interactive toolset. - The Reasonix archive contains one model call with 5,279 input tokens, while its CLI reported 10,558. The chart uses the recorded API usage. The archive does not establish why its CLI reported double.
- Codex used an explicit
model_reasoning_effortsetting for its OpenRouter Responses request. - OpenWaggle and Pi CLI used different provider routes. The recorded difference does not by itself establish that tool serialization was the only cause.
How we measured
The seven non-OpenWaggle rows used pinned versions in separate Docker containers, fresh home directories, and empty projects. The archived proxy logs contain their first-call API usage. The parser counts prompt tokens, including the probe message. For Anthropic-style usage, it adds uncached input, cache writes, and cache reads.
OpenWaggle was measured separately with scripts/benchmark-first-turn-tokens.ts, which creates temporary project and agent directories and reads assistant usage from the bundled Pi SDK. Its published value is retained in the website dataset. The container archive does not include the raw output for that separate run, so its value was not independently reverified from those logs during the documentation audit.
These are individual recorded runs, not repeated trials with error estimates. The chart measures first-call input size, not total task usage or a bill. Provider pricing and caching policies affect cost.
What this benchmark is not
This does not measure code quality, task completion, latency, or equivalent capabilities. An agent with fewer initial tokens can still use more tokens over a task or produce a worse result.
The historical OpenWaggle measurement used Pi 0.84.4. The current repository uses a newer runtime and additional app integrations. Do not use this chart as a current-release performance claim.
Run it yourself
The two measurement paths are separate:
scripts/benchmark-docker/contains the historical container runner and proxy. It also runs an older Pi 0.81.1 comparison that is not plotted here. Its parser writesscripts/benchmark-docker/results/results.json; it does not update the website dataset automatically.scripts/benchmark-first-turn-tokens.tsinspects the bundled Pi SDK. Its offline mode estimates tokens from the prompt and tool schema text. Live mode reads provider-reported usage.
The container runner currently needs fixes before it can support repeatable reruns. Its proxy also binds to all interfaces. Do not run it with credentials on an untrusted network. The archived counts remain available for inspection, but the runner is not a supported reproduction command.
For an offline SDK inspection from a development checkout with dependencies installed:
pnpm exec tsx scripts/benchmark-first-turn-tokens.ts --no-skills
A live SDK measurement requires provider credentials in ~/.pi/agent/auth.json, access to the requested model, and network access. It sends a request to the provider and can incur charges:
pnpm exec tsx scripts/benchmark-first-turn-tokens.ts --live --no-skills --second-turn \
--provider openrouter --model z-ai/glm-5.3-flash
Running today’s checkout produces a new measurement. It does not reproduce the historical runtime or establish the overhead of a configured OpenWaggle session. Keep the source revision, versions, configuration, route, and raw usage output with any new result.
Versions measured
- Pi CLI 0.84.4
- OpenWaggle 0.3.0-alpha.63
- Aider 0.86.2
- DeepSeek Harness 0.1.2a3, sdk-minimal profile
- DeepSeek Reasonix 1.35.0
- opencode 1.18.26
- Codex CLI 0.150.1
- Claude Code 2.1.247