Ten months of coding-agent tokens
$487K → $4.5K
Every LLM call from every harness on every machine, priced twice: once at the meter, once at what was actually paid.
- $487,191 Metered API-rate total What 10 months of frontier inference would cost at published per-token prices
- $4,500 I actually paid Anthropic $2,500 (MAX 20× + credits) + OpenAI $2,000 (Codex Pro)
- $38,083 Employer paid (Azure) Azure AI Foundry — billed to work tenant at full API rates
- 112.1× Personal value multiplier Every $1 of personal subscription bought ~$112 of metered inference
Between September 2025 and July 2026 I ran 483,724 assistant calls across three machines, three CLI harnesses, and forty-six models. At published API rates, the bill would have been $487,191. I actually paid $4,500 out of pocket. My employer covered another $38K in Azure compute.
123,939,382,770 total tokens 112.1× value multiplier Updated July 7, 2026
Machine → harness → billing bucket → model.
Top 10 models by API-rate spend
| Model | Bucket | Cost @ API |
|---|---|---|
| gpt-5.5 | codex-pro | $340.4K |
| claude-opus-4-8 | claude-max | $77.7K |
| claude-opus-4-7 | claude-max | $30.5K |
| gpt-5.4 | codex-pro | $14.5K |
| claude-opus-4-6 | claude-max | $6.1K |
| claude-opus-4-6-2 | azure | $3.2K |
| claude-opus-4-6-3 | azure | $3.1K |
| gpt-5.3-codex | codex-pro | $2.8K |
| gpt-5.1-codex | codex-pro | $2.8K |
| claude-fable-5 | claude-max | $2.4K |
Machine × harness spend
| Machine | Harness | Cost @ API |
|---|---|---|
| personal-mac | codex | $263.0K |
| work-mac | pi | $46.8K |
| personal-mac | claude | $43.3K |
| work-mac | codex | $41.3K |
| steambox | claude | $32.2K |
| steambox | codex | $25.2K |
| steambox | pi | $16.6K |
| work-mac | claude | $12.4K |
| personal-mac | pi | $6.4K |
The expensive part was not any one model; it was the sheer volume of turns routed through them.
Ten months of compute, stacked.
Monthly river
Stacked monthly spend at published API rates, split by billing bucket. Toggle to view the same shape as raw tokens.
Show monthly totals (data table)
| Month | Cost @ API | Tokens | Calls |
|---|---|---|---|
| 2025-09 | $15 | 8,694,437 | 225 |
| 2025-10 | $14 | 6,624,062 | 294 |
| 2025-11 | $3,403 | 2,079,606,188 | 11,421 |
| 2025-12 | $23 | 10,078,265 | 198 |
| 2026-01 | $122 | 59,757,236 | 481 |
| 2026-02 | $0 | 19,493 | 1 |
| 2026-04 | $53,849 | 21,579,985,890 | 80,336 |
| 2026-05 | $23,666 | 6,119,736,461 | 34,672 |
| 2026-06 | $365,977 | 76,093,925,622 | 296,180 |
| 2026-07 | $40,122 | 17,980,955,116 | 59,916 |
June 2026 bends the chart because the system stopped being a toy and started behaving like infrastructure.
Almost none of it was fresh input.
Cache efficiency
Modern coding agents don't send fresh input every turn. They resend a large system prompt + tool schema + running conversation, and most of that hits provider-side caches — Anthropic bills cache reads at 10% of fresh input, OpenAI at 25%. If you don't see this ratio, you don't understand the bill.
| Provider | Fresh in | Cache read | Cached ratio |
|---|---|---|---|
| anthropic | 94,597,352 tok | 34,691,361,525 tok | 99.73% cached |
| openai | 28,749,232,527 tok | 27,324,158,720 tok | 48.73% cached |
| anthropic-foundry | 60,519 tok | 11,422,899,174 tok | 100.00% cached |
| openai-codex | 364,782,913 tok | 10,155,081,984 tok | 96.53% cached |
| azure-anthropic | 25,481 tok | 6,117,802,198 tok | 100.00% cached |
| codex | 24,366,679 tok | 752,202,240 tok | 96.86% cached |
| github-copilot | 13,318,188 tok | 444,981,315 tok | 97.09% cached |
| openai-foundry | 222,642,472 tok | 332,812,032 tok | 59.92% cached |
| 14,814,312 tok | 190,641,036 tok | 92.79% cached | |
| azure-openai-responses | 14,609,654 tok | 44,438,912 tok | 75.26% cached |
| azure-openai | 19,481,120 tok | 43,799,808 tok | 69.21% cached |
| google-gemini-cli | 7,267,676 tok | 40,047,551 tok | 84.64% cached |
| azure-foundry | 5,426,430 tok | 4,986,112 tok | 47.89% cached |
| lm-studio | 63,918,191 tok | 0 tok | 0.00% cached |
| openai-foundry-kimi | 2,279,317 tok | 0 tok | 0.00% cached |
Cache reads are the invisible margin: the same context kept coming back, but at a fraction of full price.
Extract, normalize, price, aggregate, render.
This page is itself a pipeline artifact: extract, price, aggregate, render — then make the method visible.
The full ledger.
Receipts by model
Where the invoice actually landed.
Search by model, filter by billing bucket, and sort the full receipt rollup without exposing raw logs.
| Model | Bucket | Calls | Input | Output | Cache read | Cache write | Tokens | Cost |
|---|---|---|---|---|---|---|---|---|
| gpt-5.5 | codex-pro | 242,439 | 25B | 104M | 34B | 0 | 59B | $340,418 |
| claude-opus-4-8 | claude-max | 90,636 | 87M | 157M | 26B | 1.3B | 28B | $77,661 |
| claude-opus-4-7 | claude-max | 36,077 | 188K | 37M | 13B | 477M | 13B | $30,530 |
| gpt-5.4 | codex-pro | 21,965 | 2.1B | 12M | 3.0B | 0 | 5.1B | $14,505 |
| claude-opus-4-6 | claude-max | 14,781 | 140K | 6.6M | 2.6B | 89M | 2.7B | $6,067 |
| claude-opus-4-6-2 | azure | 5,706 | 7.4K | 2.2M | 1.4B | 48M | 1.5B | $3,209 |
| claude-opus-4-6-3 | azure | 5,161 | 5.8K | 1.4M | 1.6B | 32M | 1.6B | $3,073 |
| gpt-5.3-codex | codex-pro | 6,106 | 692M | 2.1M | 931M | 0 | 1.6B | $2,808 |
| gpt-5.1-codex | codex-pro | 8,260 | 892M | 1.9M | 866M | 0 | 1.8B | $2,798 |
| claude-fable-5 | claude-max | 10,586 | 4.9M | 13M | 4.8B | 191M | 5.0B | $2,377 |
| claude-sonnet-4-6 | claude-max | 11,003 | 392K | 6.1M | 1.6B | 71M | 1.7B | $832.77 |
| gpt-5.4-2 | azure | 1,899 | 120M | 2.7M | 130M | 0 | 253M | $804.32 |
| gpt-5.1-codex-max | copilot | 2,480 | 143M | 454K | 135M | 0 | 279M | $539.49 |
| claude-sonnet-5 | claude-max | 3,150 | 1.5M | 3.3M | 534M | 57M | 596M | $428.70 |
| codex-auto-review | codex-pro | 716 | 94M | 66K | 81M | 0 | 175M | $228.79 |
| claude-opus-4.6 | copilot | 236 | 285 | 95K | 75M | 4.0M | 79M | $194.53 |
| claude-sonnet-4-6-2 | azure | 1,844 | 3.0K | 771K | 426M | 9.5M | 436M | $174.96 |
| gpt-5.2-codex | copilot | 587 | 34M | 530K | 32M | 0 | 67M | $137.38 |
| claude-haiku-4-5-20251001 | claude-max | 13,873 | 798K | 2.5M | 450M | 55M | 507M | $101.33 |
| claude-opus-4.8 | copilot | 93 | 186 | 148K | 17M | 1.4M | 19M | $62.49 |
How the numbers were built.
Session JSONLs from all three machines were extracted in a single pass by a stdlib-only Python script over Tailscale SSH, then repriced against published 2026 API schedules. Anthropic Opus was modeled at $15 / $75, Sonnet at $3 / $15, Haiku at $0.80 / $4; OpenAI GPT-5.5 at $10 / $30 and GPT-5.4 at $5 / $15. Local models (LM Studio / MLX) are treated as cost-zero in the metered view.
Subscription cost is intentionally blunt: Claude MAX 20× at $200/month for ten months, plus Codex Pro at $200/month for ten months, for a flat $4,000 paid. The gap between that flat spend and the metered estimate is the whole point.
Pricing assumptions
- Anthropic cache_read = 0.10 × input, cache_write = 1.25 × input (90-day cache)
- OpenAI cached_input ≈ 0.25 × input; no cache_write concept
- Gemini cached_input ≈ 0.25 × input
- Reasoning tokens billed at output rate (Anthropic + OpenAI)
- Local models (qwen/gemma/huihui/mlx/distill/oq8/K2. HF variants) = $0
- Subscription plan modeled at $200/mo × 2 (Claude MAX 20x + Codex Pro)
Caveats
- Anthropic API-rate is what the same usage would have cost at published API pricing, NOT what was billed. Operator was on flat $200/mo MAX plan; discount is the arbitrage.
- Azure + Foundry costs are employer-paid; classified separately from personal-subscription bucket.
- Provider field is normalized upstream in extract_tokens.py; some historical rows have empty provider (bucketed as 'unknown' / 'other').
- Sankey trimmed to top-15 models per provider to keep payload <200KB.
- monthly_by_model filtered to rows with cost>0 OR tokens≥100k to bound size.
- No session IDs, source paths, cwd, or raw content ship in these aggregates.