Ten months of coding-agent tokens

$487K $4.5K

Every LLM call from every harness on every machine, priced twice: once at the meter, once at what was actually paid.

  1. $487,191 Metered API-rate total What 10 months of frontier inference would cost at published per-token prices
  2. $4,500 I actually paid Anthropic $2,500 (MAX 20× + credits) + OpenAI $2,000 (Codex Pro)
  3. $38,083 Employer paid (Azure) Azure AI Foundry — billed to work tenant at full API rates
  4. 112.1× Personal value multiplier Every $1 of personal subscription bought ~$112 of metered inference
From gross meter cost to net wallet cost.

Between September 2025 and July 2026 I ran 483,724 assistant calls across three machines, three CLI harnesses, and forty-six models. At published API rates, the bill would have been $487,191. I actually paid $4,500 out of pocket. My employer covered another $38K in Azure compute.

123,939,382,770 total tokens 112.1× value multiplier Updated July 7, 2026

§ 1 · Where the tokens went

Machine → harness → billing bucket → model.

Flow of spend from machine → harness → billing bucket → provider → model.
Top 10 models by API-rate spend
ModelBucketCost @ API
gpt-5.5 codex-pro $340.4K
claude-opus-4-8 claude-max $77.7K
claude-opus-4-7 claude-max $30.5K
gpt-5.4 codex-pro $14.5K
claude-opus-4-6 claude-max $6.1K
claude-opus-4-6-2 azure $3.2K
claude-opus-4-6-3 azure $3.1K
gpt-5.3-codex codex-pro $2.8K
gpt-5.1-codex codex-pro $2.8K
claude-fable-5 claude-max $2.4K
Machine × harness spend
MachineHarnessCost @ API
personal-mac codex $263.0K
work-mac pi $46.8K
personal-mac claude $43.3K
work-mac codex $41.3K
steambox claude $32.2K
steambox codex $25.2K
steambox pi $16.6K
work-mac claude $12.4K
personal-mac pi $6.4K

The expensive part was not any one model; it was the sheer volume of turns routed through them.

§ 2 · The month-by-month

Ten months of compute, stacked.

Monthly river

Stacked monthly spend at published API rates, split by billing bucket. Toggle to view the same shape as raw tokens.

Show monthly totals (data table)
Month Cost @ API Tokens Calls
2025-09 $15 8,694,437 225
2025-10 $14 6,624,062 294
2025-11 $3,403 2,079,606,188 11,421
2025-12 $23 10,078,265 198
2026-01 $122 59,757,236 481
2026-02 $0 19,493 1
2026-04 $53,849 21,579,985,890 80,336
2026-05 $23,666 6,119,736,461 34,672
2026-06 $365,977 76,093,925,622 296,180
2026-07 $40,122 17,980,955,116 59,916

June 2026 bends the chart because the system stopped being a toy and started behaving like infrastructure.

§ 3 · The silent leverage

Almost none of it was fresh input.

Cache efficiency

Modern coding agents don't send fresh input every turn. They resend a large system prompt + tool schema + running conversation, and most of that hits provider-side caches — Anthropic bills cache reads at 10% of fresh input, OpenAI at 25%. If you don't see this ratio, you don't understand the bill.

Provider Fresh in Cache read Cached ratio
anthropic 94,597,352 tok 34,691,361,525 tok 99.73% cached
openai 28,749,232,527 tok 27,324,158,720 tok 48.73% cached
anthropic-foundry 60,519 tok 11,422,899,174 tok 100.00% cached
openai-codex 364,782,913 tok 10,155,081,984 tok 96.53% cached
azure-anthropic 25,481 tok 6,117,802,198 tok 100.00% cached
codex 24,366,679 tok 752,202,240 tok 96.86% cached
github-copilot 13,318,188 tok 444,981,315 tok 97.09% cached
openai-foundry 222,642,472 tok 332,812,032 tok 59.92% cached
google 14,814,312 tok 190,641,036 tok 92.79% cached
azure-openai-responses 14,609,654 tok 44,438,912 tok 75.26% cached
azure-openai 19,481,120 tok 43,799,808 tok 69.21% cached
google-gemini-cli 7,267,676 tok 40,047,551 tok 84.64% cached
azure-foundry 5,426,430 tok 4,986,112 tok 47.89% cached
lm-studio 63,918,191 tok 0 tok 0.00% cached
openai-foundry-kimi 2,279,317 tok 0 tok 0.00% cached

Cache reads are the invisible margin: the same context kept coming back, but at a fraction of full price.

§ 4 · The pipeline behind this page

Extract, normalize, price, aggregate, render.

Extraction pipeline Three machines fan into a Tailscale-mediated extractor, which normalizes, prices, aggregates, precomputes JSON, and renders this Astro page. machines (JSONL sources) pipeline (stdlib → Polars → Astro) steambox joe-mac-m1 joe-mac-w extract (Tailscale SSH) extract_tokens.py · 3 schemas normalize bucket_of · provider_of price pricing.py rate table aggregate Polars group_by precompute JSON public/tokens/data/*.json Astro + ECharts this page
Fig — the pipeline behind this page

This page is itself a pipeline artifact: extract, price, aggregate, render — then make the method visible.

§ 5 · Every model, every dollar

The full ledger.

Receipts by model

Where the invoice actually landed.

Search by model, filter by billing bucket, and sort the full receipt rollup without exposing raw logs.

Model Bucket Calls Input Output Cache read Cache write Tokens Cost
gpt-5.5 codex-pro 242,439 25B 104M 34B 0 59B $340,418
claude-opus-4-8 claude-max 90,636 87M 157M 26B 1.3B 28B $77,661
claude-opus-4-7 claude-max 36,077 188K 37M 13B 477M 13B $30,530
gpt-5.4 codex-pro 21,965 2.1B 12M 3.0B 0 5.1B $14,505
claude-opus-4-6 claude-max 14,781 140K 6.6M 2.6B 89M 2.7B $6,067
claude-opus-4-6-2 azure 5,706 7.4K 2.2M 1.4B 48M 1.5B $3,209
claude-opus-4-6-3 azure 5,161 5.8K 1.4M 1.6B 32M 1.6B $3,073
gpt-5.3-codex codex-pro 6,106 692M 2.1M 931M 0 1.6B $2,808
gpt-5.1-codex codex-pro 8,260 892M 1.9M 866M 0 1.8B $2,798
claude-fable-5 claude-max 10,586 4.9M 13M 4.8B 191M 5.0B $2,377
claude-sonnet-4-6 claude-max 11,003 392K 6.1M 1.6B 71M 1.7B $832.77
gpt-5.4-2 azure 1,899 120M 2.7M 130M 0 253M $804.32
gpt-5.1-codex-max copilot 2,480 143M 454K 135M 0 279M $539.49
claude-sonnet-5 claude-max 3,150 1.5M 3.3M 534M 57M 596M $428.70
codex-auto-review codex-pro 716 94M 66K 81M 0 175M $228.79
claude-opus-4.6 copilot 236 285 95K 75M 4.0M 79M $194.53
claude-sonnet-4-6-2 azure 1,844 3.0K 771K 426M 9.5M 436M $174.96
gpt-5.2-codex copilot 587 34M 530K 32M 0 67M $137.38
claude-haiku-4-5-20251001 claude-max 13,873 798K 2.5M 450M 55M 507M $101.33
claude-opus-4.8 copilot 93 186 148K 17M 1.4M 19M $62.49
§ 6 · Method + caveats

How the numbers were built.

Session JSONLs from all three machines were extracted in a single pass by a stdlib-only Python script over Tailscale SSH, then repriced against published 2026 API schedules. Anthropic Opus was modeled at $15 / $75, Sonnet at $3 / $15, Haiku at $0.80 / $4; OpenAI GPT-5.5 at $10 / $30 and GPT-5.4 at $5 / $15. Local models (LM Studio / MLX) are treated as cost-zero in the metered view.

Subscription cost is intentionally blunt: Claude MAX 20× at $200/month for ten months, plus Codex Pro at $200/month for ten months, for a flat $4,000 paid. The gap between that flat spend and the metered estimate is the whole point.

Pricing assumptions

  • Anthropic cache_read = 0.10 × input, cache_write = 1.25 × input (90-day cache)
  • OpenAI cached_input ≈ 0.25 × input; no cache_write concept
  • Gemini cached_input ≈ 0.25 × input
  • Reasoning tokens billed at output rate (Anthropic + OpenAI)
  • Local models (qwen/gemma/huihui/mlx/distill/oq8/K2. HF variants) = $0
  • Subscription plan modeled at $200/mo × 2 (Claude MAX 20x + Codex Pro)

Caveats

  • Anthropic API-rate is what the same usage would have cost at published API pricing, NOT what was billed. Operator was on flat $200/mo MAX plan; discount is the arbitrage.
  • Azure + Foundry costs are employer-paid; classified separately from personal-subscription bucket.
  • Provider field is normalized upstream in extract_tokens.py; some historical rows have empty provider (bucketed as 'unknown' / 'other').
  • Sankey trimmed to top-15 models per provider to keep payload <200KB.
  • monthly_by_model filtered to rows with cost>0 OR tokens≥100k to bound size.
  • No session IDs, source paths, cwd, or raw content ship in these aggregates.