The anatomy of a 1M-token bill: $0.10, itemized to the token
A full agentic coding session — one million prompt tokens processed, one hundred thousand tokens written back. The itemized receipt, line by line, with the exact rates.
The anatomy of a 1M-token bill: $0.10, itemized to the token
Vendors love quoting per-token prices. Almost none will show you the actual bill a real workload produces. So here is ours: a full agentic coding session — the kind where you point an agent at a repo and let it work — with the complete itemized receipt, line by line, at the published rates.
Every number below is USD and comes from the live rate card at GET https://api.runaii.cloud/v1/models. You can reproduce this bill yourself,
request by request, from your own console ledger.
The session
An agent working a medium-sized repo for an afternoon:
- 900,000 prompt tokens processed across ~180 requests — the system prompt, file contents, and tool outputs re-sent with each turn (this is what agents do; it's also exactly what prompt caching is for)
- 100,000 completion tokens written — diffs, file edits, summaries, tool calls
- 1M-token context on
glm-5.3-flash— the whole session fits without compaction
The receipt
| Line item | Tokens | Rate | Cost |
|---|---|---|---|
| Prompt — cache miss | 200,000 | $0.15 / 1M | $0.03000 |
| Prompt — cache hit | 700,000 | $0.03 / 1M | $0.02100 |
| Completion | 100,000 | $0.50 / 1M | $0.05000 |
| Total | 1,000,000 | $0.10100 |
That's the whole bill. Ten cents for a million tokens of agentic work.
Where each fraction goes
The cache hit line is the story. 700,000 of the 900,000 prompt tokens were re-sends of prefixes the model had already seen — the system prompt, the file tree, tool outputs from earlier turns. At the published rates those billed at the cache rate, 5× cheaper than a miss:
| Bucket | Rate (glm-5.3-flash) |
|---|---|
| Cache miss (first time a prefix is seen) | $0.15 / 1M tokens |
| Cache hit (same prefix seen again) | $0.03 / 1M tokens — 5× cheaper |
Without caching, the same session costs:
0.900 × $0.15 = $0.135 (prompt, all miss)
0.100 × $0.50 = $0.050 (completion)
= $0.185 total
With caching: $0.101. The split is automatic and reported per request in
usage.prompt_tokens_details.cached_tokens — you can see exactly how many
tokens hit the cache on every request.
What the same session costs on the frontier sibling
Same session, glm-5.3 (the frontier agentic model, same 1M context):
| Line item | Tokens | Rate | Cost |
|---|---|---|---|
| Prompt — cache miss | 200,000 | $1.40 / 1M | $0.28000 |
| Prompt — cache hit | 700,000 | $0.26 / 1M | $0.18200 |
| Completion | 100,000 | $4.40 / 1M | $0.44000 |
| Total | $0.90200 |
Both rows are the same bill shape; pick the model per job. Flash for high-volume pipelines, frontier for the hard steps. The rate card carries both.
What happens when a stream dies mid-flight
You pay for delivered tokens only. Every request moves through three money events:
- Authorize — the estimated worst case is reserved from your balance, atomically. Two concurrent requests cannot double-spend the same credits.
- Settle — when the stream finishes, the reservation is settled against actual usage: prompt tokens at the full rate, cached tokens at the cache rate, completion tokens at the output rate.
- Reverse — if the request died mid-stream, the whole reservation is refunded. You never pay for a request that never delivered.
In this session, 4 of the 184 requests hit provider errors. All 4 show up as explicit refund events in the ledger — $0.00041 reversed in full, visible line items, not silent adjustments.
Verify this receipt yourself
usage.costin every response — per-request truth- Console → Usage — per-model spend, latency percentiles, daily charts
- Console → Billing — the full ledger: every charge, top-up, and refund, with the request id that caused it
Re-run this session shape yourself and compare the receipt to this post. If the ledger doesn't match the post, the receipt wins and the post gets corrected.
This post is part of the runaii.cloud transparency series. The rates above are
the live output of GET /api/v1/models as of 2026-10-01 — see the price
comparison post for the full rate
card, and the billing engine post for the
reserve → settle → refund internals.