OpenAI-compatible inference APIs in 2026: how to compare price per token (with receipts)
A worked example you can reproduce: glm-5.3-flash at $0.15/1M in, $0.03/1M cached, $0.50/1M out — and the three questions to ask any inference vendor's pricing page.
OpenAI-compatible inference APIs in 2026: how to compare price per token (with receipts)
Every inference vendor posts a "pricing" page. Very few make that page checkable. This post shows you how to audit any OpenAI-compatible inference API in ten minutes, using our own published rates as the worked example — and then gives you the checklist to run the same audit against every vendor you're considering (including us).
All numbers below are USD per 1M tokens and come from GET https://api.runaii.cloud/v1/models — the live price API. We publish rates
there rather than hardcoding them on a page because prices can change, and the
API is the source of truth.
The runaii.cloud rate card (2026-10-01)
| Model | Context | Prompt $/1M | Cached $/1M | Completion $/1M |
|---|---|---|---|---|
glm-5.3-flash |
1M | $0.15 | $0.03 | $0.50 |
glm-5.3 |
1M | $1.40 | $0.26 | $4.40 |
kimi-k3 |
1M | $3.00 | — | $15.00 |
deepseek-v4-pro |
1M | $1.32 | $0.044 | $3.96 |
deepseek-v4-flash |
1M | $0.22 | $0.022 | $0.66 |
qwen3.8-max |
262k | $2.00 | — | $6.00 |
gpt-oss-120b |
131k | $0.15 | $0.015 | $0.60 |
minimax-m3 |
512k | $0.30 | $0.03 | $1.20 |
llama-3.3-70b |
131k | $0.23 | $0.023 | $0.85 |
Also listed: qwen3-reranker-8b (reranking), whisper-v3 (transcription) and
flux-kontext-pro (image generation) — dedicated endpoints rolling out. Every
row above is the price a real request is charged; there is no "list price" vs
"actual price" gap. The API is the source of truth and prices can change.
A worked example you can reproduce
1,000 miss input tokens + 50,000 cached input tokens + 2,000 output tokens on
glm-5.3-flash:
0.001 × $0.15 = $0.00015
0.050 × $0.03 = $0.00150
0.002 × $0.50 = $0.00100
= $0.00265 total
Run the identical request yourself and check usage.cost in the response —
every response reports exactly what it cost, itemized to the token. If the
receipt doesn't match the rate card, the receipt wins and this post gets
corrected.
The three questions to ask any inference pricing page
1. "Is the cache discount real, and is it passed through 1:1?"
Many vendors claim prompt caching "supported" but bill cache hits at full
rate. On runaii.cloud, cache hits bill at the subsidized input_cache_read
rate — up to 10× cheaper than a miss — and the split is automatic. The
response reports exactly how many tokens hit:
"usage": {
"prompt_tokens": 51234,
"prompt_tokens_details": { "cached_tokens": 50176 },
"completion_tokens": 312,
"cost": 0.0192
}
To hit the cache: keep your system prompt (schemas, few-shots, RAG corpus)
static and byte-identical at the start of messages, and put variable content
last. At 50k tokens of context, that is $1.50 → $0.15 per question in prompt
costs before the model even starts answering.
2. "What happens when the stream dies mid-flight?" The honest answer is "you pay for delivered tokens only." On runaii.cloud, billing is post-paid against a prepaid wallet: an estimate is reserved before your request runs, settled against actual usage when the stream completes, and the difference refunded automatically. If the request fails upstream or you disconnect mid-stream, the reservation is refunded in full — you are only ever charged for delivered tokens.
3. "Can I verify my own bill?" You should never have to trust the vendor's arithmetic. On runaii.cloud:
usage.costin every response — per-request truthGET /api/v1/key— live balance + in-flight reservations- Console → Usage — per-model spend, latency percentiles, daily charts
- Console → Billing — full ledger (every charge, top-up, refund)
The ledger itself is append-only with a parity check that must return zero drift, always — see the billing engine post for the internals.
How to run this audit on any vendor (including us)
- Fetch their price API, not their marketing page. If there is no price API, that's your answer — prices that can't be fetched can't be checked.
- Run the worked example above at the smallest token counts and confirm
usage.costmatches the rate card exactly. - Send the same request twice with a long identical system prompt and
confirm the cached-token rate from the second response matches the
published
input_cache_readrate. - Kill a stream mid-flight and confirm you are refunded, not charged.
- Reconcile the receipt against the ledger — ask for a ledger export and confirm the sums match your own accounting of every request you sent.
If a vendor fails steps 1–5, their "cheap" rate card is a marketing number, not a price.
Footnotes
- The rate card above is the live output of
GET /api/v1/modelsas of 2026-10-01. Prices can change; the API is the source of truth. gpt-oss-120bcaches at $0.015/1M — a 10× discount on its $0.15 miss rate.qwen3.8-maxandkimi-k3do not support prompt caching (noinput_cache_readrate published — we list the field honestly as absent rather than padding it with a discount that isn't real).- For embeddings,
qwen3-embedding-8bat $0.10/1M tokens is the cheapest published embedding rate on our catalog.
This post is part of the runaii.cloud transparency series — every number published, every bill itemized. If the receipt doesn't match the post, the receipt wins and the post gets corrected.