Pricing · $1 free on signup
Serverless per-token inference, per-second dedicated GPUs, and training priced per 1M tokens. Every price on one page — pick a tier to see it live.
Per 1M tokens · Standard tier · OpenAI + Anthropic compatible
| Model | Modality | Input / 1M | Cached / 1M | Output / 1M | Context |
|---|---|---|---|---|---|
| GLM-5.3NEW | LLM | $1.40 | $0.260 | $4.40 | 1M |
| GLM 5.3 FlashNEW | Vision | $0.15 | $0.030 | $0.50 | 1M |
| Kimi K3NEW | Vision | $3.00 | — | $15.00 | 1M |
| DeepSeek V4 Pro | LLM | $1.32 | $0.044 | $3.96 | 1M |
| DeepSeek V4 Flash | LLM | $0.22 | $0.022 | $0.66 | 1M |
| Qwen3.8-Max | LLM | $2.00 | — | $6.00 | 262K |
| Qwen3.8 Flash | Vision | $0.40 | — | $1.60 | 262K |
| OpenAI gpt-oss-120b | LLM | $0.15 | — | $0.60 | 131K |
| OpenAI gpt-oss-20b | LLM | $0.05 | — | $0.20 | 131K |
| MiniMax M3 | LLM | $0.30 | — | $1.20 | 512K |
| Llama 3.3 70B Instruct | LLM | $0.23 | — | $0.85 | 131K |
| Qwen3 Embedding 8B | Embedding | $0.10 | — | — | 41K |
| Qwen3 Reranker 8B | Rerank | $0.20 | — | — | 41K |
| Whisper V3 Large | Audio | $0.25 | — | — | — |
| FLUX.1 Kontext Pro | Image | $0.04 | — | — | — |
Batch API: 50% off all serverless rates · 24h turnaround · same gateway.
Paste one key into your app — billed per token from the table above.
Per-second billing · autoscale + scale-to-zero · your weights, your VPC.
Full Precision
Popular$96/hr
NVIDIA B300 288GB ×8
FP8 · Max quality for frontier dense models
Deploy this presetThroughput
Popular$80/hr
NVIDIA B200 180GB ×8
NVFP4 · Best tokens/$ for large MoE models
Deploy this presetMinimal
Popular$48/hr
NVIDIA B300 288GB ×4
NVFP4 · Smallest frontier-capable footprint
Deploy this presetManaged SFT/DPO priced per 1M training tokens (dataset tokens × epochs). RL jobs bill per GPU-second at on-demand rates. Checkpoints deploy to inference at base-model prices.
| Model size | LoRA SFT | LoRA DPO | Full SFT | Full DPO |
|---|---|---|---|---|
| Up to 16B params e.g. gpt-oss-20b | $0.50 | $1.00 | $1.00 | $2.00 |
| 16–80B params e.g. llama-3.3-70b | $3.00 | $6.00 | $6.00 | $12.00 |
| 80–300B params e.g. qwen3.8-235b-class | $6.00 | $12.00 | $12.00 | $24.00 |
| 300B+ params e.g. deepseek-v4, kimi-k3 | $10.00 | $20.00 | $20.00 | $40.00 |
Per 1M training tokens · VLM fine-tuning billed the same way · open training console →
Early-stage teams get $500 in credits, locked beta pricing, and a direct line to engineering.
Startup program →[ credits ]
One balance across everything we meter — serverless tokens, per-second GPU deployments, and training runs. The ledger splits every call into a line item, so you always know exactly what it cost.
$1 free on signup · no card · balance never expires
Serverless tokens
15+ models, one key. Cached prefixes up to 97% off.
Dedicated GPU-seconds
B200/B300/GB300, autoscale, scale-to-zero. Same balance.
Training runs
SFT/DPO per 1M tokens, RL per GPU-second. Ships to prod in seconds.
Never expires
Sit on your balance for a year — it stays yours.
Usage is always metered separately. Plans buy limits, support, and guarantees.
$0 + usage
For developers getting a first agent into production.
$250/mo + usage
For startups scaling inference with guardrails.
Custom
For regulated scale: residency, VPC, and contracts.
No — your balance never expires. Top up once and spend it across serverless tokens, GPU deployments, and training runs whenever you're ready.
About 700K input tokens on GLM-5.3, or ~6.6M on GLM Flash. Enough to prototype a full agent loop before paying.
Repeat prompt prefixes (system prompts, RAG context) are billed at the cached rate — up to 97% off on DeepSeek V4 Pro. Cache hits are automatic.
Serverless for variable traffic and prototyping. Dedicated (B200/B300/GB300, per-second) when you need guaranteed throughput, custom weights, or VPC isolation.
Yes — async batch inference is 50% off serverless rates, same as the Fireworks model we benchmarked against.
Global routing is cheapest. US-only is the same price; EU/APAC add 15% for data-residency guarantees.
Estimate what moving your monthly inference spend to an open-model mix saves you. Rows above are the real prices.
$900
closed APIs
$25
runaii mix
$875
you keep / mo
Estimate using a Flash-heavy open mix; exact numbers on this page. Cached prefixes + batch cut it further. Start with $1 free →