How to Run Kimi K3: API, Local, and Quantized (2026)
Three verified paths to run Kimi K3 — OpenAI-compatible API with per-token pricing, vLLM cluster deploy, or quantized local. Hardware requirements, model ID, 1M context, and a cost table with receipts.
How to Run Kimi K3: API, Local, and Quantized (2026)
Kimi K3 is a 2.8T-parameter MoE with 1M-token context. That size creates an honest split at the start: you will not run the full weights on a workstation — the paths are (1) an OpenAI-compatible API, (2) a multi-GPU cluster with vLLM, or (3) a heavily quantized local build with real accuracy trade-offs. This guide covers all three with working code, the hardware table, the model ID, and the per-token cost of each path.
TL;DR — pick your path
| Path | Needs | Cost | Best for |
|---|---|---|---|
| API (this guide) | Any machine | $3.00 / 1M in, $15.00 / 1M out | Production, agents, 1M-context RAG |
| vLLM cluster | 8×80GB GPUs minimum | Hardware + ops | Full-weight self-hosting |
| Quantized local | 192GB+ RAM (CPU offload) | Accuracy trade-off | Experimentation, privacy |
Kimi K3 hardware requirements (full weights)
| Precision | Weights on disk | Minimum GPUs |
|---|---|---|
| BF16 (full) | ~1.4–1.56 TB | 8×80GB (H100/B200 class) with tensor parallelism |
| 4-bit quantized | ~594 GB | 4×80GB, or CPU-offload with 512GB+ RAM |
| GGUF 1–2 bit | ~150–250 GB | Workstation with 192GB+ RAM + CPU offload |
The 1-bit GGUF build runs at ~78.9% of full accuracy — fine for summarization, risky for code and math. Quantization tables and per-size files are published by the Unsloth team; their docs are the reference for local builds.
Path 1: Run Kimi K3 through an OpenAI-compatible API (10 minutes)
If you already use the OpenAI SDK, this is a two-line change. The model ID is
kimi-k3.
from openai import OpenAI
client = OpenAI(
api_key="sk-runaii-...",
base_url="https://api.runaii.cloud/v1", # ← the only change
)
resp = client.chat.completions.create(
model="kimi-k3",
messages=[{"role": "user", "content": "Summarize this 800k-token corpus: ..."}],
)
print(resp.choices[0].message.content)
print(resp.usage.cost) # exactly what the request charged
Or plain curl:
curl https://api.runaii.cloud/v1/chat/completions \
-H "Authorization: Bearer sk-runaii-..." \
-H "Content-Type: application/json" \
-d '{"model": "kimi-k3", "messages": [{"role": "user", "content": "Hello, fleet."}]}'
Streaming and tool calling work unchanged — same wire format as OpenAI. The 1M-token context is enabled by default; no flags.
What Kimi K3 costs per token
Published live at GET /api/v1/models (USD per 1M tokens):
| Bucket | Rate |
|---|---|
| Prompt (cache miss) | $3.00 |
| Prompt (cache hit) | — (no cache discount on kimi-k3; GLM-5.3-flash has one at $0.03) |
| Completion | $15.00 |
Worked example — a 100k-token document analyzed 10 times (1M prompt tokens, 20k completion tokens):
1.0 × $3.00 = $3.00 (prompt)
0.02 × $15.00 = $0.30 (completion)
= $3.30 total
If the same workload runs on a caching model (GLM-5.3-flash, 1M ctx), the re-sent document bills at the cache-hit rate — up to 10× less on the repeated prefix. Pick per job: Kimi for the hard reasoning step, flash models for the high-volume steps. See the rate card for the full table.
Path 2: Self-host the full weights with vLLM
For a multi-node GPU cluster:
pip install vllm
vllm serve moonshotai/kimi-k3 \
--tensor-parallel-size 8 \
--max-model-len 131072
Then point any OpenAI SDK at http://localhost:8000/v1. Expect the
vLLM recipes repo to carry a tested config for Kimi K3 — tensor-parallel
width and max context are the two knobs that matter at this weight class.
Reality check: at 8×80GB you are renting roughly $16–24/hr of H100/B200 capacity before it serves a single token. The API path is cheaper until your sustained volume passes ~5–10B tokens/month. Run the math for your workload with the savings calculator.
Path 3: Quantized local (the honest version)
1-bit/2-bit GGUF builds via llama.cpp run on big-RAM workstations with CPU offload:
# after downloading a GGUF build (see Unsloth's quantization tables)
llama-server -m kimi-k3-1bit.gguf --n-gpu-layers 0 --ctx-size 8192
Trade-offs, stated plainly: 1-bit ≈ 78.9% accuracy retention; context windows shrink to what your RAM holds; throughput is CPU-bound. Good for local experimentation and privacy-sensitive drafts. Not a production substitute.
Quantization tables and per-size GGUF files are maintained by the Unsloth team — their docs are the community reference for local builds, and the vLLM recipes repo carries tested serving configs for the full-weights path.
FAQ
What is the Kimi K3 model ID on OpenAI-compatible endpoints?
kimi-k3. On runaii.cloud the full ID list is at GET /api/v1/models.
Can you run Kimi K3 locally? The full 2.8T MoE needs an 8×80GB GPU cluster. Quantized GGUF builds run on 192GB+-RAM workstations with an accuracy trade-off (~78.9% at 1-bit).
How much VRAM does Kimi K3 need? ~1.4–1.56 TB at BF16; ~594 GB at 4-bit; ~150–250 GB as 1–2 bit GGUF.
How much does Kimi K3 cost per token?
$3.00 per 1M prompt tokens and $15.00 per 1M completion tokens on
runaii.cloud, billed per request with an itemized receipt (usage.cost).
Is there a free Kimi K3 API? New runaii.cloud accounts get $5 in free credits — enough for ~1.6M prompt tokens on Kimi K3. No card required.
Does Kimi K3 support 1M context?
Yes. On the API path it's enabled by default; self-hosted, set
--max-model-len per your VRAM budget.
Rates verified against the live price API on 2026-10-01. If a receipt disagrees with this post, the receipt wins and the post gets corrected. Next in the series: How to run GLM-5.3-flash · 1M-context models compared · What prompt caching saves.