runaiicloud
ModelsAppsGPUsPricingDocsConnectCompareEnterprisePlayground
Log inGet started
← Blog
2026-10-01·runaii engineering·9 min read

How to Run Kimi K3: API, Local, and Quantized (2026)

Three verified paths to run Kimi K3 — OpenAI-compatible API with per-token pricing, vLLM cluster deploy, or quantized local. Hardware requirements, model ID, 1M context, and a cost table with receipts.

tutorialskimiguidespricing

How to Run Kimi K3: API, Local, and Quantized (2026)

Kimi K3 is a 2.8T-parameter MoE with 1M-token context. That size creates an honest split at the start: you will not run the full weights on a workstation — the paths are (1) an OpenAI-compatible API, (2) a multi-GPU cluster with vLLM, or (3) a heavily quantized local build with real accuracy trade-offs. This guide covers all three with working code, the hardware table, the model ID, and the per-token cost of each path.

TL;DR — pick your path

Path Needs Cost Best for
API (this guide) Any machine $3.00 / 1M in, $15.00 / 1M out Production, agents, 1M-context RAG
vLLM cluster 8×80GB GPUs minimum Hardware + ops Full-weight self-hosting
Quantized local 192GB+ RAM (CPU offload) Accuracy trade-off Experimentation, privacy

Kimi K3 hardware requirements (full weights)

Precision Weights on disk Minimum GPUs
BF16 (full) ~1.4–1.56 TB 8×80GB (H100/B200 class) with tensor parallelism
4-bit quantized ~594 GB 4×80GB, or CPU-offload with 512GB+ RAM
GGUF 1–2 bit ~150–250 GB Workstation with 192GB+ RAM + CPU offload

The 1-bit GGUF build runs at ~78.9% of full accuracy — fine for summarization, risky for code and math. Quantization tables and per-size files are published by the Unsloth team; their docs are the reference for local builds.

Path 1: Run Kimi K3 through an OpenAI-compatible API (10 minutes)

If you already use the OpenAI SDK, this is a two-line change. The model ID is kimi-k3.

from openai import OpenAI

client = OpenAI(
    api_key="sk-runaii-...",
    base_url="https://api.runaii.cloud/v1",   # ← the only change
)

resp = client.chat.completions.create(
    model="kimi-k3",
    messages=[{"role": "user", "content": "Summarize this 800k-token corpus: ..."}],
)

print(resp.choices[0].message.content)
print(resp.usage.cost)   # exactly what the request charged

Or plain curl:

curl https://api.runaii.cloud/v1/chat/completions \
  -H "Authorization: Bearer sk-runaii-..." \
  -H "Content-Type: application/json" \
  -d '{"model": "kimi-k3", "messages": [{"role": "user", "content": "Hello, fleet."}]}'

Streaming and tool calling work unchanged — same wire format as OpenAI. The 1M-token context is enabled by default; no flags.

What Kimi K3 costs per token

Published live at GET /api/v1/models (USD per 1M tokens):

Bucket Rate
Prompt (cache miss) $3.00
Prompt (cache hit) — (no cache discount on kimi-k3; GLM-5.3-flash has one at $0.03)
Completion $15.00

Worked example — a 100k-token document analyzed 10 times (1M prompt tokens, 20k completion tokens):

1.0 × $3.00  = $3.00   (prompt)
0.02 × $15.00 = $0.30  (completion)
              = $3.30 total

If the same workload runs on a caching model (GLM-5.3-flash, 1M ctx), the re-sent document bills at the cache-hit rate — up to 10× less on the repeated prefix. Pick per job: Kimi for the hard reasoning step, flash models for the high-volume steps. See the rate card for the full table.

Path 2: Self-host the full weights with vLLM

For a multi-node GPU cluster:

pip install vllm
vllm serve moonshotai/kimi-k3 \
  --tensor-parallel-size 8 \
  --max-model-len 131072

Then point any OpenAI SDK at http://localhost:8000/v1. Expect the vLLM recipes repo to carry a tested config for Kimi K3 — tensor-parallel width and max context are the two knobs that matter at this weight class.

Reality check: at 8×80GB you are renting roughly $16–24/hr of H100/B200 capacity before it serves a single token. The API path is cheaper until your sustained volume passes ~5–10B tokens/month. Run the math for your workload with the savings calculator.

Path 3: Quantized local (the honest version)

1-bit/2-bit GGUF builds via llama.cpp run on big-RAM workstations with CPU offload:

# after downloading a GGUF build (see Unsloth's quantization tables)
llama-server -m kimi-k3-1bit.gguf --n-gpu-layers 0 --ctx-size 8192

Trade-offs, stated plainly: 1-bit ≈ 78.9% accuracy retention; context windows shrink to what your RAM holds; throughput is CPU-bound. Good for local experimentation and privacy-sensitive drafts. Not a production substitute.

Quantization tables and per-size GGUF files are maintained by the Unsloth team — their docs are the community reference for local builds, and the vLLM recipes repo carries tested serving configs for the full-weights path.

FAQ

What is the Kimi K3 model ID on OpenAI-compatible endpoints? kimi-k3. On runaii.cloud the full ID list is at GET /api/v1/models.

Can you run Kimi K3 locally? The full 2.8T MoE needs an 8×80GB GPU cluster. Quantized GGUF builds run on 192GB+-RAM workstations with an accuracy trade-off (~78.9% at 1-bit).

How much VRAM does Kimi K3 need? ~1.4–1.56 TB at BF16; ~594 GB at 4-bit; ~150–250 GB as 1–2 bit GGUF.

How much does Kimi K3 cost per token? $3.00 per 1M prompt tokens and $15.00 per 1M completion tokens on runaii.cloud, billed per request with an itemized receipt (usage.cost).

Is there a free Kimi K3 API? New runaii.cloud accounts get $5 in free credits — enough for ~1.6M prompt tokens on Kimi K3. No card required.

Does Kimi K3 support 1M context? Yes. On the API path it's enabled by default; self-hosted, set --max-model-len per your VRAM budget.


Rates verified against the live price API on 2026-10-01. If a receipt disagrees with this post, the receipt wins and the post gets corrected. Next in the series: How to run GLM-5.3-flash · 1M-context models compared · What prompt caching saves.

Keep reading
OpenAI-compatible inference APIs in 2026: how to compare price per token (with receipts)2026-10-01The anatomy of a 1M-token bill: $0.10, itemized to the token2026-10-01Your prompts are cached — here is what that saves you2026-09-14
runaiicloud

Serverless inference, dedicated GPUs, and training for open models. OpenAI- and Anthropic-compatible APIs.

© 2026 runaii

Platform

Model libraryAppsGPUsPricingCompare providersSavings calculatorDocsServerlessDeploymentsTrainingBatch API

Developers

PlaygroundCookbookCLIAgents / MCPResearch notesUI/UX systemUse casesTutorialsModel advisorBlogCustomersFAQ

Company

EnterpriseStartupsAboutCareersPartnersTrust centerSLAStatusChangelogrunaii chatSupportAPI keysTermsPrivacy