[ docs — api reference ]
An OpenAI-compatible, per-token-metered inference API on our own GPU fleet. Swap your base_url, keep your SDK — pay per token with automatic prompt-cache discounts.
base_url
new here?
Make your first metered API call in 60 seconds.