Every model.
One key.
Pick a paid plan and keep one key and one endpoint. Launch starts at $2 with $30 of included usage, Starter carries $60, Pro $300, and Max $800. Higher tiers add catalog breadth, premium execution budget, and more room to run.
Real monthly headroom at the smallest possible commitment.
- $30 of model usage included every month
- Open models + GPT-5.6 Luna
- 100 requests / 5 min · 5 concurrent
- In-flight runs never cut off
A serious month of agent work for the price of a coffee.
- $60 of model usage included every month
- Adds GPT-5.6 Terra & the Sonnet class
- 300 requests / min · 10 concurrent
- In-flight runs never cut off
For a daily driver — coding agents that work all day.
- $300 of model usage included every month
- Adds GPT-5.6 Sol & the Opus class
- 1,000 requests / min · 40 concurrent
- In-flight runs never cut off
5x Pro. For agents that never sleep and teams of one that ship like ten.
- $800 of model usage included every month
- Every model — including Astra & Fable
- 3,000 requests / min · 100 concurrent
- In-flight runs never cut off
No subscription. Load credit and spend it whenever you like — at each model's published rate, zero markup. Credit doesn't expire and doesn't reset, and it keeps working after a monthly pool runs out.
Need more than Max — higher limits, custom usage pools, procurement? Talk to us and we'll size a plan to your fleet.
The dashboard and allowance headers show what each request deducts. Model-page prices are the PAYG/list-price reference — see Plans & limits for how the meter works.
- What counts as included usage?
- Every successful request draws from your plan's included-usage allowance. The amount depends on the model, token shape, cache reuse, and plan policy; the dashboard request ledger and allowance headers show the enforced meter. Model-page prices are the PAYG/list-price reference.
- Does repeated context burn my included usage?
- No — on supported models, repeated context is cached automatically and counts at just 10% of the model's input list rate. A long agent session that resends the same system prompt and history every turn draws far less from your pool than its raw token count suggests. See Prompt caching in the docs.
- Which models are included?
- Launch includes open-weight models plus GPT-5.6 Luna. Starter adds GPT-5.6 Terra and the Sonnet class; Pro adds GPT-5.6 Sol and the Opus class; Max adds GPT-6 Astra, GPT-5.5, Fable, and the complete frontier catalog. Every tier requirement is marked on the model catalog.
- What happens when I run out?
- In-flight runs finish — nothing is killed mid-stream. As you approach the edge of your plan, the heaviest models pace to your plan while everything else keeps full speed; at your plan's monthly ceiling, new requests pause until the pool refreshes. If you need it back immediately, upgrade takes effect right away; otherwise it refreshes at the start of your next monthly cycle.
- How do rate limits work?
- Each plan carries a request window and a number of simultaneous in-flight requests. Launch allows 100 requests every 5 minutes; higher tiers use per-minute windows. Past either limit, the API returns 429 with a retry-after header telling your client exactly how long to wait. The full table is on the Plans & limits page in the docs.
- Can I try it before paying?
- Yes — every new account gets $2 in one-time trial credit across the open-model catalog, with both API dialects and streaming. Point one agent at the endpoint and see how it runs before you spend a dollar.
Your agent stays the same.
Its inference gets better.
Start free · no card · Launch from $2/mo