Inference for AI agents.
An LLM API for coding agents. Run Claude Code cheaper on GLM, MiniMax, Kimi — or Codex, Cline, OpenCode — through one OpenAI-compatible endpoint, built for long-running agent loops. One key, one bill.
- 01Sixteen harnesses · one command
- 02Four protocols · one key
- 03Fails over mid-stream
from openai import OpenAI
client = OpenAI(
base_url="https://api.layerx1.com/v1",
api_key="lx1_your_key",
)
response = client.chat.completions.create(
model="lx1-gpt-oss-120b",
messages=[{"role": "user", "content": "Say hello."}],
)lx1_ key · every modelFull API reference →Point any agent here in one command.
Run npx layerx1 and pick your tools — the installer writes each config where the tool expects it, key and model included; GUI-configured tools get their exact paste-in values. No SDK, no code changes. The session below is the setup.
Three guarantees, drawn to scale.
You point an agent at an endpoint. We do the rest. You never see the machinery — you see what it guarantees. Open one for the full story.
One engine underneath. Every dial in your hands.
You never see the machinery — you see what it guarantees. And you see everything it does for you: usage, spend, every request, caps and keys, live in the dashboard from the first call.
Wired into the agents you already run.
Six of the harnesses people point here every day — open one for the exact config. Each is a base-URL swap and a key: no SDK, no code changes, and the model picker fills from the catalog.
Anything that speaks the OpenAI or Anthropic API speaks to us.
Point your SDK at the endpoint and every model in the catalog is a string away — streaming, tool calls, and structured output included. Free tier on the same key.
Real monthly headroom at the smallest possible commitment.
- $30 of model usage included every month
- Open models + GPT-5.6 Luna
- 100 requests / 5 min · 5 concurrent
- In-flight runs never cut off
A serious month of agent work for the price of a coffee.
- $60 of model usage included every month
- Adds GPT-5.6 Terra & the Sonnet class
- 300 requests / min · 10 concurrent
- In-flight runs never cut off
For a daily driver — coding agents that work all day.
- $300 of model usage included every month
- Adds GPT-5.6 Sol & the Opus class
- 1,000 requests / min · 40 concurrent
- In-flight runs never cut off
5x Pro. For agents that never sleep and teams of one that ship like ten.
- $800 of model usage included every month
- Every model — including Astra & Fable
- 3,000 requests / min · 100 concurrent
- In-flight runs never cut off
No subscription. Load credit and spend it whenever you like — at each model's published rate, zero markup. Credit doesn't expire and doesn't reset, and it keeps working after a monthly pool runs out.
Need more than Max — higher limits, custom usage pools, procurement? Talk to us and we'll size a plan to your fleet.
The dashboard and allowance headers show what each request deducts. Model-page prices are the PAYG/list-price reference — see Plans & limits for how the meter works.
Every agent you run.
One key, one endpoint.
One command wires Claude Code, Codex, Cursor and a dozen more to the same gateway — no per-tool accounts, no per-tool bills.
Start free · no card · Launch from $2/mo