All models
Model card

Nemotron 3 Ultra

lx1-nemotron-3-ultraReasoningLive in the catalog

NVIDIA's largest Nemotron: a 550B mixture-of-experts (55B active) that reasons on demand and keeps its tool output clean, priced well under what its scale suggests. The natural step up when the Super 120B runs out of depth but the task doesn't justify frontier rates.

Context window202K tokens
Input · per Mtok$0.6
Output · per Mtok$2.4
Served byLayer X1 engine
Where it earns its keep[01/03]
  • 01550B-scale reasoning on demand
  • 02Clean structured tool output
  • 03202K context at a mid-tier price
Capabilities
Tool callingYes

Strict, schema-faithful tool calls, enforced by the engine on every request — safe to build an agent loop on.

ReasoningYes

Thinks before it answers. Budget max_tokens generously — hidden reasoning counts against it.

VisionNo

Text-only.

Behind the endpoint[02/03]

One endpoint. Served by our engine.

Nemotron 3 Ultra is served through the Layer X1 engine — zero-downtime serving is the design target, not a status-page apology. You request it by name; everything else is our problem.

Call it by name
curl https://api.layerx1.com/v1/messages \
  -H "x-api-key: lx1_your_key" \
  -H "content-type: application/json" \
  -d '{
    "model": "lx1-nemotron-3-ultra",
    "max_tokens": 1024,
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

OpenAI-style clients work too — send the same model name to /v1/chat/completions with a Bearer key. See the docs for both dialects.

Put your agent on inference built for the work

Your agent stays the same.
Its inference gets better.

$export ANTHROPIC_BASE_URL=https://api.layerx1.com

Start free · no card · Starter from $5/mo