All models
Model card

Gemini 3.8 Flash

lx1-gemini-3.8-flashGeneral purposeStarter planLive in the catalog

Gemini 3.8 Flash combines a million-token context window with vision, reasoning, and tool calling. It is served through Vertex AI and is available by explicit model choice on Starter, Pro, Max, and Scale.

Context window1M tokens
Input · per Mtok$0.75
Output · per Mtok$3.75
Served byLayer X1 engine
Where it earns its keep[01/03]
  • 01Long-horizon coding and agent workflows
  • 02Vision, reasoning, and forced tool calls
  • 031M-token context with up to 65K output
Capabilities
Tool callingYes

Strict, schema-faithful tool calls, enforced by the engine on every request — safe to build an agent loop on.

ReasoningYes

Thinks before it answers. Budget max_tokens generously — hidden reasoning counts against it.

VisionYes

Reads images — screenshots, diagrams, and UI states — inline in the conversation.

Behind the endpoint[02/03]

One endpoint. Served by our engine.

Gemini 3.8 Flash is served through the Layer X1 engine — zero-downtime serving is the design target, not a status-page apology. You request it by name; everything else is our problem.

Call it by name
curl https://api.layerx1.com/v1/messages \
  -H "x-api-key: lx1_your_key" \
  -H "content-type: application/json" \
  -d '{
    "model": "lx1-gemini-3.8-flash",
    "max_tokens": 1024,
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

OpenAI-style clients work too — send the same model name to /v1/chat/completions with a Bearer key. See the docs for both dialects.

Put your agent on inference built for the work

Your agent stays the same.
Its inference gets better.

$export ANTHROPIC_BASE_URL=https://api.layerx1.com

Start free · no card · Launch from $2/mo