Production model routing

One endpoint.
Every model.

Route OpenAI-compatible requests across model providers with one API, predictable budgets, and automatic failover.

OpenAI SDK compatible · Streaming supported · Per-key controls
quickstart.py
from openai import OpenAI

client = OpenAI(
  base_url="https://gateway.example/v1",
  api_key="sk-..."
)

response = client.chat.completions.create(
  model="chat-default",
  messages=[{
    "role": "user",
    "content": "Summarize this incident."
  }],
  max_tokens=300
)

print(response.choices[0].message.content)
OpenAI compatibleDrop-in SDK base URL
StreamingServer-sent events
Key budgetsUsage and rate controls
Automatic routingAliases and fallbacks
A smaller gateway surface

Routing without rebuilding your application.

Keep one client integration while model aliases, provider selection, and access policy evolve behind the endpoint.

01 / ROUTE

Stable model aliases

Point applications at names such as chat-default and change the backing model without shipping new client configuration.

02 / CONTROL

Keys with boundaries

Issue project credentials with model access, request limits, and budgets defined at the gateway.

03 / OBSERVE

Request-level visibility

Review latency, status, token use, and routing outcomes from one consistent operational surface.

Model catalog

Use the names your code already knows.

Provider-specific identifiers and fallbacks stay behind a compact OpenAI-compatible API.

Gateway operational
chat-defaultGeneral chat and extraction
available
gpt-6-astraFrontier reasoning alias
available
gpt-6-solGeneral reasoning alias
available
gpt-6-lunaFast inference alias
available
claude-opus-5-5Compatibility alias
available
claude-sonnet-5-5Compatibility alias
available
gpt-4o-miniCompatibility alias
available
claude-3-5-sonnetCompatibility alias
available
gemini-2.0-flashCompatibility alias
available