Skip to main content
Request a model by name with the model field. Every paid plan includes the whole catalog; the Free plan includes the open-model lineup. Prices are USD per one million tokens — the public list rates your plan’s included usage is measured at. Model names link to the full page for that model: description, strengths, and capabilities.

Live and preview

The Availability column is the one to read before you pin an id in production.
  • Live — the gateway holds a routed placement for this exact model. It is what GET /v1/models returns, and its published context window, capabilities and rates are that model’s own.
  • Preview — the id is accepted and answered, but no dedicated placement is routed for it yet, so the request is served by a capability-comparable model from the serving pool. Its row below is the card we will bill against and the shape we are working toward; it is not a guarantee that the named model itself is what responds today. A conversation stays on one pick, so behaviour is stable within a session.
Pin Live ids for anything whose exact behaviour you depend on — evals, structured extraction, a golden-path agent step. Treat Preview ids as a forward-compatible name to develop against, and re-check this page before you depend on one.
This page is a snapshot, and it is the broader of the two lists. Call GET /v1/models to read the catalog programmatically — it returns the routed models only, each entry self-describing down to its capabilities, limits and pricing. The gateway is always authoritative over this page.

Frontier

The largest frontier models, reachable by name through the same endpoint.
  • lx1-fable-5 — The narrative flagship — deepest creative and long-form reasoning. Max plan only.
  • lx1-gemini-3-pro — Frontier scale — a 1M-token window and strong multimodal reasoning.
  • lx1-gpt-5.4 — The prior GPT-5 flagship — most of 5.5’s strength for less.
  • lx1-gpt-5.5 — Frontier generalist — broad, sharp, and steady on long tool chains.
  • lx1-inkling — Thinking Machines’ frontier debut — text, images, and audio in one model.
  • lx1-kimi-k3 — Moonshot’s flagship — deep reasoning, vision, and a 1M-token window.
  • lx1-opus-4.7 — Prior Opus flagship — frontier reasoning, pinned in place.
  • lx1-opus-5 — The frontier ceiling — deepest reasoning and vision for the hardest problems. Pro and Max plans.
  • lx1-qwen3.7-max — The current Qwen flagship — frontier scale with a 1M-token window.

Premium

The ceiling — for the steps where nothing else will do.
  • lx1-grok-4.3 — Deep reasoning over a very large window, at the low end of premium.
  • lx1-haiku-4.5 — The fast Claude — premium-family quality at a fraction of the latency.
  • lx1-sonnet-4.6 — The premium ceiling — top reasoning and vision for the hardest work.
  • lx1-sonnet-5 — The newest Sonnet — near-frontier quality for everyday premium work, on every paid plan.

Coding

Heavy, high-stakes coding. The main line of a serious session.
  • lx1-deepseek-v4-pro — Agentic coding flagship — thinks out loud, and holds a million tokens while it does.
  • lx1-glm-5 — Fast, reliable coding flagship for everyday heavy work.
  • lx1-glm-5.2 — Highest-quality GLM. Flagship tuned for quality over raw speed — higher, variable latency.
  • lx1-kimi-k2.7-code — Coding-specialist flagship. Quality-first; higher, variable latency.
  • lx1-longcat-2 — Meituan’s flagship MoE — a 1M-token window with an unusually deep cached-input discount.
  • lx1-qwen3-coder-480b — Heavy coding flagship — large MoE built for complex code.
  • lx1-qwen3-max — The prior Qwen flagship — a heavyweight that codes exceptionally well.

Reasoning

Models that think before they answer — analysis, planning, hard problems.
  • lx1-deepseek-v3.2 — Strong reasoning. Best on open-ended analysis, not strict tool loops.
  • lx1-ernie-x1 — Baidu’s reasoning specialist — deep deliberation at a low price.
  • lx1-hunyuan-t1 — Tencent’s reasoning model — strong long-form thinking, very cheap.
  • lx1-kimi-k2-thinking — Extended reasoning; budget output tokens for its hidden chain-of-thought.
  • lx1-minimax-m2.5 — Cheap reasoning with a large context window.
  • lx1-nemotron-3-ultra — The top Nemotron — 550B of deliberate reasoning at a mid-tier price.
  • lx1-qwen3-235b — Big reasoning model at a cheap-tier price — standout value.
  • lx1-qwen3.5-397b — Qwen’s biggest open-weight reasoner — flagship thinking, mid-tier price.
  • lx1-qwen3.8-max — Flagship reasoning with images and a near-million-token window.
  • lx1-step-3 — StepFun’s multimodal reasoner — thinks, sees, and calls tools.

General purpose

Capable generalists priced to carry an agent’s daily traffic.
  • lx1-deepseek-v4-flash — Default. The fast half of the V4 line — a million-token window at everyday prices.
  • lx1-ernie-5.1 — Baidu’s flagship generalist — broad knowledge with vision.
  • lx1-gemini-3-flash — Google’s fast frontier model — 1M context, multimodal, cheap.
  • lx1-gemma-4-31b — Google’s open workhorse — reasoning and a 256K window near the price floor.
  • lx1-glm-4.7 — The GLM workhorse — flagship instincts at an everyday price.
  • lx1-gpt-oss-120b — Strong general-purpose agent model with fast responses. Reachable by name.
  • lx1-hunyuan-hy3 — Tencent’s newest generalist — reasoning and tools near the price floor.
  • lx1-hunyuan-turbos — Tencent’s fast generalist — quick answers at a rock-bottom price.
  • lx1-kimi-k2.6 — Kimi’s vision generalist — thinks when asked, sees what you show it.
  • lx1-llama-3.3-70b — Meta’s dependable open workhorse — a known quantity everywhere.
  • lx1-minimax-m3 — MiniMax’s frontier agent model — 1M context and vision at a workhorse price.
  • lx1-mistral-large-3-675b — Large general-purpose model, fast and capable.
  • lx1-nemotron-3-120b — Hybrid MoE, strong on multi-agent, 256K context.
  • lx1-nemotron-super-3-120b — Strong all-round workhorse with a very large context.
  • lx1-qwen3-next-80b — Efficient workhorse — big context at a low price.
  • lx1-qwen3-vl-235b — Affordable eyes — a 235B vision model at workhorse money.
  • lx1-qwen3.7-plus — Qwen’s balanced mid-tier — a 1M-token window at an everyday price.
  • lx1-qwen3.8-27b — The compact Qwen 3.8 — vision, tools, and reasoning in a 27B dense model.

Everyday coding

Everyday coding hands for routine changes and fast loops.
  • lx1-devstral-2-123b — Coding-specialist tuned for software tasks.
  • lx1-kat-coder-pro — Kuaishou’s coding specialist — direct edits, no reasoning overhead.
  • lx1-kimi-k2.5 — Fast coding-general model.
  • lx1-minimax-m2.7 — MiniMax’s coding workhorse — agentic edits at a budget rate.
  • lx1-qwen3-coder-30b — Cheap coding offload for routine changes.
  • lx1-qwen3-coder-next — Balanced coding model with a large context.

Long context

Huge context at small-model prices — for the jobs that eat tokens.
  • lx1-gemma-4-26b — 256K context for the price of a small model.
  • lx1-mimo-v2.5 — Xiaomi’s efficiency play — 1M context and vision for pocket change.
  • lx1-qwen3.5-flash — A million tokens of context at one flat, tiny price.

Fast

Quick turns, glue steps, dispatch — at all-day-volume prices.
  • lx1-glm-4.7-flash — Near the price floor — GLM-family quality in the budget tier.
  • lx1-gpt-oss-20b — Smallest and fastest tier.
  • lx1-nemotron-nano-3-30b — Cheapest tool + reasoning capable model.
  • lx1-qwen-turbo — The catalog’s cheapest chat tokens — quick, direct answers at volume.
  • lx1-step-3.7-flash — StepFun’s quick multimodal — sees, thinks, and answers fast for very little.

Embeddings

Turn text into vectors — for semantic search, RAG, clustering, and dedupe.
  • lx1-bge-base-en — The long-standing default English embedding.
  • lx1-bge-large-en — The most accurate English embedding in the catalog.
  • lx1-bge-m3 — BAAI’s versatile embedding — multilingual, multi-granularity.
  • lx1-bge-small-en — 384 dimensions — the smallest index and the fastest search.
  • lx1-embed-gemma-300m — Compact Gemma-family embedding with a 2K input window.
  • lx1-plamo-embed-1b — Japanese-specialist embedding — the widest vector in the catalog.
  • lx1-qwen3-embed-0.6b — Multilingual retrieval with an 8K input window.