model field. Every paid plan includes the whole
catalog; the Free plan includes the open-model lineup. Prices are USD per one million
tokens — the public list rates your plan’s included usage is measured at. Model names
link to the full page for that model: description, strengths, and capabilities.
Live and preview
The Availability column is the one to read before you pin an id in production.- Live — the gateway holds a routed placement for this exact model. It is what
GET /v1/modelsreturns, and its published context window, capabilities and rates are that model’s own. - Preview — the id is accepted and answered, but no dedicated placement is routed for it yet, so the request is served by a capability-comparable model from the serving pool. Its row below is the card we will bill against and the shape we are working toward; it is not a guarantee that the named model itself is what responds today. A conversation stays on one pick, so behaviour is stable within a session.
This page is a snapshot, and it is the broader of the two lists. Call
GET /v1/models
to read the catalog programmatically — it returns the routed models only, each entry
self-describing down to its capabilities, limits and pricing. The gateway is always
authoritative over this page.Frontier
The largest frontier models, reachable by name through the same endpoint.lx1-fable-5— The narrative flagship — deepest creative and long-form reasoning. Max plan only.lx1-gemini-3-pro— Frontier scale — a 1M-token window and strong multimodal reasoning.lx1-gpt-5.4— The prior GPT-5 flagship — most of 5.5’s strength for less.lx1-gpt-5.5— Frontier generalist — broad, sharp, and steady on long tool chains.lx1-inkling— Thinking Machines’ frontier debut — text, images, and audio in one model.lx1-kimi-k3— Moonshot’s flagship — deep reasoning, vision, and a 1M-token window.lx1-opus-4.7— Prior Opus flagship — frontier reasoning, pinned in place.lx1-opus-5— The frontier ceiling — deepest reasoning and vision for the hardest problems. Pro and Max plans.lx1-qwen3.7-max— The current Qwen flagship — frontier scale with a 1M-token window.
Premium
The ceiling — for the steps where nothing else will do.lx1-grok-4.3— Deep reasoning over a very large window, at the low end of premium.lx1-haiku-4.5— The fast Claude — premium-family quality at a fraction of the latency.lx1-sonnet-4.6— The premium ceiling — top reasoning and vision for the hardest work.lx1-sonnet-5— The newest Sonnet — near-frontier quality for everyday premium work, on every paid plan.
Coding
Heavy, high-stakes coding. The main line of a serious session.lx1-deepseek-v4-pro— Agentic coding flagship — thinks out loud, and holds a million tokens while it does.lx1-glm-5— Fast, reliable coding flagship for everyday heavy work.lx1-glm-5.2— Highest-quality GLM. Flagship tuned for quality over raw speed — higher, variable latency.lx1-kimi-k2.7-code— Coding-specialist flagship. Quality-first; higher, variable latency.lx1-longcat-2— Meituan’s flagship MoE — a 1M-token window with an unusually deep cached-input discount.lx1-qwen3-coder-480b— Heavy coding flagship — large MoE built for complex code.lx1-qwen3-max— The prior Qwen flagship — a heavyweight that codes exceptionally well.
Reasoning
Models that think before they answer — analysis, planning, hard problems.lx1-deepseek-v3.2— Strong reasoning. Best on open-ended analysis, not strict tool loops.lx1-ernie-x1— Baidu’s reasoning specialist — deep deliberation at a low price.lx1-hunyuan-t1— Tencent’s reasoning model — strong long-form thinking, very cheap.lx1-kimi-k2-thinking— Extended reasoning; budget output tokens for its hidden chain-of-thought.lx1-minimax-m2.5— Cheap reasoning with a large context window.lx1-nemotron-3-ultra— The top Nemotron — 550B of deliberate reasoning at a mid-tier price.lx1-qwen3-235b— Big reasoning model at a cheap-tier price — standout value.lx1-qwen3.5-397b— Qwen’s biggest open-weight reasoner — flagship thinking, mid-tier price.lx1-qwen3.8-max— Flagship reasoning with images and a near-million-token window.lx1-step-3— StepFun’s multimodal reasoner — thinks, sees, and calls tools.
General purpose
Capable generalists priced to carry an agent’s daily traffic.lx1-deepseek-v4-flash— Default. The fast half of the V4 line — a million-token window at everyday prices.lx1-ernie-5.1— Baidu’s flagship generalist — broad knowledge with vision.lx1-gemini-3-flash— Google’s fast frontier model — 1M context, multimodal, cheap.lx1-gemma-4-31b— Google’s open workhorse — reasoning and a 256K window near the price floor.lx1-glm-4.7— The GLM workhorse — flagship instincts at an everyday price.lx1-gpt-oss-120b— Strong general-purpose agent model with fast responses. Reachable by name.lx1-hunyuan-hy3— Tencent’s newest generalist — reasoning and tools near the price floor.lx1-hunyuan-turbos— Tencent’s fast generalist — quick answers at a rock-bottom price.lx1-kimi-k2.6— Kimi’s vision generalist — thinks when asked, sees what you show it.lx1-llama-3.3-70b— Meta’s dependable open workhorse — a known quantity everywhere.lx1-minimax-m3— MiniMax’s frontier agent model — 1M context and vision at a workhorse price.lx1-mistral-large-3-675b— Large general-purpose model, fast and capable.lx1-nemotron-3-120b— Hybrid MoE, strong on multi-agent, 256K context.lx1-nemotron-super-3-120b— Strong all-round workhorse with a very large context.lx1-qwen3-next-80b— Efficient workhorse — big context at a low price.lx1-qwen3-vl-235b— Affordable eyes — a 235B vision model at workhorse money.lx1-qwen3.7-plus— Qwen’s balanced mid-tier — a 1M-token window at an everyday price.lx1-qwen3.8-27b— The compact Qwen 3.8 — vision, tools, and reasoning in a 27B dense model.
Everyday coding
Everyday coding hands for routine changes and fast loops.lx1-devstral-2-123b— Coding-specialist tuned for software tasks.lx1-kat-coder-pro— Kuaishou’s coding specialist — direct edits, no reasoning overhead.lx1-kimi-k2.5— Fast coding-general model.lx1-minimax-m2.7— MiniMax’s coding workhorse — agentic edits at a budget rate.lx1-qwen3-coder-30b— Cheap coding offload for routine changes.lx1-qwen3-coder-next— Balanced coding model with a large context.
Long context
Huge context at small-model prices — for the jobs that eat tokens.lx1-gemma-4-26b— 256K context for the price of a small model.lx1-mimo-v2.5— Xiaomi’s efficiency play — 1M context and vision for pocket change.lx1-qwen3.5-flash— A million tokens of context at one flat, tiny price.
Fast
Quick turns, glue steps, dispatch — at all-day-volume prices.lx1-glm-4.7-flash— Near the price floor — GLM-family quality in the budget tier.lx1-gpt-oss-20b— Smallest and fastest tier.lx1-nemotron-nano-3-30b— Cheapest tool + reasoning capable model.lx1-qwen-turbo— The catalog’s cheapest chat tokens — quick, direct answers at volume.lx1-step-3.7-flash— StepFun’s quick multimodal — sees, thinks, and answers fast for very little.
Embeddings
Turn text into vectors — for semantic search, RAG, clustering, and dedupe.lx1-bge-base-en— The long-standing default English embedding.lx1-bge-large-en— The most accurate English embedding in the catalog.lx1-bge-m3— BAAI’s versatile embedding — multilingual, multi-granularity.lx1-bge-small-en— 384 dimensions — the smallest index and the fastest search.lx1-embed-gemma-300m— Compact Gemma-family embedding with a 2K input window.lx1-plamo-embed-1b— Japanese-specialist embedding — the widest vector in the catalog.lx1-qwen3-embed-0.6b— Multilingual retrieval with an 8K input window.