Docs · Getting started
Overview
MAX API is an enterprise LLM API platform: one API key and one set of OpenAI- and Anthropic-compatible endpoints give you models from OpenAI, Anthropic, Google, DeepSeek and more.
What you get
- OpenAI-compatible endpoints — chat completions, image generation, embeddings and model listing — that work with the official OpenAI SDKs.
- The Anthropic Messages endpoint (
/v1/messages), so Claude models work with the official Anthropic SDKs. - One API key for every model in the catalog. Switching models means changing the
modelfield. - Multi-channel routing: each model is served by several upstream channels with automatic failover.
- Official list prices, pay as you go, with every request itemized in the console.
- 151
- models
- 11
- vendors
- 3
- model types
Endpoints
| Endpoint | Method & path | Notes |
|---|---|---|
| Chat completions | POST /v1/chat/completions | OpenAI format; streaming, tools, JSON output, vision |
| Anthropic Messages | POST /v1/messages | Native Claude format |
| Image generation | POST /v1/images/generations | Text to image |
| Embeddings | POST /v1/embeddings | Text to vectors |
| List models | GET /v1/models | Models available to your key |
There is no proprietary SDK to learn. Use the official OpenAI / Anthropic SDKs, or any framework or tool that lets you set a custom base URL.
Start here
Getting started
Integrate
- OpenAI SDKIntegrate with the official OpenAI SDKs (Python / Node.js) or plain HTTP. Existing OpenAI code only needs a new base URL and API key.
- Chat completionsPOST /v1/chat/completions sends a list of messages and returns the model's reply. Every chat model in the catalog uses this endpoint.
- StreamingSet stream: true to receive tokens as server-sent events (SSE) while they are generated. You get the first token sooner and avoid client timeouts on long outputs.
- Anthropic formatClaude models can be called in the native Anthropic Messages format at POST /v1/messages, so the official Anthropic SDKs work unchanged.
- Image generationPOST /v1/images/generations creates images from a text prompt. Pick a model of type "Image" from the catalog.
- EmbeddingsPOST /v1/embeddings turns text into vectors for semantic search, RAG, clustering and deduplication. Pick a model of type "Embedding".
- Models & selectionGET /v1/models lists the models available to your key. Prices, context length and capabilities are in the model catalog.
Capabilities
- Tool callingTool calling (function calling) lets the model return structured calls to functions you define. Your code runs them and hands the results back so the model can answer. Works with models that have the Tool calling capability.
- Structured outputGet JSON you can parse directly — for extraction, classification, form filling and more.
- VisionModels with the Vision capability understand images alongside text: receipts, charts, screenshots, photos and more.
- ReasoningReasoning models think before they answer, which helps with math, coding, planning and multi-step analysis. You control how hard they think, trading quality against latency and cost.
Production
- Best practicesPractical advice for running in production: more reliable, faster and cheaper.
- Error codesErrors use standard HTTP status codes. OpenAI-format endpoints return OpenAI-style error bodies; the Anthropic-format endpoint returns Anthropic-style ones.
- Rate limitsTo keep the platform stable for everyone, request rates and concurrency are limited. Default limits suit most production workloads.
- TroubleshootingCommon problems, by symptom.
- FAQCommon questions about integration, billing and capabilities.