Docs · Integrate

Anthropic format

Claude models can be called in the native Anthropic Messages format at POST /v1/messages, so the official Anthropic SDKs work unchanged.

The Anthropic SDK base URL does not include /v1 (the SDK appends /v1/messages). The OpenAI SDK base URL does include /v1.

Install

pip install --upgrade anthropic

Basic request

curl https://<your-endpoint>/v1/messages \
  -H "x-api-key: $API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Streaming

import anthropic

client = anthropic.Anthropic(base_url="https://<your-endpoint>", api_key="YOUR_API_KEY")

with client.messages.stream(
    model="claude-sonnet-4-5",
    max_tokens=4096,
    messages=[{"role": "user", "content": "Explain TCP slow start."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
    final = stream.get_final_message()
print("\n", final.usage)

CLI and IDE tools

Tools that accept a custom Anthropic endpoint usually only need these two variables:

# Tools that accept a custom Anthropic endpoint (CLIs, IDE plugins…) usually read these
export ANTHROPIC_BASE_URL="https://<your-endpoint>"
export ANTHROPIC_API_KEY="YOUR_API_KEY"

Differences from the OpenAI format

OpenAI formatAnthropic format
Path/v1/chat/completions/v1/messages
AuthAuthorization: Bearerx-api-key + anthropic-version
System promptrole: system messageTop-level system field
Output capOptionalmax_tokens required
Replychoices[0].message.contentcontent[] blocks
Usageprompt_tokens / completion_tokensinput_tokens / output_tokens / cache_*

Claude models can also be called in the OpenAI format. For Claude-specific features such as thinking or cache_control breakpoints, use the native format; parameters follow Anthropic's official documentation.