文档 · 接入
对话补全
POST /v1/chat/completions:发送消息列表,返回模型回复。目录中所有对话模型都使用这个接口。
curl https://<your-endpoint>/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4.1-mini",
"messages": [{"role": "user", "content": "Hello!"}]
}'多轮对话与 system 消息
API 是无状态的:每次请求都要带上完整的对话历史。system 消息用于设定角色和规则,放在最前面。
resp = client.chat.completions.create(
model="gpt-4.1-mini",
messages=[
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is a vector database?"},
{"role": "assistant", "content": "A database optimized for similarity search over embeddings."},
{"role": "user", "content": "Give me two use cases."},
],
temperature=0.3,
max_tokens=800,
)
print(resp.choices[0].message.content)
print(resp.usage)常用参数
| 参数 | 说明 |
|---|---|
model | 必填。目录中的模型 ID,区分大小写 |
messages | 必填。消息数组,role 为 system / user / assistant / tool |
max_tokens / max_completion_tokens | 输出 tokens 上限;推理模型请用 max_completion_tokens |
temperature / top_p | 采样随机性;推理模型通常不支持 |
stop | 停止序列,最多若干个字符串 |
stream | true 时以 SSE 流式返回,见流式输出 |
tools / tool_choice | 工具调用,见工具调用 |
response_format | JSON 输出,见JSON 输出 |
reasoning_effort | 推理深度,见推理参数 |
参数会原样传给模型。模型不支持的参数可能被忽略,也可能返回 400,以具体模型为准。
响应结构
{
"id": "chatcmpl-...",
"object": "chat.completion",
"model": "gpt-4.1-mini",
"choices": [{
"index": 0,
"message": { "role": "assistant", "content": "Hello! How can I help you today?" },
"finish_reason": "stop"
}],
"usage": { "prompt_tokens": 9, "completion_tokens": 10, "total_tokens": 19 }
}| finish_reason | 含义 |
|---|---|
stop | 正常结束 |
length | 达到 max_tokens 上限,输出被截断 |
tool_calls | 模型请求调用工具 |
content_filter | 内容被安全策略拦截 |
usage 给出本次请求的输入、输出 tokens,费用按模型价格计算。