文档 · 接入
流式输出
设置 stream: true,以服务器推送事件(SSE)边生成边返回。首字更快,也能避免长输出触发客户端超时。
curl https://<your-endpoint>/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4.1-mini",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": true
}'事件格式
每个事件是一行 data: {JSON},增量文本在 choices[0].delta.content;流以 data: [DONE] 结束。
data: {"id":"chatcmpl-...","choices":[{"index":0,"delta":{"role":"assistant","content":""}}]}
data: {"id":"chatcmpl-...","choices":[{"index":0,"delta":{"content":"Hel"}}]}
data: {"id":"chatcmpl-...","choices":[{"index":0,"delta":{"content":"lo!"},"finish_reason":"stop"}]}
data: [DONE]在流式中获取用量
设置 stream_options.include_usage,最后一个 chunk 会携带整次请求的 usage。
stream = client.chat.completions.create(
model="gpt-4.1-mini",
messages=[{"role": "user", "content": "Write a haiku about routers."}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
if chunk.usage: # final chunk: empty choices, carries usage
print("\n", chunk.usage)注意事项
- 如果你在自己的 Nginx 等反向代理后转发流式响应,请关闭响应缓冲(如
proxy_buffering off),否则内容会攒到最后一次性到达。 - 流式请求的读超时按「两次数据之间的间隔」设置,而不是整次请求时长。
- 推理模型在输出第一个字之前可能先思考较长时间,这是正常现象。
- Anthropic 格式的流式事件与 Anthropic 官方一致(
message_start、content_block_delta等),见Anthropic 格式。