文档 · 生产实践

最佳实践

让生产环境更稳、更快、更省的实践建议。

重试

可重试:429、500、502、503、504、超时和连接错误。不要重试:400、401、403、404——参数或权限问题,重试也不会成功。

使用指数退避加随机抖动,最多 3–5 次;响应带 Retry-After 头时按它等待。OpenAI / Anthropic 官方 SDK 已内置重试,设置 max_retries 即可。

import random, time
import openai

RETRYABLE = (openai.RateLimitError, openai.APITimeoutError, openai.APIConnectionError, openai.InternalServerError)

def chat_with_retry(**kwargs):
    for attempt in range(5):
        try:
            return client.chat.completions.create(**kwargs)
        except RETRYABLE:
            if attempt == 4:
                raise
            time.sleep(min(30, 2 ** attempt) + random.random())   # exponential backoff + jitter
        # 400 / 401 / 403 / 404 are not retried

超时

  • 长输出和推理模型可能需要数分钟,客户端总超时建议不低于 300 秒。
  • 流式请求按「两次数据之间的间隔」设置读超时(例如 60 秒),比限制整次时长更合理。
  • 不要设置很短的超时再反复重试:上游可能仍在生成,既增加负载,也可能产生重复费用。

并发

用信号量或连接池限制同时在途的请求数,逐步加压;批量任务尽量错峰运行,并对 429 做退避。

import asyncio
from openai import AsyncOpenAI

client = AsyncOpenAI(base_url="https://<your-endpoint>/v1", api_key="YOUR_API_KEY", max_retries=3)
sem = asyncio.Semaphore(8)            # max requests in flight

async def one(prompt):
    async with sem:
        r = await client.chat.completions.create(
            model="gpt-4.1-mini", messages=[{"role": "user", "content": prompt}])
        return r.choices[0].message.content

async def main(prompts):
    return await asyncio.gather(*(one(p) for p in prompts))

提示词缓存

OpenAI 等模型会自动缓存重复的长前缀(通常 1024 tokens 以上),命中部分按更低的「缓存读取」价格计费,可在 usage.prompt_tokens_details.cached_tokens 中看到:

"usage": {
  "prompt_tokens": 12840,
  "completion_tokens": 312,
  "prompt_tokens_details": { "cached_tokens": 12288 }
}

Claude 需要用 cache_control 显式标记缓存断点:

msg = client.messages.create(
    model="claude-sonnet-4-5",
    max_tokens=1024,
    system=[{
        "type": "text",
        "text": LONG_STABLE_INSTRUCTIONS,           # long, stable prefix
        "cache_control": {"type": "ephemeral"},     # cache breakpoint
    }],
    messages=[{"role": "user", "content": question}],
)
print(msg.usage.cache_creation_input_tokens, msg.usage.cache_read_input_tokens)
  • 固定内容(系统提示词、工具定义、参考文档)放在最前面,每次变化的内容放在最后。
  • 前缀中不要出现时间戳、随机 ID 等每次都变的内容,任何字节变化都会让缓存失效。

成本控制

  • 按任务选模型:先用小模型,效果不够再升级。
  • 设置合理的输出上限,避免无意义的长输出。
  • 精简上下文:对长对话做截断或摘要,只传需要的文档片段。
  • 推理模型按需调低 effort。
  • 为不同项目使用不同 API Key,在控制台分别查看用量与费用。
  • 上线前在模型详情页估算单次请求成本。

安全

  • API Key 只放在服务端,前端请求通过你自己的后端转发。
  • 记录每次请求的时间、模型、状态码和耗时,排查问题时会非常有用。