Resources

Rate Limits

Per-key limits, 600 requests per minute and 20 concurrent by default. OpenAI header names. 429 with Retry-After.

Last reviewed 2026-08-25

Defaults

LimitDefaultScope
Requests per minute600, sliding one-minute windowPer key
Concurrent requests20Per key

Raise them through support, or set a lower rpm_limit on a key you hand to a less trusted service.

Headers

Every response carries the OpenAI header names.

x-ratelimit-limit-requests: 600
x-ratelimit-remaining-requests: 584
x-ratelimit-reset-requests: 12s

When limited

429
{ "error": { "message": "Rate limit exceeded. Retry after 12 seconds.", "type": "rate_limit_error", "param": null, "code": "rate_limit_exceeded" } }

The response carries Retry-After in seconds. Wait that long, then retry. Do not retry faster. The OpenAI SDKs do this on their own.

Capacity is not a rate limit

A 503 with model_booting or no_capacity means the model is scaling up, not that your key is throttled. See Capacity errors.

On this page

Share feedback