Codex Rate Limit Exceeded: What It Means and How to Fix It

In this article

When OpenAI Codex returns a rate limit error, it means you've exceeded the number of requests or tokens allowed within a given time window. The fix depends on which limit you hit: requests-per-minute (RPM), tokens-per-minute (TPM), or daily task quotas enforced at the account level. Free-tier users hit these walls much faster than paid users.

  • Free Codex access is limited to a small number of tasks per day (OpenAI hasn't published an exact number, but reports from users suggest single-digit daily tasks)
  • API-based Codex usage is governed by your OpenAI API tier: Tier 1 starts at 500 RPM and 200,000 TPM for most models
  • Limits reset on a rolling per-minute basis (RPM/TPM) or at midnight UTC (daily quotas)

What does the "Codex rate limit exceeded" error mean?

The error means your account has sent too many requests, consumed too many tokens, or exhausted a daily task quota in the current window. OpenAI enforces rate limits at multiple levels simultaneously: per-minute request counts, per-minute token consumption, and (for the Codex CLI agent) per-day task limits tied to your subscription. Any one of these can trigger the error independently.

There are three distinct surfaces where you can hit this:

  • Codex CLI (the agentic tool): Task-based limits on free and Plus accounts. OpenAI has not published hard numbers, but users on the free tier report being cut off after a handful of tasks per day.
  • OpenAI API (model access): If you're calling Codex models via the API, limits are tiered by your OpenAI usage tier, starting at Tier 1 with 500 RPM.
  • ChatGPT / web UI: If Codex is surfaced inside ChatGPT, it inherits ChatGPT's message limits per plan.

What are the Codex rate limits by plan?

For the Codex CLI agent (the agentic coding tool launched by OpenAI in 2025), access is currently research-preview gated. Free and paid users get different task allocations, though OpenAI has not committed to specific public numbers. For the underlying API, limits scale with your account's spending history across five tiers.

TierRPM (requests/min)TPM (tokens/min)Requirement
FreeVery low (model-dependent)Very lowAccount created
Tier 1500200,000$5 spend
Tier 25,0002,000,000$50 spend + 7 days
Tier 35,0004,000,000$100 spend + 7 days
Tier 410,00010,000,000$250 spend + 14 days

Source: OpenAI rate limits documentation. These tiers apply to API access. The Codex CLI product has separate usage policies not fully documented at the time of writing.

Why does Codex hit rate limits so fast?

Codex is an agentic tool. Unlike a single-turn chat request, each task it runs can spawn dozens of internal API calls as it reads files, edits code, runs tests, and iterates. A single "add feature X" prompt might consume 50,000 tokens or more behind the scenes. This burns through per-minute and daily limits far faster than a typical API integration.

Other contributing factors:

  • Large codebases passed as context inflate token counts significantly
  • Multi-step tasks (write + test + fix) chain multiple model calls
  • Parallel task execution (if enabled) compounds the RPM hit
  • Agentic loops that self-correct multiply token usage unpredictably

How do you fix the Codex rate limit exceeded error?

The right fix depends on which limit you've hit. For RPM/TPM limits on the API, you wait and retry with exponential backoff. For daily CLI task quotas, you wait until midnight UTC or upgrade your plan. For persistent problems, the solutions below address the root cause rather than just the symptom.

1. Wait for the window to reset

Per-minute limits (RPM and TPM) reset every 60 seconds. If you hit one of these, a short pause is enough. Check the Retry-After header in the API response — it tells you exactly how many seconds to wait. See when Codex usage limits reset for a full breakdown of reset timing by limit type.

2. Upgrade your API tier

The fastest path to higher limits is spending more on the OpenAI API. Each tier requires a minimum spending history. Adding $5 to your account gets you to Tier 1 immediately; $50 and 7 days unlocks Tier 2. Check your current tier in the OpenAI usage dashboard.

3. Reduce context size per request

Don't pass your entire repository into every Codex task. Use .codexignore (similar to .gitignore) to exclude irrelevant files. Keeping context tight reduces TPM consumption and lets you run more tasks within the same window.

4. Batch or queue your tasks

If you're running multiple tasks programmatically via the API, implement a queue with rate-aware backoff rather than fire-and-forget parallel requests. OpenAI provides a rate limit mitigation guide with example retry logic in Python and Node.

5. Switch to a different model temporarily

If you're hitting limits on a specific model (e.g., codex-1), check whether a related model has headroom. Rate limits are tracked per-model, so temporarily routing to an alternative can unblock you while you wait for the primary quota to reset.

What are the alternatives when Codex is rate-limited?

If you're blocked waiting for Codex limits to reset, other agentic coding tools can pick up the slack. Each has its own rate limit model. Claude Code, for instance, uses a 5-hour rolling window with a separate weekly cap, which can be easier to predict and plan around than per-minute API quotas.

How do you monitor Codex usage before hitting limits?

For API usage, OpenAI exposes usage data in the platform usage dashboard, and rate limit headers (x-ratelimit-limit-requests, x-ratelimit-remaining-requests, x-ratelimit-reset-requests) are included in every API response. Parsing these headers in your integration lets you back off proactively before an error occurs.

For the Codex CLI product, there is no built-in real-time usage indicator at the time of writing. The tool surfaces an error only after you've exhausted your quota, not before. If you also use Claude Code and want proactive limit tracking there, Usagebar sits in your macOS menu bar and shows your remaining 5-hour window, weekly cap, and context usage with notifications at 50%, 75%, and 90%.

Also see: OpenAI Codex pricing breakdown and is Codex free?

Key takeaways

  1. Codex rate limit errors come from three sources: per-minute API limits, daily CLI task quotas, or account-tier restrictions.
  2. Per-minute limits reset within 60 seconds; check the Retry-After header. Daily quotas reset at midnight UTC.
  3. Agentic tasks burn limits faster than single-turn requests. Expect 10-50x the token usage of a plain API call.
  4. Adding $5 to your OpenAI account unlocks Tier 1 API access with 500 RPM and 200K TPM.
  5. Reduce context size and implement retry logic to stay within limits without upgrading.
  6. When blocked, tools like Claude Code, Kiro, or Hermes can fill the gap while Codex quotas reset.

Sources

Never Get Locked Out Mid-Task Again

Never hit your usage limits unexpectedly. Usagebar lives in your menu bar and shows your 5-hour and weekly limits at a glance.

Get Usagebar

$9 — one-time, lifetime updates