Claude rate limits span two separate systems: API tiers (RPM/TPM by spend) and Code session caps. Every tier's numbers, 429 handling, and how to stretch throughput.
Claude rate limits are not one system. They are two completely different regimes that happen to share a name, and conflating them is where most developer confusion starts.
The API enforces per-organization quotas measured in requests and tokens per minute, gated by how much you have spent. Claude Code enforces session-level compute caps measured in rolling time windows, gated by your subscription plan. The 429 error you hit on one has nothing to do with the other.
This breakdown covers both, the spend thresholds that unlock each API tier, and the practical techniques that stretch your effective throughput before you need to negotiate an enterprise contract.
The Claude API rate-limits your organization based on three metrics: requests per minute (RPM), input tokens per minute (ITPM), and output tokens per minute (OTPM). Your tier determines the ceiling for each. Limits reset on a rolling basis rather than on a fixed clock, so capacity recovers continuously.
Claude Code is separate. It enforces a 5-hour rolling window plus a weekly cap on active compute hours, and that quota is shared across Claude Code, Claude.ai, and Cowork. Using Claude.ai for a long conversation eats into your Claude Code budget and vice versa.
The distinction matters because the fix for each is different. API limits respond to architectural changes: caching, batching, model tiering. Code limits respond to plan upgrades or workflow changes. Treating one like the other wastes time.
Every API request is metered against three independent counters: RPM, ITPM, and OTPM. Exceeding any one of them returns a 429 error with a retry-after header indicating how long to wait before your next attempt.
Limits are set at the organization level, and the published ceilings vary by model: Haiku's token ceilings run higher than Sonnet's at every tier. Check your organization's actual per-model limits in the console before assuming traffic on one model leaves headroom on another.
There are three usage tiers (Start, Build, Scale) plus a Custom tier for enterprise accounts whose limits are managed directly with Anthropic. You can request a rate limit increase once you are using at least 50% of your current limits.
Anthropic's own docs name the tiers Start, Build, and Scale and publish no spend thresholds for them. The numbered tiers below follow TokenCalculator's April 2026 breakdown of the classic tier ladder: a payment method plus $5 of spend unlocks Tier 1, $500 cumulative spend unlocks Tier 2, $5,000 unlocks Tier 3, and Tier 4 and enterprise limits are negotiated directly.
Here is what that looks like in practice for the two models most production apps use:
| Tier | RPM | TPM | TPD |
|---|---|---|---|
| Free / Build | 5 | 40,000 | 1,000,000 |
| Tier 1 ($5) | 50 | 80,000 | 5,000,000 |
| Tier 2 ($500) | 1,000 | 160,000 | 40,000,000 |
| Tier 3 ($5,000) | 2,000 | 320,000 | Unlimited* |
| Tier 4 / Enterprise | 4,000+ | 800,000+ | Unlimited* |
| Tier | RPM | TPM | TPD |
|---|---|---|---|
| Free / Build | 5 | 50,000 | 5,000,000 |
| Tier 1 ($5) | 50 | 100,000 | 25,000,000 |
| Tier 2 ($500) | 1,000 | 200,000 | Unlimited* |
| Tier 3 ($5,000) | 2,000 | 400,000 | Unlimited* |
| Tier 4 / Enterprise | 4,000+ | 1,000,000+ | Unlimited* |
*Unlimited daily tokens are subject to fair-use policies and can be throttled under peak load.
The Free tier's 5 RPM means one request every 12 seconds on average. Any interactive app with multiple concurrent users will hit that wall immediately. Haiku carries the highest token limits of any model tier, reaching 1,000,000+ TPM at Tier 4, which matters for the model-tiering strategy covered below.
For a full breakdown of what each token costs at these tiers, see Claude API pricing.
Claude Code does not use RPM or TPM. It enforces a dual-layer system: a 5-hour rolling window that caps how many prompts you can send in a burst, plus a weekly cap on total active compute hours.
The limits exist because a small share of power users were running 24-hour sessions and sharing credentials across teams, burning thousands of dollars of compute on $20 subscriptions and degrading service for everyone else. The caps are fair-use enforcement, not a product decision.
What each plan gets you:
| Plan | Price | 5-Hour Window | Weekly Cap | Models |
|---|---|---|---|---|
| Pro | $20/mo | ~10-45 prompts | ~40-80 Sonnet hours | Sonnet 4.6 only |
| Max 5x | $100/mo | ~50-225 prompts | ~140-240 Sonnet + 15-35 Opus hours | Sonnet + Opus |
| Max 20x | $200/mo | ~200-900 prompts | 240-480 Sonnet + 24-40 Opus hours | Sonnet + Opus |
These ranges depend on prompt complexity. A short "rename this variable" prompt consumes far less compute than "refactor this module and write tests." The tilde is doing real work in those numbers.
In March 2026, Anthropic reduced 5-hour limits during weekday peak hours (5-11 AM PT) and acknowledged users were hitting limits faster than expected. That throttling was reversed on May 6, 2026 for Pro and Max accounts, alongside a doubling of the base 5-hour rate limits.
The limits are an infrastructure constraint, not a pricing lever. When Anthropic gets more compute, limits go up.
Anthropic attributes the May 2026 increases to a new agreement with SpaceX, alongside its other recent compute deals. The SpaceX deal gives Anthropic the full capacity of the Colossus 1 data center, over 300 megawatts and 220,000+ NVIDIA GPUs, with the capacity arriving within the month of the announcement.
This pattern is worth understanding. Teams that assume current limits are permanent overbuild retry infrastructure and under-invest in prompt efficiency. Teams that assume limits will vanish get burned when a new model launches and demand spikes. The realistic position: limits will rise, but the ceiling at any given moment is real.
Before upgrading your tier or negotiating an enterprise contract, four techniques can multiply what you get from your current allocation.
Route by model complexity. Sending simple classification, extraction, or summarization tasks to Haiku instead of Opus can multiply effective throughput 5-10x. Haiku's published token ceilings are higher at every tier, and its per-token price is a fraction of Opus's. The pricing difference between Anthropic and OpenAI models makes this even more attractive when you compare across providers.
Cache repeated context. If your system prompt exceeds 1,000 tokens, prompt caching is straightforward. Cached input tokens cost roughly 10% of the normal input rate and do not consume your full TPM the way fresh tokens do. For apps that send the same system prompt on every request, that is a large recurring saving on both spend and rate-limit headroom.
Batch non-urgent work. The Batch API processes requests asynchronously within a 24-hour window at 50% of the standard API rate. Batch requests have separate, higher rate limits that do not count against your synchronous limits. Evaluation runs, bulk classification, nightly content generation: anything that does not need a sub-second response belongs here.
Handle 429s with exponential backoff and jitter. When you do hit a limit, the retry pattern matters. Start at 1 second, double each retry up to 5-6 attempts, and add ±20% random jitter to prevent thundering herd problems where all your clients retry simultaneously and spike the limit again. Read the retry-after header first; if the API tells you how long to wait, use that value instead of guessing.
Twenty developers on Max 20x plans cost $4,000 per month and still hit the 5-hour rolling window during crunch periods. The per-seat subscription model scales linearly with headcount but not with actual compute needs. Some developers barely use their allocation. Others exhaust it by noon.
An infrastructure layer can centralize observability and policy, but it cannot create provider capacity. Handle Claude 429s with explicit client-side backoff and capacity planning; cross-provider fallback needs its own request construction and validation.
Past Tier 3, this trade is worth pricing out. Not because the gateway is cheaper per token, but because it turns a per-seat capacity problem into an infrastructure routing problem, and routing problems have known solutions.
shim preserves Anthropic's native Messages route, status codes, and safe response headers while adding tenant admission limits. It does not pool provider capacity, add semantic caching, or fall back across providers; handle retries and fallback policy in the application.