A single fallback model is not an LLM routing strategy. Learn how retries, fallback chains, circuit breakers, hedged requests, and smart routing each solve a different failure mode.
Forty-seven incidents across major AI systems in December 2025 alone. Anthropic logged 20 of them, totaling 184.5 hours of impact. OpenAI logged 22, totaling 182.7 hours. That is one month.
A single fallback model is not an LLM routing strategy. It is one layer of five. Retries, fallbacks, circuit breakers, hedged requests, and smart routing each target a different failure mode. Skip any one of them and you have a gap that will surface at the worst possible time.
The argument for multi-provider routing used to be about optimization. Now it is about basic availability.
Research from the AtLarge Research Group at TU Delft found that OpenAI's median time-to-repair for API incidents was 1.23 hours, roughly 1.6 times longer than Anthropic's 0.77 hours. Both numbers exceed what most customer-facing systems can tolerate. And these are median figures. The tails are worse.
Infrastructure failures compound the problem. The October 2025 AWS US-EAST-1 outage demonstrated how a DNS race condition in DynamoDB cascaded across 70+ AWS services for over 15 hours. Any LLM application depending on that region for inference, authentication, or coordination went dark with it.
Even a provider advertising 99.99% uptime still accumulates 52 minutes of downtime per year. Enterprises processing millions of LLM requests daily hedge across two or more providers as a baseline design principle, not as an optimization. No single vendor can be the only path.
Most teams treat LLM failures as binary: the API is up, or it is down. Production reality is more granular. A failover layer must handle several distinct failure modes:
Each mode calls for a different routing response. A retry fixes a transient 5xx. It makes a 429 storm worse. A fallback handles a total outage. It does nothing for latency degradation. Content-filter rejections are not a transient failover case at all and should route through a policy remediation path, not a retry loop.
This is why LLM routing is rarely a single rule and more often a layered policy.
Retries are the first instinct. A request fails, so you send it again. For a genuine transient error (a momentary 502, a dropped connection), this works.
For rate limits, it backfires. Naive immediate retries turn a 429 into a thundering-herd storm that sustains the overload you are trying to escape. Every retry adds to the queue depth that caused the rate limit in the first place.
The fix is exponential backoff with jitter, combined with honoring the provider's Retry-After and x-ratelimit-* headers. Back off longer on each attempt. Add randomness so retries from multiple clients do not synchronize. Respect the provider's explicit signal about when capacity will return.
Retries are the correct response to transient glitches. They are the wrong response to sustained degradation, quota exhaustion, or content policy blocks.
When retries are exhausted, fallbacks take over. A sequential fallback chain defines an ordered list of alternative providers: if Claude fails, try GPT-4; if GPT-4 fails, try Gemini.
Two problems surface at scale. First, every failed attempt adds latency to the user-facing request. If the primary is degraded rather than fully down, the request sits waiting for a timeout before the fallback is tried. A 30-second timeout on the primary plus inference time on the fallback means the user waits over a minute for what should be a two-second response.
Second, fallback switching is not free. The same prompt can behave differently on the fallback model. Output format, reasoning depth, tone, even JSON structure can shift when the request lands on a different model family. Production fallback chains need output validation on the fallback path, not just on the primary.
There is a subtler trap. Fallbacks can share the same failure domain as the primary. If your fallback runs on the same cloud infrastructure, behind the same provider or private endpoint, it may fail in exactly the same way. A fallback chain that routes from OpenAI on AWS to another OpenAI deployment on AWS is not real redundancy during a regional outage.
An LLM proxy that normalizes request formats across providers makes fallback chains practical. Without format normalization, each fallback target requires its own request construction logic.
Retries and fallbacks are reactive. They try to recover after a failure has already happened. Neither prevents a failing provider from being hammered repeatedly with new requests.
Circuit breakers add proactive protection. They monitor failure patterns per provider and automatically cut off traffic to unhealthy components before the rest of the system is affected.
The mechanism follows a three-state machine:
The value is in what does not happen. Without a circuit breaker, a degraded provider that returns errors 60% of the time still receives 100% of its traffic share. Retries make it worse. The breaker fails fast so the degraded provider does not cascade into your own latency and queues.
A chatbot that takes 6 to 8 seconds to respond feels broken, even when the API technically succeeded. For customer support workflows, that delay means longer handle times. For trading or healthcare assistants, decisions arrive too late to be useful.
Parallel hedged requests address latency as a failure mode. The gateway sends the same request to two providers simultaneously and returns whichever response arrives first. This is the pattern Google SRE documented for tail latency reduction. The trade-off is straightforward: hedging doubles the per-request cost during the hedge window. For latency-critical paths (real-time chat, synchronous tool calls), the trade-off is worth it. For batch processing, it is not.
Load balancing across providers matters too, but the naive approach fails. Traditional round-robin is poorly suited for LLM workloads because LLM requests stream over seconds and spike unpredictably. Long-running requests block subsequent ones in the queue, the balancer has no awareness of cache state or prompt context, and the result is severe load imbalance.
Better approaches exist. Consistent hashing with bounded loads (CHWBL) showed a 95% reduction in Time to First Token and a 127% increase in throughput over naive round-robin in benchmarks. The difference comes from routing awareness: CHWBL accounts for which instances are already busy and which have warm caches.
One common misconception about load balancing: extra API keys inside one organization usually share a quota and do not multiply it. Spreading requests across three keys on the same org account does not give you three times the rate limit. Real headroom comes from spreading across independent quota pools, which usually means separate provider accounts or separate providers entirely.
The mechanisms above keep requests flowing when things break. Smart routing asks a different question: which provider should handle this request when everything is working?
In 2026, 37% of enterprises use five or more models in production environments. Not every request needs the most expensive model. A simple classification query does not need GPT-4. A complex reasoning task does not belong on a lightweight model to save a fraction of a cent.
The RouteLLM paper (ICLR 2025, researchers from UC Berkeley, Anyscale, and Canva) quantified what smart routing can achieve: 85% cost reduction while maintaining 95% of GPT-4 performance through trained routing models. Their matrix factorization router achieved 95% of GPT-4's quality with only 26% of calls routed to GPT-4. With augmented training data, that dropped to only 14% GPT-4 calls, 75% cheaper than the random baseline.
This is where LLM routing intersects with AI FinOps. The routing layer becomes a cost optimization surface, not just a reliability mechanism.
Each of these mechanisms (retries, fallbacks, circuit breakers, hedging, smart routing) could theoretically be implemented in application code. In practice, that approach breaks down.
Retries, fallback, load balancing, and circuit breaking belong at the gateway because the gateway sees every provider and key and normalizes the API. A circuit breaker implemented in one microservice has no visibility into the failure rates that other services are experiencing with the same provider. A retry policy in application code cannot coordinate backoff across 15 services hitting the same rate limit. The gateway is the only point in the architecture with full cross-service, cross-provider visibility.
An AI proxy that centralizes these policies means each application team writes a single API call. The routing, failover, and cost optimization happen underneath, managed once and applied everywhere.
Streaming failover illustrates why gateway-level implementation matters. Streaming is the hard case because partial output may already have been sent to the user before a failure is detected mid-stream. Handling this gracefully (buffering, detecting mid-stream failures, replaying from a fallback provider) requires infrastructure that sits between every provider and every consumer. That is a gateway, by definition.
| Mechanism | Failure mode targeted | Trade-off | When to skip |
|---|---|---|---|
| Retries with backoff | Transient 5xx, dropped connections | Adds latency per attempt | 429s without backoff, content policy blocks |
| Fallback chains | Total provider outage, model deprecation | Output variance across models | Single-provider-only contracts |
| Circuit breakers | Sustained degradation, cascading failures | Brief unavailability during cooldown | Low-traffic systems where manual intervention is fast enough |
| Hedged requests | Latency tail, slow responses | Doubled per-request cost | Batch workloads, cost-sensitive pipelines |
| Smart routing | Cost-quality mismatch, model over-provisioning | Router training, classification overhead | Single-model deployments |
No single mechanism covers all five failure modes. Retries without circuit breakers amplify storms. Fallbacks without output validation introduce silent quality regressions. Circuit breakers without fallbacks just drop requests. Hedging without cost controls doubles your bill.
The stack works when all five layers operate together, at the gateway, with shared state across providers and consumers.