Contact Us

The enterprise-grade AI Gateway for security-conscious teams. Protect your data, govern spend, and account for usage.

Read Documentation→

Product

  • Features
  • Security
  • Pricing
  • Docs

Company

  • About Us
  • Blog
  • Playground
  • Contact Us

© 2026 Shim. All rights reserved.

Trust · Care · Precision
SecurityPrivacy PolicyTerms of Service
Contact Us

The enterprise-grade AI Gateway for security-conscious teams. Protect your data, govern spend, and account for usage.

Read Documentation→

Product

  • Features
  • Security
  • Pricing
  • Docs

Company

  • About Us
  • Blog
  • Playground
  • Contact Us

© 2026 Shim. All rights reserved.

Trust · Care · Precision
SecurityPrivacy PolicyTerms of Service
Contact Us
Back to Blog|Home
AI Infrastructure

LiteLLM Is Migrating to Rust. Here Is What Changes for Gateway Users

LiteLLM's Rust migration cuts memory 11x and latency 150x. Config, database, and API stay the same. Here is the timeline and what gateway users need to know.

July 7, 20269 min read

On June 22, 2026, LiteLLM announced it is moving its entire AI gateway hot path to Rust, committing to sub-1ms overhead and a sub-100MB binary. The headline benchmarks are dramatic: 15x throughput, 150x lower per-request overhead, 11x less memory.

But the number that matters most is not the speed gain. It is the memory drop. Under concurrent load, the Python proxy peaked at 358.9MB. The Rust target sits at roughly 31.7MB. That gap is the difference between a gateway that stays up and one that gets OOM-killed during a traffic spike.

For gateway users, the short answer: nothing changes on your side. Your config, your database, your API calls all stay exactly where they are. The runtime underneath shifts gradually, route by route, behind parity tests you never have to think about.

What the migration is actually solving

The Python GIL gets blamed for a lot of performance limitations, but LiteLLM's gateway is mostly I/O-bound. The GIL only constrains CPU-bound work on the request path. LiteLLM already scaled by running multiple workers.

The real production risk was memory. Under concurrent load, Python processes climb in memory usage and trigger OOM kills. That is not a latency problem. It is an availability problem. The Rust migration moves request transforms, streaming, and routing into a Rust core outside the GIL, with no first-party Python on the forwarding path in the end state.

The benchmark numbers put this in perspective. At 10 concurrent clients against the same mock upstream, the Rust gateway adds about 0.05ms of overhead per request; the Python path adds about 7.5ms. Throughput rises from 453 to 6,782 requests per second. Memory drops from 358.9MB to 31.7MB.

Zero migration burden for gateway users

This is not a v2 rewrite. LiteLLM's announcement is explicit: config, database schema, and the client API contract stay the same. The runtime under the hot path changes gradually, route by route, behind passing parity and end-to-end tests.

If you self-host LiteLLM, you keep your config.yaml. If you use the Python SDK, you keep your imports. If you call the proxy over HTTP, your request format does not change. The providers you have configured today remain configured tomorrow.

That said, a multi-month migration touching the core request path of a production gateway is not a zero-risk event, regardless of how carefully it is staged. Teams running LiteLLM in production may want to evaluate whether the transition period changes their risk calculus, especially those who have encountered LiteLLM security vulnerabilities in the past. For those exploring the landscape, our LiteLLM alternatives comparison covers the options.

Four stages, one route at a time

The litellm rust migration follows a four-stage architecture, each shifting more of the request path from Python to Rust:

StageArchitectureWhat runs in Rust
Stage 0 (today)Pure Python SDK + FastAPI proxyNothing
Stage 1Python drives Rust transforms via PyO3Data transforms
Stage 2FastAPI thin shell, hot path all RustAuth, rate limiting, transforms, streaming
Stage 3Pure Rust axum server, Python in sidecarEverything on the forwarding path

The Rust core is designed as a pure data-transform layer. It turns your request into a provider request, turns the provider response back, transforms stream chunks, counts tokens, and normalizes errors. It never opens a socket, reads a secret, or writes to your database. The host process handles all I/O until Stage 3.

Each route follows the same cadence: prove one provider first, then expand to all providers, then move the route into the Rust core. OCR goes first as the lowest-risk route. Then /v1/messages adds the streaming axis. Then /chat/completions, the route with the largest parameter surface.

Timeline: what is live and what is coming

As of the June 2026 townhall, LiteLLM has shipped an async-first Mistral OCR bridge and an experimental Axum realtime gateway in the Rust workspace.

The published migration dates:

MilestoneTarget date
litellm.ocr() in RustAugust 15
/messages and /chat/completions in RustSeptember 1
Router (load balancing, fallbacks, retries) in RustSeptember 15
Full server: axum replaces FastAPIDecember 1

December 1 is when the Python FastAPI process is fully replaced by an axum server, with Python relegated to a sidecar for any remaining non-hot-path work.

When the performance gains matter for you

Gateway overhead is usually a small fraction of total model latency. If your typical call takes 2 seconds round-trip to Claude or GPT-4, shaving 7ms off the gateway hop is invisible.

The gains become meaningful for high-throughput, low-latency workloads: classification batches processing thousands of short inputs, embeddings at scale, coding agents making rapid sequential calls. In those patterns, the gateway sits on the critical path and the 150x overhead reduction compounds across every request.

The memory improvement matters everywhere. A gateway that uses 32MB instead of 359MB means smaller instance sizes, more headroom before scaling, and fewer OOM events during traffic bursts. That translates directly to infrastructure cost and uptime.

The benchmark is reproducible

One reasonable concern with migration announcements is whether the numbers are cherry-picked. LiteLLM addressed this by publishing the benchmark harness alongside the results. The setup consists of a mock upstream, a thin Rust gateway, and a load client that times each request in microseconds. The summarized CSV is checked into the repo under benchmark/. The only variable between runs is Python versus Rust.

That level of transparency makes the claims verifiable. Anyone can clone the repo, run the harness, and confirm the numbers on their own hardware.

Stability work running in parallel

A migration of this scope would be risky if it were happening in isolation. LiteLLM's June townhall indicates it is not. The team shipped 94 bug fixes and 24 security fixes in June 2026, alongside a public commitment to zero reported regressions by August 29th.

That August 29 date is two weeks after the OCR route is scheduled to move to Rust, which means the zero-regression commitment covers the first production Rust route. Whether that timeline holds will be the first real signal of how the broader migration will go.

For teams that depend on a gateway staying boring and predictable, the next six months are worth watching. The litellm rust migration promises a meaningfully better runtime without asking users to change anything. The question is whether the transition itself introduces the kind of instability it aims to eliminate.

Back to all articlesGet Started Free

The enterprise-grade AI Gateway for security-conscious teams. Protect your data, govern spend, and account for usage.

Read Documentation→

Product

  • Features
  • Security
  • Pricing
  • Docs

Company

  • About Us
  • Blog
  • Playground
  • Contact Us

© 2026 Shim. All rights reserved.

Trust · Care · Precision
SecurityPrivacy PolicyTerms of Service