Contact Us

The enterprise-grade AI Gateway for security-conscious teams. Protect your data, govern spend, and account for usage.

Read Documentation→

Product

  • Features
  • Security
  • Pricing
  • Docs

Company

  • About Us
  • Blog
  • Playground
  • Contact Us

© 2026 Shim. All rights reserved.

Trust · Care · Precision
SecurityPrivacy PolicyTerms of Service
Contact Us

The enterprise-grade AI Gateway for security-conscious teams. Protect your data, govern spend, and account for usage.

Read Documentation→

Product

  • Features
  • Security
  • Pricing
  • Docs

Company

  • About Us
  • Blog
  • Playground
  • Contact Us

© 2026 Shim. All rights reserved.

Trust · Care · Precision
SecurityPrivacy PolicyTerms of Service
Contact Us
Back to Blog|Home
AI Infrastructure

Designing an LLM Gateway for Kubernetes Workloads

An LLM gateway for Kubernetes needs two layers: a control plane for secrets and budgets, and an inference plane for GPU-aware routing. Here is how to design the split.

July 8, 202611 min read

Kubernetes gives you primitives for routing HTTP traffic: Services, Ingress, Gateway API. None of them understand what an LLM request is. They see opaque bytes, apply round-robin or least-connections, and move on. An LLM gateway operates at a different layer entirely, offering token-aware rate limiting, model-aware routing, prompt and response logging, cost tracking at the token level, semantic caching, and PII detection that generic API gateways simply do not have.

That distinction matters, but it is only half the design problem. Most guides treat an LLM gateway Kubernetes deployment as a single proxy: pick a tool, point your pods at it, done. In practice, clusters running real AI workloads need two distinct layers with different responsibilities. Getting the boundary between them right is the actual architecture decision.

Why a Generic API Gateway Fails LLM Workloads

LLM inference sessions are long-running, resource-intensive, and partially stateful, unlike typical short-lived web requests. A single GPU-backed model server may keep multiple sessions active and maintain in-memory token caches. Traditional load balancers lack the specialized capabilities needed for these workloads and do not account for model identity or request criticality, such as the difference between interactive chat and batch jobs.

Standard Kubernetes Services distribute requests without considering any of this. Round-robin across inference replicas ignores KV cache prefix affinity, leading to wasted expensive GPU resources and suboptimal latency. The request lands on whichever pod is next in line, regardless of whether that pod already has the relevant context loaded in GPU memory.

This is the gap. Not "we need a proxy," but "we need a proxy that understands tokens, models, and GPU state." The architecture that solves it cleanly has two layers.

The Two-Layer Architecture

The Envoy AI Gateway makes this split explicit with a two-tier pattern: a Tier 1 centralized entry point handling authentication, top-level routing, and global rate limiting, and a Tier 2 gateway handling ingress traffic to self-hosted model serving clusters with endpoint picker support. kgateway mirrors this with a sharded deployment pattern where a central ingress gateway applies common traffic management, resiliency, and security rules, then forwards to second-layer gateway proxies that apply app-specific policies.

The two layers serve fundamentally different concerns:

ConcernLayer 1: Control PlaneLayer 2: Inference Plane
Primary jobWho can spend what on which providerWhich pod serves this request most efficiently
Key inputsAPI keys, budgets, routing rulesGPU queue depth, KV cache state, loaded adapters
Changes whenPolicy changes, provider additionsModel deployments, scaling events
Operated byPlatform/security teamML/infrastructure team

Collapsing both into a single component forces one team to own concerns they should not, and makes every policy change a potential inference disruption.

Layer 1: Control Plane (Secrets, Budgets, Policy)

The control-plane gateway sits between your application pods and the outside world. Its job is governance.

Centralized secrets. In the gateway pattern, the gateway is the only service in the cluster that needs access to provider API keys; application pods never touch provider credentials directly. Without this, keys scatter. A typical enterprise with 200 developers discovers 15 to 30 different API keys spread across personal accounts, CI/CD environments, and production deployments, each an untracked connection to an external AI provider. The gateway pattern eliminates that sprawl. For AI agents, this extends further: with agentgateway, the agent holds zero secrets, provider keys live only in the gateway, and the agent cannot call tools not explicitly permitted.

ConfigMap-driven routing. Routing rules live in a ConfigMap, not in code. Which model handles requests tagged as "fast"? Which provider is the fallback? All of this updates without redeploying any application service. Without a gateway, hardcoding provider URLs into services means switching models or providers requires a code change and redeployment.

OpenAI-compatible interface. App pods communicate with the gateway using a standard OpenAI-compatible API format, so existing LLM client libraries work without modification. You change the base URL. That is it.

Token budgets. agentgateway enforces per-route token budget limits using token bucket rate limiting, where each user or API key gets a budget measured in tokens rather than requests. This is where per-namespace budget enforcement becomes practical. One gotcha worth knowing: rate limiting is evaluated before prompt guards, so requests rejected by guardrails still consume the user's token budget quota. Unauthenticated requests, by contrast, do not consume quota because authentication runs first. Design your evaluation chain accordingly.

Layer 2: Inference Plane (Model-Aware Routing)

The inference-plane gateway sits between the control plane and your GPU-backed model servers. Its job is efficiency.

The Gateway API Inference Extension introduces two new CRDs for this layer. InferencePool defines a pool of pods running on shared compute (GPU nodes), managed by the platform admin for deployment, scaling, and balancing. InferenceModel maps a public name (like "gpt-4-chat") to the actual model within an InferencePool, managed by AI/ML owners.

The critical component is the Endpoint Selection Extension. Instead of forwarding to any available pod, it examines live pod metrics, including queue lengths, memory usage, and loaded adapters, to pick the ideal pod for each request. This is what standard Kubernetes Services cannot do.

The performance difference is measurable. In benchmarks, the Extension showed significantly lower p90 latency at 500+ QPS compared to a standard Kubernetes Service, because model-aware routing reduces queueing and resource contention as GPU memory approaches saturation.

llm-d, an open-source distributed inference platform launched by Red Hat, Google, and IBM, pushes this further with disaggregated prefill/decode serving and intelligent routing. Its routing layer delivered up to 3x improvements in time-to-first-token and doubled throughput under SLO constraints in benchmarks. Those gains come specifically from routing decisions that account for GPU state, something no generic load balancer attempts.

What Breaks Without the Split

Skip the two-layer design and three problems compound:

Secret sprawl. Without a centralized gateway, API keys scatter across dozens of Kubernetes Secrets per namespace or deployment. When a key rotates or gets compromised, you hunt through dozens of manifests to update it. There is no single place to revoke access or audit who is calling what.

Provider lock-in. Hardcoded provider URLs mean switching models or providers requires a code change and redeployment. Multiply this across every service calling an LLM, and provider migration becomes a multi-sprint project instead of a ConfigMap update.

No cost visibility. Modern enterprises use 3 to 5 or more LLM providers simultaneously, each with its own pricing model. Without a single point tracking token consumption, there is no unified view of AI spend. Budgets become guesswork. This is the AI FinOps problem at the infrastructure layer.

Tooling Map

Different tools fit different layers. The choice depends on whether you are routing to external APIs, self-hosted models, or both.

Control-plane gateways (Layer 1): LiteLLM Proxy, Portkey, OpenRouter, and Envoy with custom filters all serve this role. They handle auth, routing rules, provider abstraction, and cost tracking. For enterprise-scale LLM deployments, this is where multi-provider management lives.

Inference-plane gateways (Layer 2): kgateway and agentgateway handle model-aware routing to self-hosted inference. The Gateway API Inference Extension with its InferencePool/InferenceModel CRDs slots in here. llm-d combines routing with disaggregated serving for clusters running vLLM at scale.

Overhead is not a concern. The Envoy AI Gateway adds roughly 1 to 3ms of latency per request, negligible against typical LLM response times of hundreds of milliseconds to several seconds.

Which Pattern Fits Your Cluster

Single cluster, external APIs only. You need Layer 1 but not Layer 2. Deploy a control-plane gateway (LiteLLM Proxy, Portkey, or similar) as a Kubernetes Service. App pods point their base URL at it. You get centralized secrets, routing via ConfigMap, and token-level budget enforcement. This covers most teams getting started with LLM workloads.

Multi-team cluster, external APIs. Add the sharded gateway pattern from kgateway: a central ingress gateway with common policies, forwarding to team-specific or app-specific second-layer proxies. This isolates noisy neighbors. Each team's gateway can have its own budget limits and routing rules while sharing the central auth and policy layer.

Self-hosted GPU inference. Now you need both layers. Layer 1 handles the external-facing concerns (secrets, budgets, provider abstraction). Layer 2, using InferencePool/InferenceModel CRDs or llm-d, routes to GPU-backed pods based on live metrics. The boundary between them is clear: Layer 1 decides whether a request should proceed, Layer 2 decides where it lands.

The design mistake is treating this as a single-tool problem. The control plane and inference plane have different operators, different change cadences, and different failure modes. Splitting them is not over-engineering. It is the architecture that lets each layer evolve without breaking the other.

Back to all articlesGet Started Free

The enterprise-grade AI Gateway for security-conscious teams. Protect your data, govern spend, and account for usage.

Read Documentation→

Product

  • Features
  • Security
  • Pricing
  • Docs

Company

  • About Us
  • Blog
  • Playground
  • Contact Us

© 2026 Shim. All rights reserved.

Trust · Care · Precision
SecurityPrivacy PolicyTerms of Service