Bifrost AI gateway benchmarks show 9.5x throughput over LiteLLM, but raw speed misses the point. Here is where the real security gaps in LLM gateways show up.
Gateway benchmarks measure the wrong thing. A proxy that adds 0.99ms of overhead sounds impressive until you realize the breach came through an ungoverned MCP tool call that the gateway never inspected. The bifrost ai gateway posts strong numbers against LiteLLM on raw throughput, latency, and memory. Those numbers are real. They are also incomplete as a selection criterion, because the security surface that matters most in 2026 is the one most gateways do not govern at all.
Every LLM gateway sits between your application and the model provider. The overhead it adds, measured in milliseconds per request, compounds across every call in a chain. For agentic workflows that make dozens of LLM calls per user interaction, a 40ms tax per hop adds up fast.
Bifrost published benchmarks against LiteLLM on identical AWS EC2 t3.medium instances (2 vCPU, 4 GB RAM) with 500 concurrent virtual users over a 60-second test window. The results:
| Metric | Bifrost | LiteLLM | Difference |
|---|---|---|---|
| Throughput | 424 req/s | 44.84 req/s | 9.5x higher |
| P99 latency | 1.68s | 90.72s | 54x lower |
| Gateway overhead | 0.99ms | 40ms | 40x less |
| Peak memory | 120 MB | 372 MB | 68% less |
| Success rate | 100% | 88.78% | 11.2% gap |
The throughput gap is architectural. Bifrost is built in Go, compiled to native machine code, using goroutines for lightweight concurrency. LiteLLM runs on Python, constrained by the Global Interpreter Lock, asyncio overhead, and higher memory consumption from dynamic typing and garbage collection. At 500 RPS, LiteLLM's success rate drops to 88.78% while Bifrost holds at 100%.
Scale the hardware up and the overhead shrinks further. On a t3.xlarge (4 vCPU, 16 GB RAM), Bifrost's internal overhead drops to 11 microseconds at 5,000 RPS with 100% success.
These are meaningful numbers for capacity planning. They are not, by themselves, a security argument.
An AI gateway handles two distinct traffic planes. The first is the LLM inference plane: prompt in, completion out, with token counting, rate limiting, and content filtering along the way. Most gateway benchmarks, including the ones above, test this plane exclusively.
The second is the MCP tool plane. When an agent calls an external tool through the Model Context Protocol, that call carries its own authentication, its own data exposure, and its own injection surface. Most LLM governance projects start and stop at the model plane, leaving tool calls ungoverned. That gap matters because the two planes carry different risks, and an incident on either plane can affect production.
The numbers back this up. According to IBM's 2025 Cost of a Data Breach Report, 97% of organizations that suffered an AI-related breach lacked proper AI access controls. A Cloud Security Alliance survey found that 82% of organizations discovered an AI agent or workflow in the past year that security or IT did not previously know about.
A gateway that governs only the inference plane leaves the larger attack surface unmonitored.
MCP is young infrastructure with old problems. Authentication was added to the specification only in March 2025 and is still frequently neglected in practice. Researchers found over 1,800 MCP servers exposed on the public internet without authentication.
The scale of unvetted tooling compounds the risk. One research dataset cataloged 13,875 MCP servers and 300 MCP clients crawled from the open web, and very few of those servers were ever vetted by the teams whose machines now run them. This is shadow MCP: employees connecting AI tools to servers without security review, the same pattern that made shadow IT a compliance nightmare a decade ago.
The attack vectors are concrete. Tool poisoning allows an attacker who controls or compromises an MCP server to embed hidden directives in tool metadata (names, descriptions, parameter schemas) that the model reads as instructions. This is indirect prompt injection through a channel most teams are not monitoring. The OWASP Top 10 for LLM Applications 2025 places prompt injection at the top and sensitive information disclosure at number two, with both mitigated primarily through input and output validation rather than prompt engineering alone.
These are not theoretical concerns. Real incidents have shown agents with privileged database access reading and leaking integration tokens after processing attacker-supplied input.
The instinct is to add guardrails in application code, per service. Four failure modes make that approach brittle in practice:
These failure modes are exactly what regulated industries cannot afford under the 2 August 2026 application date for most provisions of the EU AI Act, which requires demonstrable policy enforcement and tamper-evident audit trails for high-risk AI systems. With that deadline less than a month away, gateway-layer enforcement is the fastest path to compliance.
Bifrost addresses the inference and tool planes through a single guardrail layer integrating six providers: Bifrost-native Secrets Detection (Gitleaks-backed), Custom Regex with a PII Detection template, AWS Bedrock Guardrails, Azure AI Content Safety, GraySwan Cygnal, and Patronus AI. All six run behind one configuration interface.
On the MCP plane, Bifrost applies deny-by-default tool filtering by virtual key. Each key can only access the tools explicitly allowed for it. No blanket access, no unvetted servers.
The Go architecture keeps this governance layer invisible in the latency budget. At 0.99ms of overhead on a t3.medium and 11 microseconds on a t3.xlarge, the guardrail processing does not register as a meaningful cost.
Deployment is straightforward. Bifrost is open source under the Apache 2.0 license and operates as a drop-in replacement for OpenAI, Anthropic, LiteLLM, LangChain, and PydanticAI SDKs. You change the base URL. The same SDKs, request formats, and response structures work without modification.
Not every gateway covers both planes equally. Here is where the options diverge on dimensions beyond raw speed.
OpenRouter provides multi-model routing but has no self-hosting or in-VPC deployment option, which is a blocker for regulated industries and air-gapped environments. Its compliance posture depends on the underlying provider routed to, not the gateway itself. For teams subject to the EU AI Act or similar frameworks, that dependency is hard to audit.
LiteLLM offers broad model support and an active community, but the Python architecture limits throughput under load, as the benchmarks above demonstrate. Production deployments require PostgreSQL, Redis, salt-key management, and tuned connection pools. Several enterprise features, including SSO and audit logs, sit behind a commercial license. For a deeper look at these trade-offs, see our LiteLLM alternatives comparison.
| Capability | Bifrost | LiteLLM | OpenRouter |
|---|---|---|---|
| Self-hosted | Yes (Apache 2.0) | Yes (open core) | No |
| Gateway overhead | 0.99ms | 40ms | N/A (hosted) |
| MCP tool governance | Deny-by-default filtering | No native MCP governance | No native MCP governance |
| Guardrail integrations | 6 providers | Via callbacks | Provider-dependent |
| Production dependencies | Single binary | PostgreSQL, Redis, salt-key | N/A (managed) |
The performance gap matters less than the governance gap. A gateway that routes fast but leaves the tool plane ungoverned is optimizing the wrong layer. For teams evaluating AI security posture ahead of the EU AI Act deadline, the question is not how many milliseconds the proxy adds. It is whether the proxy can see, filter, and audit the tool calls that agents make on your behalf.
Benchmark numbers are table stakes. They tell you whether a gateway can keep up with your traffic. They do not tell you whether it can keep your data from leaking through an unvetted MCP server that an intern connected last Tuesday.
Evaluate on three dimensions: inference-plane governance (content filtering, PII redaction, rate limiting), tool-plane governance (MCP authentication, deny-by-default tool access, tool-call auditing), and operational footprint (deployment complexity, licensing, dependency chain). The bifrost ai gateway is the only open-source option that covers all three with sub-millisecond overhead. Whether that combination fits your stack depends on your compliance requirements and how many of those 13,875 cataloged MCP servers your teams have already connected to.