Developer documentation

Shim AI Gateway

Add tenant isolation, PII controls, quotas, spend limits, usage accounting, and audit evidence to native OpenAI, Anthropic, and Gemini API traffic.

Introduction

Shim is an SDK-native governance gateway. It validates each request, resolves the authenticated tenant, applies privacy and admission policy, calls the selected provider, then settles usage and audit records. OpenAI and Anthropic clients can switch by changing the base URL and API key. Gemini clients keep their native routes. Compatibility is scoped to the documented endpoints, not every SDK resource.

Native provider APIs

OpenAI Responses and Chat, Anthropic Messages, and Gemini generateContent stay provider-native.

PII Scrubbing

Redact configured sensitive data before it reaches the selected provider, then restore it in the response.

Durable Governance

Enforce model, quota, and spend policy with tenant-scoped usage and audit records.

Authentication

Authenticate requests using your Shim API Key. You can generate keys in Gateway keys.

Option 1: Bearer Token
Authorization: Bearer sk-shim-12345...
Option 2: Custom Header
x-shim-key: sk-shim-12345...

Both methods are equivalent. The Anthropic SDK may instead send the Shim key as x-api-key on its native Messages and Models routes. That header is gateway authentication, never Anthropic BYOK.

Quick Start

The Python examples use the official OpenAI and Anthropic SDKs. Gemini requests use its native endpoints and payloads. Shim does not translate, route, or fail over requests between providers.

Set optional OpenAI and Gemini output limits so quota and spend reservation match the response size you intend. Omitting one reserves the platform ceiling, and Anthropic always requires max_tokens.

# Tested with openai==2.53.0
from openai import OpenAI

client = OpenAI(
    base_url="https://api.getshim.tech/v1",
    api_key="sk-shim-..."
)

response = client.responses.create(
    model="gpt-5.6-luna",
    input="Hello world",
    max_output_tokens=128
)

chat = client.chat.completions.create(
    model="gpt-5.6-luna",
    messages=[{"role": "user", "content": "Hello world"}]
)

models = client.models.list()

Coding agents

The gateway above covers the traffic your applications send. Your developers also run coding agents all day, and those calls never pass through it. Shim CLI covers that side. It runs on the developer machine, with no account, API key or network call to Shim, and it works with Claude Code, Codex and GitHub Copilot CLI.

Install

The package is shim on PyPI. It supports CPython 3.10 to 3.13 on macOS and Linux.

uv tool install --python 3.12 --compile-bytecode shim
# or
pipx install --python python3.12 shim

Connect your agent

Preview the change, install the hook, then check it. Replace claude with codex or copilot.

shim install claude --dry-run
shim install claude
shim doctor claude

Codex keeps a trust record for each hook and skips any hook without one. After installing, open /hooks in Codex and enable the shim entry.

What each client does

ClientYour typed promptFiles and tool results
Claude CodeReported, and blocked under enforce modeMasked before the model reads them
CodexReported, and blocked under enforce modeNot installed yet
Copilot CLIReplaced with the redacted textNot installed yet

In the default configuration a secret you type into a Claude Code or Codex prompt is reported, not stopped. Set user-prompt = "enforce" to block it instead.

See a session

shim watch runs Claude Code behind a local proxy for one command and reports the tokens, the spend and every sensitive value that went past. It changes nothing and writes no request body to disk.

shim watch -- claude
printf '%s' 'Mail alice@example.com' | shim redact

Native Inference APIs

Each documented route preserves provider-owned request fields and native success JSON or SSE. Compatibility is endpoint-scoped rather than a promise that every SDK resource works, and Shim never silently moves a request to another provider. OpenAI Responses background execution is excluded because Shim does not expose its retrieval lifecycle.

ResponsesPOST /v1/responses

Supports native Responses input, tools, continuations, JSON, and SSE.

Chat CompletionsPOST /v1/chat/completions

Preserves native OpenAI messages, tools, JSON, and SSE.

Anthropic MessagesPOST /v1/messages

Supports the Anthropic SDK's regular and beta Messages JSON and SSE semantics.

ModelsGET /v1/models[/{model_id}]

Returns the native OpenAI or Anthropic catalog shape for the calling SDK.

Gemini Developer APIPOST /v1beta/models/{model}:generateContent

Use :streamGenerateContent?alt=sse for native Gemini streaming.

Request Governance

On the request path

Shim applies controls before an upstream call, then finalizes usage and audit state exactly once.

1. Authenticate

Resolve the tenant from the Shim API key and load its policy.

2. Admit

Check the model catalog, RPM/TPM, quotas, loop protection, and spend policy.

3. Settle

Record actual tokens, cost, lifecycle status, and audit evidence.

Usage, Cost & Budgets

Durable records

Shim records daily requests, model usage, input/output tokens, and settled cost. Budgets can target an organization, team, or request tag with USD or token limits and alert thresholds.

Cost attribution

Assign an API-key cost center or send request tags for filterable activity and budget evaluation.

x-shim-tag: team:research,environment:prod

Budget enforcement

Reserve estimated spend before the selected provider, settle actual usage afterward, and surface thresholds through the management API.

PII Protection

Security

When enabled by tenant policy, Shim scans prompts for PII (Personally Identifiable Information) at the gateway before sending them to the selected provider.

Original Prompt

"My email is alice@example.com"

Sent to AI

"My email is <EMAIL_ADDRESS_a6f1c49b2d807e35c84a917f0b6e2d43>"

Randomized placeholders are request-local and repeated values reuse the same placeholder within a request. The value is restored only in the client-facing response. OpenAI Responses continuations additionally use encrypted, tenant-bound, time-limited mappings. Other provider routes preserve their native conversation semantics. While scrubbing is enabled, opaque image, audio, and file inputs are rejected because Shim cannot inspect them safely.

Detected PII Categories

CategoryDetails
FinancialLuhn-validated credit cards and IBAN identifiers
PersonalEmail Addresses, Phone Numbers, TCKN (Turkey ID), SSN
Secrets & KeysAWS keys, OpenAI keys, GitHub tokens, and RSA, EC, OpenSSH, or encrypted private keys
InfrastructureIP Addresses, MAC Addresses, Database Connection Strings

Rate Limits

Tenant policy

Shim applies configured request-per-minute and token-per-minute burst limits, durable daily/monthly quotas, and exact-repeat loop protection before the selected provider is called.

Burst controls

RPM and TPM are evaluated per authenticated API key.

Durable quotas

Request and token reservations prevent concurrent requests from overspending a period.

Loop protection

Repeated identical requests can be rejected before another billable call.

When a configured limit is exceeded, Shim returns 429 Too Many Requests. Admission errors may also name the rejected dimension.

Streaming

Shim preserves each provider's streaming response format: set stream: true for OpenAI or Anthropic, and call Gemini :streamGenerateContent?alt=sse.

Audit & Compliance Evidence

Shim keeps tenant-scoped request lifecycle and audit records and exposes audit-chain verification, evidence reports, findings, and oversight workflows through the management and compliance control planes.

Evidence

Review request status, model, usage, cost, privacy facts, audit chain, and generated reports without persisting raw prompt bodies in request records.

Oversight

Evaluate oversight policies, route flagged requests to a decision queue, and record human decisions when that control is enabled.

Compliance reports map observed evidence to controls. They are not legal advice, certification, or attestation.

Health Check

Use the health endpoint for monitoring and uptime checks.

GET https://api.getshim.tech/health
{
  "status": "ok",
  "version": "0.1.0",
  "database": "connected",
  "redis": "connected"
}

Dependency failures return a "degraded" status. A database failure also makes the health endpoint return HTTP 503. Redis health is reported separately.

Supported Models Reference

Chat Completions and Anthropic Messages require an approved model. OpenAI Responses may omit model and input to use the provider default, but an explicit model must be approved. GET /v1/models is provider-aware: the OpenAI and Anthropic SDKs receive native responses for their matching enabled catalogs. Gemini model admission remains on its native routes.

ProviderCatalogAdmission
OpenAIUse the OpenAI Models resource at /v1/models.Unapproved models fail before the provider call.
AnthropicUse the Anthropic Models resource at /v1/models.Unapproved models fail before the provider call.
GoogleConfigured Gemini models. Use the native Gemini Developer API routes.Unapproved models fail before the provider call.

Error Codes

CodeMeaningSolution
400 MODEL_NOT_PRICEDModel Not PricedChoose a model listed by the matching OpenAI or Anthropic SDK.
400 INVALID_REQUESTInvalid RequestUse a valid native output limit and candidate count.
400 PRIVACY_POLICY_BLOCKEDPrivacy Policy BlockedRemove input that the enabled privacy policy cannot safely inspect or redact.
401UnauthorizedUse a Shim key through Authorization: Bearer, x-shim-key, or the Anthropic SDK's x-api-key.
413Payload Too LargeThe request exceeds the configurable body cap: 32,000,000 bytes by default (about 32 MB decimal).
429Rate Limit ExceededA configured RPM, TPM, quota, or spend limit blocked the request.
502Provider Request FailedThe selected provider rejected or could not complete the request. The response body is sanitized.
503Service UnavailableA required service or the selected provider is temporarily unavailable. Honor Retry-After when present.
504Provider TimeoutThe selected provider timed out. Retry only when your application's policy permits it.

Provider failures preserve the upstream status and safe request ID but sanitize their bodies into the calling SDK's native envelope: OpenAI {"error": {...}} or Anthropic {"type": "error", "error": {...}, "request_id": "..."}.

Response Headers

Completed inference responses include the Shim lifecycle identifier. Provider responses and sanitized provider errors include a safe upstream identifier when one is available.

HeaderDescription
X-Shim-Request-IdShim lifecycle identifier on completed inference responses.
x-request-idSafe upstream OpenAI request identifier, when available.
request-idSafe upstream Anthropic request identifier, when available.
x-goog-request-idSafe upstream Google request identifier, when available.

Provider Credentials

Store an OpenAI, Anthropic, or Google key for the tenant in the dashboard, or supply one for a single request. Shim consumes an invocation-scoped key as the provider credential and excludes it from durable request metadata. Authenticate Shim separately with Authorization, x-shim-key, or the Anthropic SDK's gateway x-api-key. Only documented Shim and provider-credential headers are interpreted. Arbitrary provider-specific caller headers are not forwarded upstream.

Invocation-scoped OpenAI key
x-openai-api-key: sk-proj-...
Invocation-scoped Anthropic key
x-provider-key: sk-ant-...
Invocation-scoped Google key
x-goog-api-key: AIza...

x-provider-key is also a generic fallback for OpenAI and Google. The Anthropic SDK's x-api-key contains the Shim key and authenticates the gateway, and it is never BYOK. Every invocation-scoped Anthropic provider key uses x-provider-key.