Developer documentation
Shim AI Gateway
Add tenant isolation, PII controls, quotas, spend limits, usage accounting, and audit evidence to native OpenAI, Anthropic, and Gemini API traffic.
Introduction
Shim is an SDK-native governance gateway. It validates each request, resolves the authenticated tenant, applies privacy and admission policy, calls the selected provider, then settles usage and audit records. OpenAI and Anthropic clients can switch by changing the base URL and API key. Gemini clients keep their native routes. Compatibility is scoped to the documented endpoints, not every SDK resource.
Native provider APIs
OpenAI Responses and Chat, Anthropic Messages, and Gemini generateContent stay provider-native.
PII Scrubbing
Redact configured sensitive data before it reaches the selected provider, then restore it in the response.
Durable Governance
Enforce model, quota, and spend policy with tenant-scoped usage and audit records.
Authentication
Authenticate requests using your Shim API Key. You can generate keys in Gateway keys.
Security Best Practice
Authorization: Bearer sk-shim-12345...x-shim-key: sk-shim-12345...Both methods are equivalent. The Anthropic SDK may instead send the Shim key as x-api-key on its native Messages and Models routes. That header is gateway authentication, never Anthropic BYOK.
Quick Start
The Python examples use the official OpenAI and Anthropic SDKs. Gemini requests use its native endpoints and payloads. Shim does not translate, route, or fail over requests between providers.
Set optional OpenAI and Gemini output limits so quota and spend reservation match the response size you intend. Omitting one reserves the platform ceiling, and Anthropic always requires max_tokens.
# Tested with openai==2.53.0
from openai import OpenAI
client = OpenAI(
base_url="https://api.getshim.tech/v1",
api_key="sk-shim-..."
)
response = client.responses.create(
model="gpt-5.6-luna",
input="Hello world",
max_output_tokens=128
)
chat = client.chat.completions.create(
model="gpt-5.6-luna",
messages=[{"role": "user", "content": "Hello world"}]
)
models = client.models.list()Coding agents
The gateway above covers the traffic your applications send. Your developers also run coding agents all day, and those calls never pass through it. Shim CLI covers that side. It runs on the developer machine, with no account, API key or network call to Shim, and it works with Claude Code, Codex and GitHub Copilot CLI.
Install
The package is shim on PyPI. It supports CPython 3.10 to 3.13 on macOS and Linux.
uv tool install --python 3.12 --compile-bytecode shim
# or
pipx install --python python3.12 shimConnect your agent
Preview the change, install the hook, then check it. Replace claude with codex or copilot.
shim install claude --dry-run
shim install claude
shim doctor claudeCodex keeps a trust record for each hook and skips any hook without one. After installing, open /hooks in Codex and enable the shim entry.
What each client does
| Client | Your typed prompt | Files and tool results |
|---|---|---|
| Claude Code | Reported, and blocked under enforce mode | Masked before the model reads them |
| Codex | Reported, and blocked under enforce mode | Not installed yet |
| Copilot CLI | Replaced with the redacted text | Not installed yet |
In the default configuration a secret you type into a Claude Code or Codex prompt is reported, not stopped. Set user-prompt = "enforce" to block it instead.
See a session
shim watch runs Claude Code behind a local proxy for one command and reports the tokens, the spend and every sensitive value that went past. It changes nothing and writes no request body to disk.
shim watch -- claude
printf '%s' 'Mail alice@example.com' | shim redactNative Inference APIs
Each documented route preserves provider-owned request fields and native success JSON or SSE. Compatibility is endpoint-scoped rather than a promise that every SDK resource works, and Shim never silently moves a request to another provider. OpenAI Responses background execution is excluded because Shim does not expose its retrieval lifecycle.
POST /v1/responsesSupports native Responses input, tools, continuations, JSON, and SSE.
POST /v1/chat/completionsPreserves native OpenAI messages, tools, JSON, and SSE.
POST /v1/messagesSupports the Anthropic SDK's regular and beta Messages JSON and SSE semantics.
GET /v1/models[/{model_id}]Returns the native OpenAI or Anthropic catalog shape for the calling SDK.
POST /v1beta/models/{model}:generateContentUse :streamGenerateContent?alt=sse for native Gemini streaming.
Request Governance
On the request pathShim applies controls before an upstream call, then finalizes usage and audit state exactly once.
1. Authenticate
Resolve the tenant from the Shim API key and load its policy.
2. Admit
Check the model catalog, RPM/TPM, quotas, loop protection, and spend policy.
3. Settle
Record actual tokens, cost, lifecycle status, and audit evidence.
Approved models only
400 MODEL_NOT_PRICED before the selected provider is called.Usage, Cost & Budgets
Durable recordsShim records daily requests, model usage, input/output tokens, and settled cost. Budgets can target an organization, team, or request tag with USD or token limits and alert thresholds.
Cost attribution
Assign an API-key cost center or send request tags for filterable activity and budget evaluation.
x-shim-tag: team:research,environment:prodBudget enforcement
Reserve estimated spend before the selected provider, settle actual usage afterward, and surface thresholds through the management API.
PII Protection
SecurityWhen enabled by tenant policy, Shim scans prompts for PII (Personally Identifiable Information) at the gateway before sending them to the selected provider.
Original Prompt
"My email is alice@example.com"
Sent to AI
"My email is <EMAIL_ADDRESS_a6f1c49b2d807e35c84a917f0b6e2d43>"
Randomized placeholders are request-local and repeated values reuse the same placeholder within a request. The value is restored only in the client-facing response. OpenAI Responses continuations additionally use encrypted, tenant-bound, time-limited mappings. Other provider routes preserve their native conversation semantics. While scrubbing is enabled, opaque image, audio, and file inputs are rejected because Shim cannot inspect them safely.
Detected PII Categories
| Category | Details |
|---|---|
| Financial | Luhn-validated credit cards and IBAN identifiers |
| Personal | Email Addresses, Phone Numbers, TCKN (Turkey ID), SSN |
| Secrets & Keys | AWS keys, OpenAI keys, GitHub tokens, and RSA, EC, OpenSSH, or encrypted private keys |
| Infrastructure | IP Addresses, MAC Addresses, Database Connection Strings |
Rate Limits
Tenant policyShim applies configured request-per-minute and token-per-minute burst limits, durable daily/monthly quotas, and exact-repeat loop protection before the selected provider is called.
Burst controls
RPM and TPM are evaluated per authenticated API key.
Durable quotas
Request and token reservations prevent concurrent requests from overspending a period.
Loop protection
Repeated identical requests can be rejected before another billable call.
When a configured limit is exceeded, Shim returns 429 Too Many Requests. Admission errors may also name the rejected dimension.
Streaming
Shim preserves each provider's streaming response format: set stream: true for OpenAI or Anthropic, and call Gemini :streamGenerateContent?alt=sse.
Streaming remains governed
Audit & Compliance Evidence
Shim keeps tenant-scoped request lifecycle and audit records and exposes audit-chain verification, evidence reports, findings, and oversight workflows through the management and compliance control planes.
Evidence
Review request status, model, usage, cost, privacy facts, audit chain, and generated reports without persisting raw prompt bodies in request records.
Oversight
Evaluate oversight policies, route flagged requests to a decision queue, and record human decisions when that control is enabled.
Compliance reports map observed evidence to controls. They are not legal advice, certification, or attestation.
Health Check
Use the health endpoint for monitoring and uptime checks.
GET https://api.getshim.tech/health{
"status": "ok",
"version": "0.1.0",
"database": "connected",
"redis": "connected"
}Dependency failures return a "degraded" status. A database failure also makes the health endpoint return HTTP 503. Redis health is reported separately.
Supported Models Reference
Chat Completions and Anthropic Messages require an approved model. OpenAI Responses may omit model and input to use the provider default, but an explicit model must be approved. GET /v1/models is provider-aware: the OpenAI and Anthropic SDKs receive native responses for their matching enabled catalogs. Gemini model admission remains on its native routes.
| Provider | Catalog | Admission |
|---|---|---|
| OpenAI | Use the OpenAI Models resource at /v1/models. | Unapproved models fail before the provider call. |
| Anthropic | Use the Anthropic Models resource at /v1/models. | Unapproved models fail before the provider call. |
| Configured Gemini models. Use the native Gemini Developer API routes. | Unapproved models fail before the provider call. |
Error Codes
| Code | Meaning | Solution |
|---|---|---|
| 400 MODEL_NOT_PRICED | Model Not Priced | Choose a model listed by the matching OpenAI or Anthropic SDK. |
| 400 INVALID_REQUEST | Invalid Request | Use a valid native output limit and candidate count. |
| 400 PRIVACY_POLICY_BLOCKED | Privacy Policy Blocked | Remove input that the enabled privacy policy cannot safely inspect or redact. |
| 401 | Unauthorized | Use a Shim key through Authorization: Bearer, x-shim-key, or the Anthropic SDK's x-api-key. |
| 413 | Payload Too Large | The request exceeds the configurable body cap: 32,000,000 bytes by default (about 32 MB decimal). |
| 429 | Rate Limit Exceeded | A configured RPM, TPM, quota, or spend limit blocked the request. |
| 502 | Provider Request Failed | The selected provider rejected or could not complete the request. The response body is sanitized. |
| 503 | Service Unavailable | A required service or the selected provider is temporarily unavailable. Honor Retry-After when present. |
| 504 | Provider Timeout | The selected provider timed out. Retry only when your application's policy permits it. |
Provider failures preserve the upstream status and safe request ID but sanitize their bodies into the calling SDK's native envelope: OpenAI {"error": {...}} or Anthropic {"type": "error", "error": {...}, "request_id": "..."}.
Response Headers
Completed inference responses include the Shim lifecycle identifier. Provider responses and sanitized provider errors include a safe upstream identifier when one is available.
| Header | Description |
|---|---|
| X-Shim-Request-Id | Shim lifecycle identifier on completed inference responses. |
| x-request-id | Safe upstream OpenAI request identifier, when available. |
| request-id | Safe upstream Anthropic request identifier, when available. |
| x-goog-request-id | Safe upstream Google request identifier, when available. |
Provider Credentials
Store an OpenAI, Anthropic, or Google key for the tenant in the dashboard, or supply one for a single request. Shim consumes an invocation-scoped key as the provider credential and excludes it from durable request metadata. Authenticate Shim separately with Authorization, x-shim-key, or the Anthropic SDK's gateway x-api-key. Only documented Shim and provider-credential headers are interpreted. Arbitrary provider-specific caller headers are not forwarded upstream.
x-openai-api-key: sk-proj-...x-provider-key: sk-ant-...x-goog-api-key: AIza...x-provider-key is also a generic fallback for OpenAI and Google. The Anthropic SDK's x-api-key contains the Shim key and authenticates the gateway, and it is never BYOK. Every invocation-scoped Anthropic provider key uses x-provider-key.