Contact Us

The enterprise-grade AI Gateway for security-conscious teams. Protect your data, govern spend, and account for usage.

Read Documentation→

Product

  • Features
  • Security
  • Pricing
  • Docs

Company

  • About Us
  • Blog
  • Playground
  • Contact Us

© 2026 Shim. All rights reserved.

Trust · Care · Precision
SecurityPrivacy PolicyTerms of Service
Contact Us

The enterprise-grade AI Gateway for security-conscious teams. Protect your data, govern spend, and account for usage.

Read Documentation→

Product

  • Features
  • Security
  • Pricing
  • Docs

Company

  • About Us
  • Blog
  • Playground
  • Contact Us

© 2026 Shim. All rights reserved.

Trust · Care · Precision
SecurityPrivacy PolicyTerms of Service
Contact Us
Back to Blog|Home
Cost Optimization

Grok API Pricing Breakdown for Developers

Grok API pricing for every current model, including cached input rates, tool invocation fees, batch discounts, and the hidden cost drivers in agentic workflows.

July 6, 202610 min read

Grok 4.3 charges $1.25 per million input tokens and $2.50 per million output tokens. Competitive with other frontier models on paper. But grok api pricing has layers that the headline numbers do not capture: server-side tool calls billed per invocation, a 2x priority processing multiplier, and context re-injection in multi-turn agents that compounds input costs across every turn.

This breakdown covers each billing surface, from base token rates through storage fees, and identifies where costs actually accumulate in production.

Grok API Pricing at a Glance

ModelContext WindowInput ($/M tokens)Cached Input ($/M)Output ($/M tokens)
grok-4.31M$1.25$0.20$2.50
grok-4.20-multi-agent-03092M$1.25$0.20$2.50
grok-4.20-0309-reasoning1M$1.25$0.20$2.50
grok-4.20-0309-non-reasoning1M$1.25$0.20$2.50
grok-build-0.1256K$1.00$0.20$2.00

All 4.20 variants share identical token pricing. The differentiation is functional, not financial. The multi-agent variant gets a 2M context window for orchestration across long conversations. The reasoning and non-reasoning variants split on whether the model uses chain-of-thought.

The Two Production Models: Grok 4.3 vs. Grok Build 0.1

xAI splits its lineup by task type rather than capability tier.

Grok 4.3 handles general-purpose work. xAI describes it as "the most intelligent and fastest model we've built", intended for everything except code, audio, image, and video. The 1M token context window accommodates large document analysis and extended conversations. Compared to the legacy Grok 4 ($3.00/M input, $15.00/M output), Grok 4.3 is 58% cheaper on input and 83% cheaper on output.

Grok Build 0.1 is the dedicated coding model, trained specifically for agentic coding workflows. At $1.00/M input and $2.00/M output, it runs 20% cheaper than Grok 4.3 on both sides. The trade-off is a smaller 256K context window, which constrains how much codebase context you can inject per request.

The split matters for routing decisions. A pipeline that sends every request to Grok 4.3 overpays for code tasks. One that routes coding requests to Build 0.1 saves 20% on those tokens while using a model purpose-built for the job. For more on this pattern, see our guide to multi-provider LLM routing.

How Cached Input Pricing Works, and Why It Matters for Agents

The $0.20/M cached input rate applies across every model in the current lineup. For Grok 4.3 and the 4.20 variants, that represents an 84% discount versus the standard $1.25/M input rate. For Build 0.1, it is an 80% discount off the $1.00/M standard rate.

Where this hits hardest: agent loops. Any agentic workflow that prepends the same system prompt, tool definitions, or shared context on every API call will see large portions of input tokens served from cache at $0.20/M instead of the full rate. The discount is automatic on repeated prefixes.

But caching only helps with the tokens that repeat. The tokens that change on every turn (new user messages, tool results, growing conversation history) still bill at full price. That distinction becomes critical in multi-turn sessions, covered below.

Server-Side Tool Pricing: The Cost Layer Most Benchmarks Ignore

xAI offers five server-side tools billed per invocation, on top of token costs:

ToolCost per 1,000 Calls
Web Search$5.00
X Search$5.00
Code Execution$5.00
File Attachments$10.00
Collections Search (RAG)$2.50

These charges sit outside the token billing entirely. A single API request that triggers a web search, parses the results, and then calls code execution to process them incurs $0.005 + $0.005 = $0.01 in tool fees before counting a single token.

The compounding factor: the agent autonomously decides how many tools to call, so costs scale with query complexity. A straightforward question might invoke zero tools. A research-heavy query might trigger multiple web searches and an X search in a single turn. You cannot predict per-request tool costs from the prompt alone.

This is why setting budget limits at the gateway layer matters. Without per-request or per-user caps, a spike in complex queries can drive tool costs well beyond what token-only projections would suggest.

Batch API vs. Priority Processing: 20% Off or 2x On

xAI offers two processing modifiers that push costs in opposite directions.

Batch API gives a 20% discount on all token types for grok-4.3 and the grok-4.20 variants. Requests are queued and processed asynchronously, with most completing within 24 hours. The discount applies only to those models; Grok Build 0.1 has no batch discount.

Batch works for offline workloads: content generation pipelines, bulk classification, evaluation runs. Anything where a 24-hour turnaround is acceptable.

Priority Processing goes the other direction, charging a 2x premium over standard rates for higher scheduling priority and lower latency. The multiplier applies to all token types: input, output, cached, and reasoning. A Grok 4.3 request under priority processing costs $2.50/M input and $5.00/M output.

Two restrictions limit where priority applies. It is available only for Chat Completions and Responses endpoints. Image generation, video generation, and Batch API requests cannot use it. So latency-sensitive applications that need priority on text generation will pay double, while their media generation calls stay at standard rates.

Voice and Imagine API Pricing

Beyond text, xAI prices three media categories separately.

Voice API: Realtime agent conversations run $0.05/min ($3.00/hr). Text-to-speech costs $15.00 per million characters. Speech-to-text is $0.10/hr via REST or $0.20/hr for streaming.

Image generation: The quality tier (grok-imagine-image-quality) costs $0.05 per image at 1K resolution and $0.07 at 2K. The standard tier (grok-imagine-image) drops to $0.02 per image at both 1K and 2K.

Video generation: grok-imagine-video-1.5 charges $0.08/sec at 720p and $0.25/sec at 1080p. A 10-second 1080p clip runs $2.50.

These are all separate billing lines from token costs. A multimodal application that combines text, images, and voice will see three distinct cost categories on its invoice.

Storage and File Costs

Files uploaded to the xAI platform incur daily storage fees. Standard file storage costs $0.025/GiB/day. Collection storage (used for RAG) costs $0.10/GiB/day, a 4x premium that reflects the indexing and retrieval infrastructure behind collections.

For a 10 GiB document collection, that is $1.00/day or roughly $30/month in storage alone, before any query costs.

Subscription Plans vs. API Credits

xAI sells two separate products that share the Grok name.

Subscription plans are chat products. Free ($0/month) gives limited access. SuperGrok ($30/month) adds higher rate limits and frontier model access. Business runs $30/user/month. Enterprise offers custom pricing with dedicated infrastructure, SSO/SCIM, data residency, and volume discounts.

API billing is entirely separate and usage-based. Token costs, tool invocations, storage, and media generation all deduct from an API credit balance. A SuperGrok subscription does not include API credits, and API usage does not require a subscription.

Teams sometimes confuse the two. If you are building applications against the Grok API, subscription plan features (chat UI, higher consumer rate limits) are irrelevant to your cost model.

The Hidden Cost Driver: Context Re-Injection in Multi-Turn Agents

Grok's token prices look manageable in isolation. The compounding problem appears in multi-turn agentic workflows.

Every API call in a stateless agent re-sends the full conversation history. A 20-turn session where each turn includes 30K tokens of context accumulates 600K input tokens from re-injected history alone. At Grok 4.3's standard input rate, that is $0.75 in input costs for a single conversation.

Cached input pricing helps with the portion of those 600K tokens that repeats identically (the system prompt, fixed tool definitions). But the growing conversation history changes on every turn, so those tokens bill at full price. The longer the session runs, the larger the uncacheable portion becomes.

This is where grok api pricing diverges most from what a per-token comparison would suggest. Two models with similar per-token rates can produce very different bills depending on how each handles context in agentic loops.

For teams comparing across providers, our Gemini API pricing breakdown covers Google's approach to the same problem.

What This Means for Production Deployments

Grok's billing has at least four independent cost surfaces: token pricing, tool invocations, processing tier multipliers, and storage. Controlling spend requires visibility and policy across all four.

Three patterns reduce Grok costs in practice:

Route by task type. Grok Build 0.1 at $1.00/M input handles coding workloads 20% cheaper than Grok 4.3. Sending every request to the flagship model overpays for specialized tasks. A routing layer that matches request type to model tier captures that spread automatically. See LLM routing patterns for implementation approaches.

Enforce per-request and per-user budget limits. Tool invocations scale with query complexity, and the agent decides how many tools to call. Without caps, a handful of complex queries can spike costs unpredictably. Budget limits at the gateway layer catch these before they compound.

Manage context windows deliberately. The difference between 600K re-injected input tokens and a well-summarized 50K context is an order of magnitude in cost. Prompt caching absorbs the fixed prefix, but only aggressive context management (summarization, sliding windows, selective history) controls the variable portion.

None of these are Grok-specific problems. They are the standard challenges of running any LLM API at scale. But Grok's combination of per-invocation tool billing, the 2x priority multiplier, and large context windows makes them surface faster and hit harder than teams expect from the headline token prices alone.

Plan Grok Traffic With Its Native API

SHIM does not proxy Grok. Use Grok's native API or an integration that explicitly supports it, and keep budget, routing, and failover policy in your application.

Compare with Gemini pricingHow multi-provider routing worksSet up budget limits
Back to all articlesGet Started Free

The enterprise-grade AI Gateway for security-conscious teams. Protect your data, govern spend, and account for usage.

Read Documentation→

Product

  • Features
  • Security
  • Pricing
  • Docs

Company

  • About Us
  • Blog
  • Playground
  • Contact Us

© 2026 Shim. All rights reserved.

Trust · Care · Precision
SecurityPrivacy PolicyTerms of Service