xAI Grok API Review: Models, Features, Examples, and Pricing in 2026
Information current as of September 29, 2026.
Grok started as an assistant inside the X ecosystem, so for a long time its API was easy to see as a rather specialized alternative to OpenAI. That is no longer the case in 2026. xAI is building a full developer platform: the Grok API provides reasoning models, images, files, web search and X search, Python code execution, RAG over your own documents, MCP, image and video generation, real-time voice, speech recognition, and speech synthesis.
There is another practical difference from many competitors: the xAI REST API is compatible with the OpenAI format. This significantly lowers the cost of a first experiment for an existing project—in many cases, all you need to change is base_url, the API key, and the model ID. At the same time, xAI itself recommends the Responses API for new integrations, while the old Chat Completions interface is now considered legacy.
This article looks at current Grok models, the Responses API, reasoning, web/X search, file handling, Grok Imagine, Voice API, pricing, prompt caching, privacy, and connecting to Grok from PHP.
What is the Grok API?
The Grok API is xAI's programmatic interface for accessing Grok models and related services. The main inference endpoint is at https://api.x.ai, and authentication uses a regular Bearer API key.
The API currently provides several broad areas:
- text and multimodal Grok models;
- Responses API and legacy Chat Completions;
- function calling and Structured Outputs;
- Web Search and X Search;
- Code Execution;
- Files and Collections for working with documents;
- Remote MCP;
- Grok Imagine for images and video;
- Voice API for speech-to-speech, TTS, and STT;
- Batch API and deferred requests;
- prompt caching and context compaction.
An API key is created in the xAI Console. The most common billing option is prepaid: a team buys credits in advance, which are then deducted as the API is used. Monthly invoiced billing is available to some organizations.
Responses API: the primary interface
Older xAI examples often use /v1/chat/completions. For a new integration, it is better to start with:
POST /v1/responses
xAI describes the Responses API as its primary interface for text generation, reasoning, and tool use. Chat Completions remains compatible and functional, but new capabilities appear in the Responses API first.
A minimal request looks like this:
curl https://api.x.ai/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-d '{
"model": "grok-4.7",
"input": "Briefly explain the difference between REST and GraphQL."
}'
The Responses API can store conversation state on xAI's side. The first response receives an ID, and the next request can be linked to it with previous_response_id without manually resending the entire history.
Stored Responses are available for up to 30 days. If an application wants to control the history entirely on its own, it can store it locally and pass it with the required reasoning data in later requests.
Grok 4.7 has one additional feature: the Responses API returns encrypted reasoning content. You can pass it back without decrypting it in the next turn so the model can continue with the previous reasoning state.
OpenAI-compatible API
xAI officially describes its REST API as compatible with the OpenAI REST API. This applies in particular to the Responses API and Chat Completions.
In Python, you can switch an existing OpenAI client roughly like this:
from openai import OpenAI
client = OpenAI(
api_key="XAI_API_KEY",
base_url="https://api.x.ai/v1",
)
response = client.responses.create(
model="grok-4.7",
input="Explain how dependency injection works.",
)
print(response.output_text)
This is one of the most convenient features of the Grok API for a quick A/B test. But full interchangeability is not guaranteed: xAI has its own server-side tools, reasoning behavior, X Search, Collections, and separate media and voice endpoints.
So a simple text generator often migrates with almost no changes, while a complex agentic project still needs a separate integration layer.
Current Grok text models
xAI has a rather unusual model numbering scheme. For example, Grok 4.20 came out before Grok 4.7, so the number should not be read like an ordinary decimal version where 4.20 must be “newer” than 4.7. As of late September 2026, xAI itself recommends Grok 4.7 as its most capable general model for code, agentic tasks, and knowledge work.
For most projects, these options are worth considering:
| Model | Context | Short-context input / cached / output | Long-context input / cached / output | Notes |
|---|---|---|---|---|
| Grok 4.7 | 500K | $2 / $0.50 / $6 | $4 / $1 / $12 | Current frontier model |
| Grok 4.3 | 1M | $1.25 / $0.20 / $2.50 | $2.50 / $0.40 / $5 | Less expensive, 1M context |
| Grok 4.20 Reasoning | 1M | $1.25 / $0.20 / $2.50 | $2.50 / $0.40 / $5 | Reasoning + tool calling |
| Grok 4.20 Non-Reasoning | 1M | $1.25 / $0.20 / $2.50 | $2.50 / $0.40 / $5 | Fast tasks without reasoning |
| Grok 4.20 Multi-Agent | 1M | $1.25 / $0.20 / $2.50 | $2.50 / $0.40 / $5 | Parallel work by multiple agents |
For all the models listed, long-context pricing starts at 200K prompt tokens. Once the prompt reaches that threshold, the higher rate applies to the entire request, not just the tokens after 200K.
Grok 4.5 and 4.6 are also available, but for a new project without special requirements it makes more sense to test 4.7 and the less expensive 4.3/4.20 first.
Grok 4.7
Grok 4.7 was released on September 21, 2026 and is now positioned by xAI as a frontier model for coding, agentic tasks, and knowledge work.
The model accepts text and images and returns text. Its context window is 500K tokens; xAI's documentation currently does not specify a fixed text output limit. Its knowledge cutoff is May 2026.
Grok 4.7 supports function calling, Structured Outputs, and reasoning at these levels:
low;medium;high;xhigh.
The default is high. You cannot turn reasoning off completely in 4.7.
For prompts under 200K tokens, pricing is $2 per million input tokens, $0.50 per million cached input tokens, and $6 per million output tokens. With a long context, the rates double to $4/$1/$12.
Grok 4.7 supports the Responses API, Chat Completions, web search, X search, and code execution. xAI currently recommends it as the main choice for development and agentic tasks.
Grok 4.3 and Grok 4.20
If Grok 4.7 is the current choice for maximum capability, Grok 4.3 and the 4.20 family are interesting for their price and 1M-token context windows.
Grok 4.3 has a 1M-token context window and costs $1.25/$2.50 per million input/output tokens for short context. Cached input costs $0.20. For prompts of 200K tokens or more, the price rises to $2.50/$5.
Grok 4.3 lets you adjust reasoning, including using it without reasoning. This makes the model useful when complex reasoning is not needed for every request.
Grok 4.20 has at least three important variants:
- reasoning;
- non-reasoning;
- multi-agent.
The reasoning and non-reasoning versions have the same base price and 1M-token context window. The choice depends on whether the model needs to spend compute on internal analysis or a fast, direct answer matters more.
Grok 4.20 Multi-Agent
Grok 4.20 Multi-Agent, currently in beta, is especially interesting. Instead of one agent loop, the model launches several agents to research different parts of a task in parallel, after which a leader agent combines the results.
It is designed mainly for deep research and complex multi-step tasks. Multi-Agent can be used with Web Search, X Search, Code Execution, and Collections Search.
The reasoning-effort parameter has an unusual meaning here: it controls not only reasoning depth, but also the number of active agents. xAI shows scenarios with 4 and 16 agents. The 16-agent option consumes significantly more tokens and makes sense only for genuinely complex research.
The full internal discussion does not have to appear in the outward response: the API returns tool calls and the leader agent's final answer.
Reasoning and encrypted reasoning
You cannot turn reasoning off completely in Grok 4.7. Developers adjust its intensity through reasoning.effort.
For example:
{
"model": "grok-4.7",
"reasoning": {
"effort": "medium"
},
"input": "Analyze this system's architecture and find potential points of failure."
}
The higher the reasoning effort, the more compute the model may spend on a task. This affects latency and cost, because reasoning tokens count toward usage.
The Responses API returns reasoning in encrypted form. The application does not need to, and should not try to, decrypt the internal reasoning; the block is passed back in the next request as part of the previous state.
This matters for long agent loops: the model can continue multi-step work without losing the internal context from the previous step.
Function calling
Like OpenAI, Grok can choose and call application functions. The developer describes a function and its JSON Schema; the model returns a structured tool call, and the backend performs the actual action.
For example:
{
"type": "function",
"name": "get_order",
"description": "Get order details",
"parameters": {
"type": "object",
"properties": {
"order_id": {
"type": "string"
}
},
"required": ["order_id"]
}
}
A user might write: “Check order 54821 and tell me when it was shipped.” Grok decides to call get_order; the application checks the user and queries the database, after which the result can be returned to the model to prepare a regular answer.
Function calling should not replace business logic. Even if the model selects the right function, the backend must separately check authorization, input parameters, and whether the operation can be performed.
Structured Outputs
For backend scenarios, xAI supports Structured Outputs. With response_format, you can require either a regular JSON object or JSON that strictly conforms to a specified JSON Schema.
This lets you use Grok for:
- extracting fields from documents;
- classifying requests;
- creating API objects;
- parsing incoming emails;
- preparing CRM data;
- passing a result to the next step in an agentic workflow.
When supported schema elements are used, xAI guarantees that the response conforms to the specified structure. This is more reliable than a regular instruction to “return only JSON.”
Web Search
The built-in web_search gives Grok access to the current web. The model can run searches and open pages on its own, then use the results when forming an answer.
Example:
{
"model": "grok-4.7",
"input": "Find the latest changes in PHP and briefly list the main ones.",
"tools": [
{
"type": "web_search"
}
]
}
Web Search costs $5 per 1,000 calls, plus the model's regular token charges. The same tool can enable image search, after which the images found become part of the model's context.
Keep this in mind when estimating the cost of an agentic scenario: one user request can generate several searches, so the price depends on more than the tokens in the final answer.
X Search: Grok's key distinctive advantage
xAI has a built-in tool that other major LLM APIs do not offer in the same way: X Search. It lets Grok search posts, users, and threads directly on X.
For example:
{
"model": "grok-4.7",
"input": "What are developers on X saying about the latest Laravel release?",
"tools": [
{
"type": "x_search"
}
]
}
For monitoring public reaction, trends, breaking news, product discussions, or a specific event, this can be a serious advantage of the Grok API. The model gets information from the current X feed rather than relying on its own knowledge cutoff.
X Search pricing changed on September 21, 2026. Instead of a simple charge per call, it now counts retrieved items: $5 per 1,000 fetched posts and $10 per 1,000 fetched user profiles. If the same post is found by two searches, it may be counted twice.
The Responses API returns detailed usage counters, including the number of posts and profiles found, so you can track the actual search cost for each request.
Code Execution
Built-in Code Execution lets Grok run Python in a sandbox. The tool is useful for calculations, statistics, data analysis, and cases where the model needs to run a program rather than just “estimate” the result.
It costs $5 per 1,000 calls, plus model tokens.
The tool can be combined with search and files. For example, Grok can find data online, load information from an attached CSV, run a Python analysis, and then prepare a final report.
Remote MCP
The Responses API supports Remote MCP tools. This lets you connect compatible MCP servers and give the model access to external systems through a standardized protocol.
xAI currently charges no separate fee for calling a Remote MCP tool; you pay for the tokens the model uses while working with it. The external MCP service may, of course, have its own cost or limits.
MCP is especially useful when you do not want to define dozens of custom functions manually for every external system. Still, access rights and critical operations should be restricted on the server side rather than relying only on model instructions.
Working with files
The Files API lets you upload documents and attach them to Grok messages. The maximum size of a single file is 512 MB.
Supported formats include PDF, Markdown, TXT, CSV, JSON, source code, and other text formats. You can attach several documents at once and ask questions about them.
An attachment automatically turns a regular request into an agentic workflow. xAI connects the internal attachment_search, and Grok searches for relevant parts of the document instead of simply sending the entire file into the context.
This search costs $10 per 1,000 tool calls, plus regular token charges.
Files can also be used with Code Execution—for example, Grok can find the required data in a CSV, write Python, and calculate a result.
Collections: built-in RAG
Collections are designed for a persistent knowledge base. Unlike regular file attachments, a Collection is persistent document storage with an embedding index and semantic search.
You can upload and group documents in a Collection, configure chunking and embeddings, and add metadata. Search can be limited with filters such as author, category, or another attribute.
The Responses API provides a collections_search / file_search tool. It costs $2.50 per 1,000 calls, plus tokens for the retrieved context and the model's response.
For a small RAG project, this makes it possible to avoid a separate vector database. For larger systems with their own retrieval and ranking requirements, external infrastructure can still provide more control.
Prompt caching
xAI automatically caches matching prefixes in consecutive requests. If the beginning of a conversation matches, reused tokens can be served from the cache at a substantially lower price.
For Grok 4.7, regular short-context input costs $2 per million tokens, while cached input costs $0.50. The difference is even greater for Grok 4.3 and 4.20: $1.25 versus $0.20.
A cache hit is not guaranteed. Requests may reach different servers, or a cache entry may be evicted.
xAI recommends using:
prompt_cache_keyin the Responses API;- the
x-grok-conv-idheader in Chat Completions.
This helps route requests in one conversation to the same server and significantly improves the chance of a cache hit.
Context Compaction
Caching alone is not enough for a long agent loop: history keeps growing as tool outputs and old turns accumulate.
xAI offers Context Compaction. The API condenses a long history into a compact opaque item that preserves important state—system instructions, files, reasoning, and key results from previous steps.
In the next request, this compacted context is passed instead of the entire old conversation. This reduces input cost and latency, and lowers the risk of the model getting distracted by details that are no longer relevant.
For long-running autonomous processes, this feature may matter more than the nominal maximum context window.
Batch API
Batch API is available for bulk background processing. It lets you queue many requests, and results are usually ready within 24 hours.
Unlike OpenAI or Anthropic, where batch is often associated with a fixed 50% discount, xAI's discount depends on the model.
As of late September 2026, Grok 4.3 and the Grok 4.20 family receive a 20% discount on tokens through Batch API. Grok 4.7 does not yet support Batch API.
Batch supports not only text chat, but also some image/video operations and server-side tools. This makes it useful for bulk classification, document analysis, content generation, and other tasks without real-time requirements.
Deferred Chat Completions
There is a separate deferred mode for legacy Chat Completions. The application starts a request, immediately receives a request_id, and retrieves the result later.
The completed response can be retrieved once within 24 hours, after which it is deleted.
For a new architecture, it is better to use the Responses API and Batch where possible. Deferred Chat Completions is mainly useful for existing systems already built around /v1/chat/completions.
Priority Processing
For latency-sensitive tasks, xAI offers Priority Processing. You can add this to the request body:
{
"service_tier": "priority"
}
If priority capacity is available, the request receives higher priority. The response includes the service_tier actually applied, so the application can log whether the request was really handled as priority.
Priority makes sense for a user-facing endpoint where time to first token affects UX. For background processing, the regular tier or Batch is more economical.
Prompt caching and Priority can be combined: cached-token discounts are applied first, followed by the corresponding Priority pricing.
Exact Cost Tracking
xAI has a useful production feature: every inference response includes the exact cost of that request in the usage.cost_in_usd_ticks field.
One dollar is equal to 10 billion ticks:
cost_usd = cost_in_usd_ticks / 10_000_000_000
The field reflects the actual billable cost after caching and server-side tool calls. It is available not only for regular text, but also for streaming, image generation, and video generation.
This lets you calculate unit economics directly for a user action. For example, an application can determine the actual cost of one research task, one document processing job, or one agentic workflow without doing its own approximate token math.
What Grok costs in practice
Consider a regular request with 100K input tokens and 10K output tokens, with no tools or cache.
| Model | Cost |
|---|---|
| Grok 4.7 | $0.26 |
| Grok 4.3 | $0.15 |
| Grok 4.20 | $0.15 |
Now consider 250K input tokens and 10K output tokens. Long-context pricing applies to the entire request:
| Model | Cost |
|---|---|
| Grok 4.7 | $1.12 |
| Grok 4.3 | $0.675 |
| Grok 4.20 | $0.675 |
The difference shows why you should not use the maximum context window just because it is available. Once the 200K-token threshold is crossed, the rate increases for the entire prompt.
In a real agentic workflow, Search, Code Execution, file search, and other tool calls are added to the token cost. So the exact cost_in_usd_ticks field is often more useful than a theoretical calculation from the price list.
Grok Imagine: image generation
xAI uses the separate Grok Imagine family for images.
The current grok-imagine-image-2.0 supports text-to-image and image editing, including up to five reference images. The price depends on resolution and quality:
| Mode | Price per image |
|---|---|
| 1K Low | $0.04 |
| 1.5K Low | $0.05 |
| 2K Low | $0.06 |
| 1K Medium | $0.06 |
| 1.5K Medium | $0.07 |
| 2K Medium | $0.08 |
Using a source image as input costs an additional $0.01 per image.
There is also a less expensive grok-imagine-image, which generates 1K or 2K images for $0.02. For a new project, however, consider xAI's migration path: the old grok-imagine-image-quality has already been announced for retirement on November 2, 2026, and will be redirected to grok-imagine-image-2.0.
Video generation and editing
Grok Imagine supports video as well as images.
grok-imagine-video-1.5 accepts text, images, and audio references, and generates video up to 1080p:
- 480p — $0.08 per second;
- 720p — $0.14 per second;
- 1080p — $0.25 per second.
The previous grok-imagine-video is less expensive and supports video generation and editing up to 720p: $0.05/s for 480p and $0.07/s for 720p.
Imagine API supports text-to-video, image-to-video, reference-to-video, video editing, and video extension. xAI can therefore already be considered a single platform for media generation as well as LLMs.
Voice API
xAI has a full voice offering with three parts.
Speech-to-Speech works in real time through /v1/realtime. The current model costs $0.08 per minute of audio and supports tool use, so you can build conversational assistants without a separate STT → LLM → TTS chain.
Speech-to-Text is available in batch and streaming modes. REST transcription costs $0.10 per hour of audio; streaming costs $0.20 per hour. The model supports 25 languages.
Text-to-Speech costs $15 per million characters and supports expressive multilingual voices and telephony codecs.
This is already a full alternative to a separate set of speech recognition, LLM, and speech synthesis services for a voice application.
Using images as input
The main Grok 4.7, 4.3, and 4.20 models can analyze images. The maximum size for one image is 20 MiB, and JPEG and PNG are supported.
This is separate from Grok Imagine. A regular text model receives an image as part of its context and responds with text; Imagine handles image generation and editing.
Web Search and X Search can also pass images they find to the model for visual understanding. X search also supports analysis of discovered videos.
Model aliases and production stability
xAI uses several types of model ID.
A name such as grok-4.3 is usually an alias for the current stable version. The -latest suffix also follows a new version, while a model ID with a specific date pins a particular release.
For a regular application, an alias is convenient: model improvements arrive without code changes. For a workflow where full reproducibility matters—such as automated extraction or a regulated process—it is better to pin a specific version and update it after your own evaluations.
This is especially relevant for xAI, which releases and retires models fairly quickly. In May 2026, for example, Grok 3, early Grok 4, and grok-code-fast-1 were retired, and old slugs began redirecting to Grok 4.3.
Grok 4.7 Fast is not a public API model
You may come across the name Grok 4.7 Fast, but there is an important limitation. It is the same model on faster infrastructure, but is available only in Cursor and Grok Build.
You currently cannot use grok-4.7-fast in the public xAI API. Regular Grok 4.7 is available at the standard $2/$6 price, while Fast is billed separately at a higher price in the applicable products.
For an API integration, do not build the architecture around the Fast variant unless xAI makes it publicly available as a model.
Using the Grok API in PHP
For PHP, the regular REST API is the simplest option. Thanks to the OpenAI-compatible format, the integration does not require a special client library.
Example using the Responses API:
<?php
$apiKey = $_ENV['XAI_API_KEY'];
$payload = [
'model' => 'grok-4.7',
'input' => 'Briefly explain what dependency injection is.',
];
$ch = curl_init('https://api.x.ai/v1/responses');
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_RETURNTRANSFER => true,
CURLOPT_HTTPHEADER => [
'Authorization: Bearer ' . $apiKey,
'Content-Type: application/json',
],
CURLOPT_POSTFIELDS => json_encode(
$payload,
JSON_UNESCAPED_UNICODE
),
]);
$response = curl_exec($ch);
if ($response === false) {
throw new RuntimeException(curl_error($ch));
}
$data = json_decode(
$response,
true,
flags: JSON_THROW_ON_ERROR
);
foreach ($data['output'] ?? [] as $item) {
if (($item['type'] ?? null) !== 'message') {
continue;
}
foreach ($item['content'] ?? [] as $content) {
if (($content['type'] ?? null) === 'output_text') {
echo $content['text'];
}
}
}
In Laravel, it is more convenient to move xAI requests into a separate service class and use the built-in HTTP Client. The API key should be stored in .env and retrieved through config/services.php, so application code does not depend directly on environment variables.
Never pass the key to the frontend. Requests should go through a backend that controls the user, rate limits, available tools, and spending limits.
API keys, billing, and teams
Working with the API starts in the xAI Console: create an account/team, buy credits, and generate an API key.
The most common option is prepaid credits. The organization tops up the balance in advance and sees usage in Usage Explorer. Companies that need postpaid invoicing can request monthly invoiced billing separately.
Administrative operations are in a separate Management API and use a separate Management API key. This lets you keep an application's inference key separate from keys that can manage Collections, team settings, billing, and other administrative functions.
Rate limits
Rate limits depend on the model and usage tier. xAI measures at least requests per second and tokens per minute.
For example, Grok 4.7's initial published limit is substantially higher than Multi-Agent's: the frontier model is designed for high regular inference traffic, while multi-agent research launches several agents and is heavier by definition.
As usage grows, a team moves to higher tiers; xAI states that an achieved tier is not later downgraded.
For bulk processing that should not compete with real-time traffic for regular rate limits, it is better to use Batch API.
Regional endpoint in the United States
By default, https://api.x.ai is a global endpoint, and xAI may route requests across regions.
To ensure that primary inference is handled in the United States, use:
https://us.api.x.ai/v1
As of late September 2026, the US endpoint supports Grok 4.7 and 4.6. Token usage costs 10% more than on the global endpoint.
The guarantee applies to API request handling, model inference, moderation, and retained request data. Files, Collections, server-side tools, and the client's network route are not included in this guarantee.
So the US endpoint is useful as an infrastructure option, but should not automatically be treated as a complete solution for every data residency requirement.
Data privacy
xAI says it does not use API inputs and outputs to train models without the customer's explicit permission.
By default, API requests and responses are stored in encrypted form on servers for 30 days to check for possible abuse, after which they are deleted.
For teams with stricter requirements, there is Zero Data Retention. ZDR is enabled for the entire team and means that prompts and responses are not written to disk.
But ZDR comes with significant feature trade-offs. Features that need server-side storage become unavailable or limited, including stateful Responses, Files, Collections, and Batch API.
So enable ZDR not “for more privacy in general,” but when the project's compliance requirements actually call for it.
safety_identifier for user-facing services
In September 2026, xAI added the safety_identifier field. An application can pass an opaque, stable ID for the end user with a request.
If a policy violation is associated with a particular user, xAI can link the event to that identifier rather than to the API key as a whole.
For a public SaaS or API service, this is more useful than passing an email or another personal identifier: the application can use its own opaque hash and still distinguish end users.
What makes the Grok API stand out
If you compare only regular text generation, Grok is in the same market as OpenAI, Claude, Gemini, DeepSeek, and other APIs. But xAI has a few fairly distinctive strengths.
First, X Search provides native access to current X posts and threads. For tasks related to user reactions, news, trends, and social listening, this is genuinely a separate data source, not just another web search.
Second, the API is built almost entirely around a familiar OpenAI schema. This simplifies migration and makes it relatively inexpensive to add Grok as a second backend.
Third, xAI has already brought text, image, video, and voice generation, documents, RAG, web/X search, code execution, and MCP together on one platform. The Grok API is becoming a fairly broad AI infrastructure platform, not just one model.
Finally, the exact cost_in_usd_ticks in every response is very useful for applications that need to calculate spending for a specific user or operation, rather than estimate it from token usage and a constantly changing price list.
Limitations to consider
xAI has some characteristics to consider before choosing the platform.
The currently recommended Grok 4.7 has a 500K-token context window, while the less expensive Grok 4.3 and 4.20 offer 1M. At the same time, the price of the entire request goes up after 200K prompt tokens.
Not all features are available for every model: for example, Grok 4.7 does not yet support Batch API, and Multi-Agent remains a beta feature.
The model lineup changes quickly. xAI sometimes redirects old model slugs to new models, which is convenient for compatibility but can silently change quality, latency, and cost. For sensitive production workflows, pin a model version and update it in a controlled way.
Finally, server-side tools add cost and make spending less obvious in advance. Web Search, X Search, file search, or code execution may be called several times for a single user request. Agentic features should therefore be measured on a real traffic sample, not just by the price per million tokens.
Which Grok model should you choose?
For a new project where the broadest current capabilities matter, Grok 4.7 is the first model to test. xAI itself recommends it for coding, agents, and general knowledge work.
If cost matters more or you need a 1M-token context, test Grok 4.3 and Grok 4.20 in parallel. For simple operations, the non-reasoning 4.20 option may be more sensible than running reasoning all the time.
Grok 4.20 Multi-Agent is not a regular chat model; it is for deep research where parallel work by several agents is genuinely useful and the much higher token usage is justified.
For images, use grok-imagine-image-2.0; for video, the current Imagine Video; and for conversational interfaces, the separate Voice API. Do not try to solve everything with one text model just to keep a single model ID.
Conclusion
In 2026, the xAI Grok API is a full multimodal AI platform, not just programmatic access to the Grok chatbot. The Responses API provides stateful conversations, reasoning, and server-side tools; Web Search and especially X Search add current external data; Files and Collections handle documents and RAG; Grok Imagine handles images and video, while Voice API covers real-time voice, TTS, and STT.
For developers, OpenAI compatibility, built-in X Search, automatic prompt caching, context compaction, and exact cost calculation for each request are especially interesting. At the same time, keep in mind long-context pricing, the fast model lifecycle, and the fact that agentic tools can substantially increase the final cost of a user workflow.
The practical approach is the same as with other modern AI APIs: do not choose a platform based on one benchmark or token price. Take several real product tasks, compare Grok 4.7 with the less expensive 4.3/4.20, measure quality, latency, and actual cost_in_usd_ticks, and then decide which requests are really worth sending to each model.
Official sources
- Grok API overview: https://docs.x.ai/overview
- Quickstart: https://docs.x.ai/developers/quickstart
- Models: https://docs.x.ai/developers/models
- Grok 4.7: https://docs.x.ai/developers/models/grok-4.7
- Grok 4.3: https://docs.x.ai/developers/models/grok-4.3
- Grok 4.20: https://docs.x.ai/developers/models/grok-4.20
- Grok 4.20 Multi-Agent: https://docs.x.ai/developers/model-capabilities/text/multi-agent
- Pricing: https://docs.x.ai/developers/pricing
- Responses API: https://docs.x.ai/developers/rest-api-reference/inference/responses
- Chat Completions: https://docs.x.ai/developers/model-capabilities/legacy/chat-completions
- Structured Outputs: https://docs.x.ai/developers/model-capabilities/text/structured-outputs
- Reasoning: https://docs.x.ai/developers/model-capabilities/text/reasoning
- Web Search: https://docs.x.ai/developers/tools/web-search
- X Search: https://docs.x.ai/developers/tools/x-search
- Files: https://docs.x.ai/developers/files
- Collections: https://docs.x.ai/developers/files/collections
- Prompt Caching: https://docs.x.ai/developers/advanced-api-usage/prompt-caching
- Context Compaction: https://docs.x.ai/developers/advanced-api-usage/context-compaction
- Cost Tracking: https://docs.x.ai/developers/cost-tracking
- Batch API pricing: https://docs.x.ai/developers/pricing
- Deferred Chat Completions: https://docs.x.ai/developers/advanced-api-usage/deferred-chat-completions
- Priority Processing: https://docs.x.ai/developers/advanced-api-usage/priority-processing
- Imagine API: https://docs.x.ai/developers/model-capabilities/imagine
- Voice API: https://docs.x.ai/developers/model-capabilities/audio/voice
- Regional endpoints: https://docs.x.ai/developers/advanced-api-usage/regions
- Security and Zero Data Retention: https://docs.x.ai/developers/faq/security
- Billing: https://docs.x.ai/console/billing
Related: a DeepSeek API review with current models and pricing.
Our projects
- Anilau
Web development and digital product launch. - Botmarketing
Telegram bots, mini apps, storefronts and CRM for small businesses. - vietnam.anilau.com
Listings, local services and practical guides to Vietnam. - bali.anilau.com
A marketplace for goods and services in Bali. - ceylon.anilau.com
Listings, services and practical information about Sri Lanka. - mauricetop.anilau.com
A platform for listings and information about life in Mauritius. - funlab
A platform for creating and playing AI-generated games. - aura
An AI mood diary and creative space. - drained
A Telegram Mini App for daily fatigue check-ins and recovery. - vietinfodesk
Practical guides, services and help for life in Vietnam. - frau
A cozy Nha Trang cafe profile featuring waffles, breakfast and drinks.
We work in partnership with creative agency Deep.