Anthropic Claude API Review: Models, Features, Examples, and Pricing in 2026
Information current as of September 29, 2026.
Claude has long since become more than just a ChatGPT alternative in a browser. Anthropic is building a separate developer platform: through the Claude API, applications can connect Claude models to websites and apps, analyze large documents and images, return strictly structured data, give the model access to application functions and external systems, use web search and code execution, and build full AI agents for more complex workflows.
Claude API is a separate product. A Claude Pro, Max, Team, or Enterprise subscription does not automatically include API credits: developer access is configured through the Claude Console and billed separately based on actual usage.
This article looks at how the Anthropic Claude API works in 2026, which models are available, what they cost, how the Messages API differs from Managed Agents, how tool use, prompt caching, and Structured Outputs work, and how to connect to Claude from PHP.
What is the Claude API?
The Claude API is Anthropic's REST API at api.anthropic.com. It gives an application programmatic access to Claude models and Anthropic's agent infrastructure.
For regular model interactions, the primary interface is the Messages API:
POST /v1/messages
An application sends a list of messages, system instructions, a selected model, and additional parameters; Claude returns the next response. The same Messages API supports images, documents, thinking, structured outputs, and most tools.
The platform also includes:
- Message Batches API for asynchronous bulk processing at a 50% discount;
- Token Counting API for counting tokens before sending a request;
- Models API for retrieving the list of available models;
- Files API for uploading files and reusing them in requests;
- Claude Managed Agents for long-running stateful tasks in Anthropic-managed infrastructure;
- administrative APIs for usage, cost, workspaces, keys, and limits.
For most integrations, start with the Messages API. Managed Agents makes sense when an application has outgrown individual requests or a simple agent loop of its own.
Current Claude models
As of late September 2026, the main Claude lineup looks like this:
| Model | Context | Maximum output | Input, per 1M tokens | Output, per 1M tokens | Main use case |
|---|---|---|---|---|---|
| Claude Fable 5.1 | 1M | 128K | $10 | $50 | The most demanding reasoning and long-horizon agentic tasks |
| Claude Opus 5.5 | 1M | 128K | $4 | $20 | Complex development, agents, knowledge work |
| Claude Sonnet 5.5 | 1M | 128K | $2 | $10 | Fast, complex tasks with a strong balance of price and capability |
| Claude Haiku 4.5 | 200K | 64K | $1 | $5 | Speed and relatively inexpensive high-volume processing |
Fable 5.1, Opus 5.5, and Sonnet 5.5 have a reliable knowledge cutoff in June 2026. All current models support text and image input, text output, vision, multilingual capabilities, and tool use.
The token price alone does not tell you which model will be cheaper for a particular task. If Haiku needs to be checked several times by a stronger model, the final cost may exceed a single high-quality request to Sonnet or Opus. For a production project, evaluate models on your own set of real tasks rather than relying only on the pricing table.
Claude Fable 5.1
Fable 5.1 is the most expensive model in Anthropic's main public lineup. Anthropic positions it for demanding reasoning and long-horizon agentic work: complex research, extended tasks, and multi-step work where even Opus at a higher reasoning effort does not deliver the required result.
It has a 1M-token context window, a maximum regular output of 128K tokens, and costs $10 per million input tokens and $50 per million output tokens. Adaptive thinking is always enabled.
Using Fable “just in case” is expensive. It is more sensible to test Opus 5.5 on real evaluation tasks first and move to Fable only where the additional quality changes the outcome.
Claude Opus 5.5
Anthropic recommends Opus 5.5 as its primary model for most demanding workloads, including long-running agentic coding and knowledge work. It has a 1M-token context window, output up to 128K tokens, and costs $4/$20 per million input/output tokens.
Adaptive thinking is also always enabled here, but the default effort is lower than with Fable. This makes Opus a strong general-purpose model for programming, analyzing large sets of materials, multi-step tool use, and other tasks where a reliable result matters more than minimizing the cost of a single request.
For many serious AI features, Opus is a natural starting point: Fable is substantially more expensive, while the faster Sonnet can be evaluated separately for tasks with higher cost or latency requirements.
Claude Sonnet 5.5
Sonnet 5.5 was released on September 28, 2026 and, at the time this article was prepared, is the newest release in the main lineup. It keeps the 1M-token context window and output up to 128K tokens, but costs $2 per million input tokens and $10 per million output tokens.
The model is faster than Opus and supports adaptive thinking. For web apps, API features, agents, and other user-facing scenarios, Sonnet may be an especially interesting compromise: it costs half as much as Opus 5.5 while offering the same large context window and modern toolset.
Still, do not assume in advance that Sonnet is “the best value” for every project. If a task is difficult and Opus consistently needs fewer repeat calls or corrections, the difference in token price may disappear at the level of the completed workflow.
Claude Haiku 4.5
Haiku 4.5 is the fastest and least expensive model in the current main lineup. It costs $1/$5 per million tokens, has a 200K-token context window, and supports output up to 64K tokens.
Haiku is suitable for classification, extracting relatively simple data, short transformations, request routing, and other high-volume operations. At the same time, the gap from inexpensive competing models is already noticeable: some APIs on the market cost substantially less than $1 per million input tokens. So it is worth evaluating Haiku not only within the Claude family, but also against DeepSeek, Qwen, GPT Luna, and other inexpensive models.
What a typical Messages API request looks like
Here is a minimal request using cURL:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-5-5",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": "Briefly explain the difference between REST and GraphQL."
}
]
}'
Claude differs from some APIs in an important way: the system instruction is passed as a separate top-level system parameter. The native Messages API does not have a system role inside the messages array.
The Messages API is stateless. For a multi-step conversation, the application sends the conversation history again with each request: previous user messages, Claude's responses, and the current request. The context includes the system prompt, history, documents, images, tool results, tool descriptions, and generated thinking.
This approach is simple, but it has an obvious cost implication for long conversations: the same early part of the conversation is sent over and over. That is why prompt caching is especially important for Claude.
Thinking and reasoning
Reasoning is a standard part of modern Claude models. Adaptive thinking is always enabled for Fable 5.1 and Opus 5.5. Sonnet 5.5 also supports adaptive thinking and decides for itself when and how deeply to analyze a task.
For developers, the impact of reasoning on price and latency matters just as much as its availability. A complex agentic request can include a large context, several tool-use cycles, and a substantial amount of thinking. So one user instruction does not necessarily equal one inexpensive model call.
Legacy code also needs attention to sampling parameters. For Claude 4.7 and newer models, Anthropic no longer supports temperature, top_p, or top_k. When migrating an older integration, do not treat these parameters as a universal way to control the behavior of new Claude models.
Tool use: connecting Claude to application functions
Tool use is similar to function calling: the application describes an available function with JSON Schema, Claude decides when to call it, and returns a structured tool_use block.
For example, an online store could provide the model with:
{
"name": "get_order",
"description": "Get information about a customer's order",
"input_schema": {
"type": "object",
"properties": {
"order_id": {
"type": "string"
}
},
"required": ["order_id"]
}
}
If a user asks, “Where is my order 54821?”, Claude can return a get_order call with the requested ID. Your application runs the database query. The result is then sent to the model as tool_result, and Claude prepares a response for the user.
This is an important boundary of responsibility: the language model chooses an action, but it does not automatically get direct access to the database or backend. Authorization checks, valid parameters, and business rules remain the application's responsibility.
Anthropic also offers strict tool use. Setting strict: true makes tool-call arguments conform to the specified JSON Schema, which is much more convenient for production scenarios than trying to repair almost-correct JSON yourself.
Anthropic's built-in tools
The Claude API has gradually evolved from “model + function calling” into a platform with its own set of tools. Some run on Anthropic's side, some provide a standardized schema, and actual execution remains the responsibility of your application.
Server-side tools include:
- web search;
- web fetch;
- code execution;
- tool search;
- MCP connector.
For example, web search lets Claude query current online sources itself and return an answer with citations. It costs $10 per 1,000 search operations, plus the regular token charges.
Code execution runs Python and Bash in an Anthropic sandbox container. Claude can perform calculations, analyze data, create files, and use a shell. When code execution is used together with current versions of web search or web fetch, there is no separate charge for code execution; otherwise, container time is billed after the included free allowance.
There are also client-side tools and toolsets: bash, text_editor, memory, computer use, and browser use. Anthropic provides a standard schema that the model understands well, but your application must execute the actions in a controlled environment.
Computer use and browser use
Claude can work not only through regular API functions, but also through interfaces designed to control a computer or browser.
The current computer_toolset_20260801 gives the model actions for screenshots, clicks, text input, zooming, and other graphical interface operations. Your application provides Claude with snapshots of the environment and carries out the requested actions.
For tasks that take place entirely on web pages, there is a separate browser toolset. It can interact more directly with page elements, so a full desktop workflow is not needed when browser access is sufficient.
These tools can automate systems that do not have a suitable API, but they also require the most careful architecture. A model controlling a GUI should run in a restricted environment with clearly defined permissions, especially if the interface can send messages, delete data, place orders, or perform other irreversible actions.
Structured Outputs: reliable JSON instead of free-form text
If Claude's response will be processed by code, the instruction “return JSON” is a weak contract. The model might add an explanation, omit a field, or return a value of the wrong type.
The Claude API has Structured Outputs for this purpose. With output_config.format, you can provide a JSON Schema for the final response, and the API constrains generation so that the result conforms to the schema.
For tool calling, strict: true solves a similar problem. You can use both mechanisms at the same time: tool inputs are validated against the tool schemas, while the final response follows its own JSON Schema.
This is useful for extracting data from documents, classifying support requests, building objects for a backend, processing forms, and nearly any scenario where the model's result must be a reliable programmatic interface rather than polished prose.
There are limits. For example, strict JSON output cannot be used at the same time as citations, because citations require special citation blocks to be inserted into the text response.
Working with images and PDFs
All current Claude models support image input. Claude can analyze photos, screenshots, charts, diagrams, and visual elements in documents, but the main current lineup returns text—there is no native equivalent of OpenAI's image generation API as part of these Claude models.
PDFs are also well supported. A document can be provided:
- by URL;
- as base64;
- through a
file_idfrom the Files API.
When processing a PDF, Anthropic extracts the text and also converts the pages into images. Claude can therefore analyze not only the text, but also tables, charts, diagrams, and other visual elements on the page.
This affects cost: in addition to extracted text, the visual content of every page is counted. Anthropic gives an estimate of about 1,500–3,000 text tokens per page, depending on document density, plus image tokens.
Files API
When a file is used several times, sending it as base64 with every request is inconvenient. The Files API lets you upload a document once and refer to it later by its file_id.
The current maximum size for one file is 500 MB, with total storage of up to 1 TB per organization. On upload, you can set a retention period from one hour to 90 days; otherwise, the file is kept until deleted under the platform's rules.
The Files API is especially useful with PDF analysis, code execution, and agent workflows where the same set of documents is processed across multiple requests.
The Files API should not be confused with a ready-made vector database or automatic RAG. If an application needs to search a large corporate document collection, its retrieval architecture still needs to be designed separately, or implemented with your own tools and search results.
Citations and your own RAG
Claude has built-in citation support not only for web search, but also for documents and your own search results. This is useful when users need to see where a particular conclusion came from.
In a RAG application, the backend can find relevant documents itself and pass them to Claude as search_result content blocks with a title and source. The model can then attach claims in its response to the supplied sources automatically.
This does not replace retrieval: Claude does not automatically know about your internal database. It does solve a frustrating application task, though—linking the final answer to specific source passages without having to ask the model to invent source numbers manually.
Prompt caching: an important part of Claude's economics
The Messages API is stateless, so long conversation histories and system instructions are sent again with every request. In agentic scenarios, this quickly becomes a large part of the cost.
Prompt caching lets you cache a recurring prefix: tool definitions, the system prompt, and previous messages. A cache entry lives for five minutes by default; a one-hour TTL is also available.
For Sonnet 5.5, regular input costs $2 per million tokens, writing to a five-minute cache costs $2.50, writing to a one-hour cache costs $4, and reading from the cache again costs $0.20. For Opus 5.5, a cache read costs $0.20 versus $4 for regular input; for Fable 5.1, it costs $0.25 versus $10.
So the difference can be substantial when a long, unchanged context is reused. Prompt caching is especially useful for long system prompts, large tool schemas, documents, coding-agent context, and extended conversations.
Batch API
If a result is not needed immediately, the Message Batches API lets you submit many Messages requests for asynchronous processing at a 50% discount on input and output tokens.
A batch can take up to 24 hours to process. This mode is suitable for bulk classification, processing a product catalog, analyzing documents, translation, preparing descriptions, and other background operations.
Batch and prompt caching can be combined. For large recurring workloads, this can significantly change the economics compared with sequential real-time requests.
What the Claude API costs in practice
Consider a hypothetical large request with 100K input tokens and 10K output tokens.
| Model | Regular request | Through Batch API |
|---|---|---|
| Claude Haiku 4.5 | $0.15 | $0.075 |
| Claude Sonnet 5.5 | $0.30 | $0.15 |
| Claude Opus 5.5 | $0.60 | $0.30 |
| Claude Fable 5.1 | $1.50 | $0.75 |
This calculation covers tokens only. It does not include paid server-side tools, possible extra agentic cycles, or other operations.
In a real application, it is more useful to calculate cost per completed workflow than “the price of one message.” For example: the cost of processing one contract, resolving one support request, or completing one agent task. A stronger model can sometimes be cheaper if it needs fewer retries and intermediate calls.
Claude Managed Agents
The Messages API gives developers full control, but requires them to build the agent loop, store state, run client tools, and organize long-running tasks themselves. For more complex cases, Anthropic offers Claude Managed Agents, which is currently in beta.
Managed Agents is a managed agent harness with stateful sessions and persistent event history. You create an agent configuration, start a session, and the long-running work takes place in Anthropic's infrastructure.
The beta API has separate Agents, Sessions, and Environments entities. You can connect web search/web fetch, MCP servers, and other tools to an agent; a sandbox is used for files and agent actions.
This layer is unnecessary for a regular chatbot. It becomes interesting for tasks that can run for a long time, need to survive a dropped connection, and involve many sequential actions.
Claude Agent SDK and Managed Agents are not the same thing
Anthropic is also developing the Claude Agent SDK. Conceptually it solves a similar problem, but its runtime runs in a process and infrastructure that you control.
Managed Agents moves the agent harness to Anthropic's side. Anthropic's migration documentation states the distinction directly: the Agent SDK runs in your process, while Managed Agents runs in Anthropic's infrastructure.
The choice depends less on how “intelligent” the agent is and more on the architecture. If you need maximum control over the environment, execution, and data, your own runtime may be more suitable. If it matters more to get a stateful, long-running agent quickly without maintaining all the supporting infrastructure, Managed Agents reduces that work.
Claude API compatibility with the OpenAI SDK
Anthropic offers a compatibility layer for the OpenAI SDK. This makes it relatively quick to try Claude in a project already written for OpenAI: change the API key, base_url, and model name.
Anthropic explicitly warns that this layer is intended mainly for testing and comparison, rather than as the recommended production interface. The full Claude API feature set is available through the native API.
There are specific limitations, too. Prompt caching is not supported in compatibility mode, audio input is ignored, and strict in function calling does not guarantee conformance to JSON Schema. System and developer messages from the OpenAI format are combined into a single initial Claude system instruction.
So the compatibility layer is convenient for a quick A/B test. If Claude becomes a permanent backend, it is better to move to the native Messages API and use its capabilities directly.
Claude API in PHP
In 2026, Anthropic has an official PHP SDK. It requires PHP 8.1 or later and is currently in beta. Anthropic recommends installing Guzzle for streaming.
Install with Composer:
composer require "anthropic-ai/sdk" "guzzlehttp/guzzle:^7"
It is convenient to store the key in an environment variable:
ANTHROPIC_API_KEY=your-api-key
A basic request:
<?php
require __DIR__ . '/vendor/autoload.php';
use Anthropic\Client;
$client = new Client();
$message = $client->messages->create(
maxTokens: 1024,
messages: [
[
'role' => 'user',
'content' => 'Briefly explain what dependency injection is.',
],
],
model: 'claude-sonnet-5-5',
);
foreach ($message->content as $block) {
if ($block->type === 'text') {
echo $block->text;
}
}
In Laravel, it is better to get the API key through the configuration layer rather than calling env() from application code. Never pass the key to the frontend: regular requests to Claude should go through the backend, where you can control the user, limits, and available tools.
The PHP SDK also supports streaming through SSE, API error handling, retries, and different Claude platforms, including Google Cloud, Bedrock, Claude Platform on AWS, and Microsoft Foundry.
Claude is available through providers other than Anthropic
Claude models can be used through several infrastructure options:
- the direct Claude API;
- Claude Platform on AWS;
- Amazon Bedrock;
- Google Cloud;
- Microsoft Foundry.
The direct API usually gets new models and the full feature set first. Cloud platforms are convenient for companies already using the corresponding IAM, billing, compliance, and infrastructure.
The feature set is not completely identical across these options. For example, specific versions of web search, the Files API, or tool use may arrive on one platform earlier than another. When choosing Bedrock, Google Cloud, or Foundry, verify support for the specific API features your product depends on.
Claude Console, API keys, and billing
Claude Console is a separate developer platform. A paid Claude subscription for web or desktop does not include Claude API costs.
Most regular organizations pay for the API through prepaid usage credits: first, the balance is topped up, then successful requests are charged at current rates. Organizations with a separate invoicing arrangement may pay after usage.
New users may receive a small amount of test credits. Spend limits and rate limits also apply and depend on the organization's usage tier.
Workspaces are useful for separate projects and environments: they let you separate API keys, users, and expenses within one organization.
Privacy and data retention
For the commercial Claude API, Anthropic says retained data is not used to train models without the customer's explicit permission.
There is an important qualification about storage: the rules depend on the specific feature and model. Anthropic says conversation content is not retained by default where storage is not required, but Covered Models and certain features have their own retention periods. So it would be too broad to say “the Claude API stores absolutely nothing.”
For organizations with stricter requirements, Anthropic offers Zero Data Retention. With ZDR, prompts and responses for supported features are not stored at rest after the API response is returned. However, ZDR is enabled at the organization level in coordination with Anthropic, and it does not apply in exactly the same way to every feature.
If an application handles medical, financial, or other sensitive data, check the rules for the specific endpoint and tools before launch rather than relying only on the API's general policy.
What the Claude API does not include
Despite its wide feature set, the Claude API is not a complete mirror of OpenAI.
The main Claude models currently accept text and images and return text. If an application needs its own image generation, video, real-time speech-to-speech, or a full TTS/STT stack, it will need another service or a specialized model.
The Files API also does not turn Claude into a ready-made vector database. For a large RAG application, search, indexing, and data update policies remain separate architecture tasks.
Finally, a 1M-token context window does not eliminate the need to manage context. It is technically possible to send a million tokens to a model, but this is not always the best approach in terms of cost, latency, and quality. Retrieval, caching, and good task decomposition often produce more reliable results.
Which Claude model should you choose?
Anthropic offers Fable 5.1 for the most demanding reasoning and long-horizon agentic work, but it is expensive. A sensible approach is to use it only where your own evaluations show a clear advantage over Opus.
Opus 5.5 is a reasonable starting point for complex development, agents, analysis, and knowledge work. Sonnet 5.5 is worth testing separately for user-facing production scenarios where you need to reduce cost and latency without switching to a lightweight model. Haiku 4.5 is suitable for fast, high-volume, relatively simple operations.
In practice, one application can use several models at once. For example, Haiku can classify a request, Sonnet can do the main processing, and Opus or Fable can be called only for difficult cases. This kind of routing is often more cost-effective than sending all traffic to one model.
Conclusion
In 2026, the Anthropic Claude API is a full AI platform, not just an API for a chat model. The Messages API supports regular and agentic integrations; Claude can work with images and PDFs, call application functions, return strictly structured data, search the web, run code, and use MCP. A separate Claude Managed Agents layer is emerging for long-running autonomous work.
The platform's main focus is reasoning, programming, and working with tools. Its architecture also differs substantially from the OpenAI Responses API: the Messages API remains stateless, history is sent again, and prompt caching therefore plays a particularly large role.
For a new project, a sensible approach is to start with real evaluation tasks on Opus 5.5 and Sonnet 5.5, test Haiku separately for inexpensive operations, and move to Fable only where the results justify it. Then the model choice will be based not on its name in the lineup, but on the cost and quality of the completed user workflow.
Official sources
- Claude API overview: https://platform.claude.com/docs/en/api/overview
- Models overview: https://platform.claude.com/docs/en/models/overview
- Claude Sonnet 5.5: https://platform.claude.com/docs/en/models/sonnet-5-5/overview
- Claude Opus 5.5: https://platform.claude.com/docs/en/models/opus-5-5/overview
- Claude Fable 5.1: https://platform.claude.com/docs/en/models/fable-5-1/overview
- Pricing: https://platform.claude.com/docs/en/about-claude/pricing
- Messages API: https://platform.claude.com/docs/en/build-with-claude/working-with-messages
- Tool use: https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview
- Structured Outputs: https://platform.claude.com/docs/en/build-with-claude/structured-outputs
- Prompt caching: https://platform.claude.com/docs/en/build-with-claude/prompt-caching
- Web search: https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool
- Code execution: https://platform.claude.com/docs/en/agents-and-tools/tool-use/code-execution-tool
- PDF support: https://platform.claude.com/docs/en/build-with-claude/pdf-support
- Files API: https://platform.claude.com/docs/en/build-with-claude/files
- Citations: https://platform.claude.com/docs/en/build-with-claude/citations
- Managed Agents: https://platform.claude.com/docs/en/managed-agents/overview
- OpenAI SDK compatibility: https://platform.claude.com/docs/en/cli-sdks-libraries/libraries/openai-sdk
- PHP SDK: https://platform.claude.com/docs/en/cli-sdks-libraries/sdks/php
- API and data retention: https://platform.claude.com/docs/en/manage-claude/api-and-data-retention
Related: a Google Gemini API review with current models, tools, and pricing.
Related: an xAI Grok API review with current models and tools.
Related: a DeepSeek API review with current models and pricing.
Our projects
- Anilau
Web development and digital product launch. - Botmarketing
Telegram bots, mini apps, storefronts and CRM for small businesses. - vietnam.anilau.com
Listings, local services and practical guides to Vietnam. - bali.anilau.com
A marketplace for goods and services in Bali. - ceylon.anilau.com
Listings, services and practical information about Sri Lanka. - mauricetop.anilau.com
A platform for listings and information about life in Mauritius. - funlab
A platform for creating and playing AI-generated games. - aura
An AI mood diary and creative space. - drained
A Telegram Mini App for daily fatigue check-ins and recovery. - vietinfodesk
Practical guides, services and help for life in Vietnam. - frau
A cozy Nha Trang cafe profile featuring waffles, breakfast and drinks.
We work in partnership with creative agency Deep.