Moonshot Kimi API Review: Kimi K3, Models, Tools, Pricing, and Access in 2026

Information current as of September 29, 2026.

For a long time, Kimi was known primarily as a Chinese AI assistant with strong long-context capabilities. But Moonshot AI has a separate developer platform — the Kimi API Platform — through which you can connect Kimi models to your own apps, agents, and coding tools.

In 2026, the platform's main product is Kimi K3: an open-weight multimodal MoE model with 2.8 trillion parameters, a one-million-token context window, native vision, reasoning, and tool calling. Less expensive Kimi K2.6 and the specialized Kimi K2.7 Code are also available. The API follows the OpenAI Chat Completions format and supports Structured Outputs, files, automatic context caching, the Batch API, and a set of official tools for agentic use cases.

At the same time, Kimi API differs considerably from both OpenAI and other Chinese platforms such as DeepSeek or Alibaba Model Studio. The platform is comparatively compact, its main endpoint is still stateless Chat Completions, web search is currently in a transition, and its data-handling terms allow customer content to be used to improve models unless agreed otherwise.

This review explains how Moonshot Kimi API works, which models are available and what they cost, how to connect K3 through the OpenAI SDK or PHP, what caching and tools offer, where data is stored, and what to know about registration, payment, and using the service from different countries.

Moonshot AI, Kimi, and Kimi API: what is what?

Moonshot AI is the company that develops the Kimi model family.

Kimi refers to the company's user-facing products: the web assistant, Kimi Code, Kimi Business, and other interfaces.

Kimi API Platform is a separate developer platform at:

https://platform.kimi.ai

Kimi API and consumer subscriptions are billed separately. A Kimi subscription does not become API balance, and pay-as-you-go API usage does not automatically provide benefits from a consumer subscription.

The main API endpoint is:

https://api.moonshot.ai/v1

The most common interface is OpenAI-compatible Chat Completions:

POST /v1/chat/completions

Files and the Batch API are available through the same OpenAI-compatible schema.

Current Kimi models

As of late September 2026, the main public lineup looks like this:

Model Context Input Cache hit Output Main use case
Kimi K3 1M $3 / 1M $0.30 / 1M $15 / 1M complex coding, knowledge work, reasoning, and agents
Kimi K2.7 Code 256K $0.95 / 1M $0.19 / 1M $4 / 1M coding agents and long development tasks
Kimi K2.6 256K $0.95 / 1M $0.16 / 1M $4 / 1M general text, vision, and video tasks

K3 also charges separately for writing a new cache: the platform home page lists $3 per million cache-write tokens. In late September, Moonshot also added 5-minute and 1-hour cache TTL options, so check the current billing interface for the specific mode before estimating production costs.

K2.7 Code and K2.6 have the same regular input/output token prices but behave differently: K2.7 is specifically designed for programming and always uses thinking, while K2.6 is a general-purpose model that lets you turn reasoning off.

Kimi K3

Kimi K3 is Moonshot AI's current flagship model. It is a 2.8-trillion-parameter Mixture-of-Experts model that activates 16 of its 896 experts for each token.

The model is built on the Kimi Delta Attention and Attention Residuals architecture components and natively understands images and video. Its context window is one million tokens.

Moonshot positions K3 for:

  • long-horizon software engineering;
  • work with large codebases;
  • complex agentic workflows;
  • analysis of large document sets;
  • knowledge work;
  • deep reasoning;
  • tasks that connect text reasoning with images or screenshots.

K3 always runs in thinking mode. You cannot turn reasoning off completely, but you can choose its level:

low
high
max

The default is max.

This is an important detail when comparing its cost with models where reasoning can be disabled. Even a simple request is processed by K3 as a thinking model, so less expensive K2.6 can sometimes make more sense for bulk classification or simple extraction.

K3 is an open model, but not Apache 2.0

On July 27, 2026, Moonshot published the full Kimi K3 weights and technical report. The official Kimi API is therefore not the only way to use this model.

However, K3 has its own license; it is not Apache 2.0 or MIT without additional conditions.

You can use, modify, distribute, and self-host the model in ordinary applications. Additional terms apply to very large projects.

If a company runs a Model-as-a-Service business and its aggregate revenue exceeds $20 million in any consecutive 12 months, commercial use of K3 requires a separate agreement with Moonshot AI.

If a commercial product exceeds 100 million monthly active users or $20 million in monthly revenue, its interface must prominently display the name Kimi K3.

These thresholds will not apply to most ordinary SaaS or internal corporate systems, but it would be inaccurate to call K3 a model with a fully unrestricted MIT license.

Kimi K2.7 Code

Kimi K2.7 Code is a separate model for programming and coding agents.

Its context window is 256K tokens. The model accepts text, images, and video and supports long reasoning and multi-step tool calls.

Thinking cannot be disabled in K2.7 Code.

Moonshot also offers a high-speed variant:

kimi-k2.7-code-highspeed

The documentation gives an approximate speed of 180 tokens per second and up to 260 tokens per second in short-context scenarios, though the company separately warns that the high-speed model's capacity is currently limited.

K2.7 is worth using not as a generally cheaper K3, but specifically for coding: generating and editing code, refactoring, autonomous coding agents, and extended development tasks.

Kimi K2.6

Kimi K2.6 is a general-purpose 256K model for text, images, and video.

Unlike K2.7 Code and K3, you can turn off thinking:

{
  "thinking": {
    "type": "disabled"
  }
}

This makes the model useful for scenarios where complex reasoning is not always needed. For example, the same application can use K2.6 without thinking for extraction, then enable thinking for document analysis or a more complex user request.

K2.6 also supports tool calling, but with thinking enabled there are additional restrictions on tool_choice, and official web search is temporarily incompatible with thinking mode.

OpenAI-compatible API

Kimi API is compatible with the OpenAI API format. For most text integrations, this means you can use the official OpenAI SDK and change two settings:

base_url = https://api.moonshot.ai/v1
model = kimi-k3

Python example:

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_KIMI_API_KEY",
    base_url="https://api.moonshot.ai/v1",
)

response = client.chat.completions.create(
    model="kimi-k3",
    messages=[
        {
            "role": "user",
            "content": "Explain dependency injection."
        }
    ],
)

print(response.choices[0].message.content)

This makes it much easier to test Kimi in an existing app built on OpenAI Chat Completions.

There is an important difference from OpenAI in 2026, though: Kimi does not have a public equivalent to the OpenAI Responses API as its primary modern interface. The main documented endpoint is still /v1/chat/completions, and your app stores and passes conversation history itself.

A regular request with cURL

A minimal K3 call:

curl https://api.moonshot.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $MOONSHOT_API_KEY" \
  -d '{
    "model": "kimi-k3",
    "messages": [
      {
        "role": "user",
        "content": "Briefly explain the difference between REST and GraphQL."
      }
    ]
  }'

For K3, sampling parameters such as temperature, top_p, presence_penalty, and frequency_penalty are effectively fixed. The documentation recommends leaving them out.

Reasoning is configured separately:

{
  "model": "kimi-k3",
  "reasoning_effort": "high",
  "messages": [
    {
      "role": "user",
      "content": "Analyze the application's architecture."
    }
  ]
}

Preserve reasoning in multi-turn conversations

For Kimi, thinking is part of the model state.

With the streaming API, reasoning is returned separately from the final content in the reasoning_content field.

If the conversation continues or the model makes tool calls, Moonshot recommends placing the complete assistant message in the next request rather than saving only the final text.

This is especially important for K3 and K2.7 Code. If you remove the reasoning state and return only ordinary content to the model, the next step may lose an important part of the context or fail in a tool workflow.

Therefore, it is better for a Kimi conversation persistence layer to store the original response-message structure rather than a simple role + text pair.

Vision: images and video

K3, K2.7 Code, and K2.6 support text, image, and video input.

You can pass an image as base64. Public image URLs are not currently supported for K3, so use base64 or an uploaded file through ms://<file-id>.

Moonshot recommends staying below 4K for images and Full HD for video: higher resolution increases token usage and processing time without improving the model's understanding.

For large videos and files you reuse, the Files API is a better choice.

Structured Outputs

K3 can return results that strictly follow a JSON Schema.

For example:

{
  "response_format": {
    "type": "json_schema",
    "json_schema": {
      "name": "person",
      "strict": true,
      "schema": {
        "type": "object",
        "properties": {
          "name": {
            "type": "string"
          },
          "age": {
            "type": "integer"
          }
        },
        "required": ["name", "age"],
        "additionalProperties": false
      }
    }
  }
}

This is useful for extraction, classification, creating DTOs, processing documents, and any scenario where the backend consumes the response.

Make sure to parse the final message.content, not the reasoning.

There is also a simpler JSON Mode for cases that do not need an exact schema.

Function calling and Tool Calls

Kimi supports OpenAI-style function calling.

The application describes functions:

{
  "type": "function",
  "function": {
    "name": "get_order",
    "description": "Get information about an order",
    "parameters": {
      "type": "object",
      "properties": {
        "order_id": {
          "type": "string"
        }
      },
      "required": ["order_id"]
    }
  }
}

The model returns tool_calls, the backend executes the actual functions, and then the results are passed back in messages.

K3 also has tool_choice, including an option to require at least one tool call.

For multi-step work, the same key rule applies: preserve the full assistant message without losing reasoning or tool metadata.

Dynamic Tool Loading

K3 has an interesting feature called Dynamic Tool Loading.

Instead of passing all tool definitions at the beginning of a conversation, you can add a separate system message with a tool definition when that tool becomes necessary.

This is useful for large agents with dozens or hundreds of possible functions. Passing all schemas in every request uses up context and increases the token bill.

Dynamic loading lets you add the needed tools gradually and keep their definitions in later history only from the point when they are introduced.

Official Kimi tools

Kimi API has a separate ecosystem of official tools that runs through Formula.

The tools currently listed by the platform include:

  • Web Search;
  • Rethink;
  • Random Choice;
  • Memory;
  • Excel;
  • Code Runner;
  • QuickJS;
  • Date;
  • Fetch;
  • Convert;
  • Base64.

The schema differs from OpenAI built-in tools. First, the developer gets definitions through Formula /tools and passes them to the regular Chat Completions tools. Then the returned tool call is executed through Formula /fibers, and its result is returned to the model.

So this is ready-made tool infrastructure, but it still sits on top of the regular tool-calling cycle.

An important caveat about Web Search

The Kimi API home page prominently presents Web Search as an official tool, and the Help Center lists its separate price — $0.004 per call.

However, the current Kimi K3 page also warns that web_search is being updated and is not recommended for production workflows in the near future.

The K2.6 documentation has the same warning.

So, as of late September 2026, it is accurate to describe Kimi Web Search as an existing platform feature, but not as a fully stable replacement for production web search from OpenAI, Google, or Perplexity. Check its current status at the time of integration.

File API and document questions

Kimi API can upload documents and extract their contents.

A typical workflow is:

  1. Upload a file through /v1/files with purpose="file-extract".
  2. Get the extracted content.
  3. Put the extracted content itself, not the file_id, in a system message.
  4. Ask the model a question.

For images and PDFs, extraction includes OCR.

The documentation specifically emphasizes that a file ID does not by itself become an automatic retrieval tool. For text Q&A, the model receives the extracted text in its prompt.

You can work with multiple files: add each extracted document as a separate system message.

Each user is limited to up to 1,000 uploaded files, so Moonshot recommends deleting objects you no longer need after extraction and storing the extracted text yourself if necessary.

Kimi File Q&A is not a ready-made vector RAG

The file-based Q&A mechanism is considerably simpler than OpenAI File Search or Gemini File Search.

Kimi extracts a document, then sends all relevant content in the model's context. With K3's million-token context window, this can be very convenient: you can analyze a comparatively large set of documents without a separate vector database.

For tens of thousands of documents, however, this approach becomes expensive and technically inconvenient.

Moonshot itself distinguishes Context Caching from RAG: caching is especially useful for frequent questions about a fixed large context, while retrieval is better for very large collections where each request uses only a small part of the corpus.

Context Caching

Automatic caching is one of the most important economic features of Kimi API.

If several requests share the same long prefix — for example:

  • system prompt;
  • documentation;
  • codebase context;
  • tool definitions;
  • earlier conversation history,

then the platform may reuse the cached portion instead of running regular input inference again.

For K3:

regular input:  $3.00 / 1M
cache hit:      $0.30 / 1M

That makes repeated input ten times less expensive.

A cache hit costs $0.19 instead of $0.95 for K2.7 Code, and $0.16 instead of $0.95 for K2.6.

Caching works automatically

In the usual mode, Kimi identifies a matching prefix on its own. You do not need to create a cache ID or rewrite the API request.

For a request to qualify for prefix caching, the preceding prompt must contain more than 256 tokens.

The most important practical recommendation is simple: put stable, large context at the beginning of the messages and avoid changing it unnecessarily. Place the user's dynamic question after it.

Moonshot says that in suitable long-context scenarios, caching can reduce costs by up to 90% and significantly shorten time to first token.

In late September 2026, the platform also began rolling out separate 5-minute and 1-hour TTL options, making cache-write costs more explicit for longer agentic sessions.

Batch API

The Batch API is available for a large number of background requests.

It works with JSONL:

upload file → create batch → wait → download results

The processing window is set through completion_window; a typical option is 24h.

Batch saves 40% compared with realtime inference.

There is one limitation: current documentation supports only kimi-k2.6 and kimi-k2.5 in the Batch API. Kimi K3 and K2.7 Code are not supported yet.

So Batch is not currently a way to get the same K3 capability at a lower price.

What does a large request cost?

Consider a hypothetical task with 100K regular input tokens and 10K output tokens.

Model Cost without cache
Kimi K3 $0.45
Kimi K2.7 Code $0.135
Kimi K2.6 $0.135

The K3 calculation is straightforward:

100,000 × $3 / 1M = $0.30
10,000 × $15 / 1M = $0.15
Total: $0.45

If the same 100K input tokens are fully cached, the input portion drops to $0.03, and the total request costs about $0.18 plus any cache-write charge.

So K3 may look fairly expensive at headline prices, but caching can significantly change the actual economics for coding agents and repeated-document workflows.

K3 is not one of the cheapest Chinese APIs

This is where Kimi differs sharply from DeepSeek or Qwen Flash.

K3 costs $3/$15 per million input/output tokens, placing it closer in price to Western frontier models than to ultra-cheap Chinese inference.

K3's strength is not the lowest token cost, but the combination of:

  • 1M context;
  • always-on reasoning;
  • long-horizon coding;
  • multimodal understanding;
  • agentic tool use;
  • open weights.

If the task is simple, K3 may be excessive on cost. K2.6 is available for bulk processing within the platform, and there are many even less expensive models outside Kimi API.

Kimi K3 and coding agents

Moonshot actively positions Kimi as a backend for coding agents.

The documentation has dedicated instructions for:

  • Claude Code;
  • Codex;
  • OpenCode;
  • OpenClaw;
  • Hermes Agent;
  • other coding tools.

The API itself is OpenAI-compatible, so an agent that lets you configure a custom base URL can often connect without a separate SDK.

Long context, multimodal input, and tool calling are especially important for K3: the model can work with a large codebase, terminal tools, and screenshots in the same agent loop.

Kimi Code and Kimi API are different products

Kimi Code, Kimi Membership, and Kimi API have separate pricing.

This matters because the name “Kimi Code” may suggest that a paid user subscription includes an API allowance for coding tools. It does not.

If Claude Code, Codex, OpenClaw, or another tool uses your API key with api.moonshot.ai, usage is billed through pay-as-you-go Kimi API and its rate limits.

Registration and API key

Developers use:

https://platform.kimi.ai

The process is:

  1. Register for or sign in to a Kimi account.
  2. Open the API Platform.
  3. Add funds to the API balance.
  4. Create a key in API Keys.
  5. Copy the key and save it in secret storage.

K3 requires at least one successful top-up. The current minimum cumulative top-up to unlock K3 is $1.

A new API account does not automatically include free inference balance. You may be able to create a key without funds, but real requests will require a top-up.

Rate limits depend on top-up amount

Kimi uses a tier system.

As cumulative top-ups grow, the account automatically receives higher concurrency, RPM, TPM, and TPD limits.

This differs from a system where rate limits depend only on the selected paid plan. With Kimi, spending history itself becomes part of the account tier.

Production companies can request organization verification, enterprise SLAs, higher limits, and separate terms through sales.

How Kimi API is billed

API usage is pay-as-you-go.

There is an important detail: Kimi's published Help Center pages currently show slightly different payment methods depending on locale and billing setup.

The Platform Terms explicitly provide for credit/debit cards. The Russian-language Help Center lists Visa, Mastercard, and other major cards; verified organizations can arrange bank transfers through sales.

The English-language guidance for individual top-ups separately mentions WeChat Pay and Alipay QR, while business accounts may use online payments or bank transfers depending on region and account configuration.

For an international registration, check the payment methods actually shown in your API Console rather than assuming the same set is available to every account.

Invoices and enterprise billing

You can request invoices for top-ups in the Kimi API Console. Verified organizations can arrange separate commercial agreements, bank transfers, volume discounts, and enterprise SLAs.

This distinguishes the API Platform from an ordinary Kimi subscription: the developer product is designed for both individuals and direct B2B contracts with Moonshot AI.

The international Kimi OpenPlatform is currently operated by Moonshot AI PTE. LTD., a Singapore company.

Can Kimi API be used from Russia?

It is best to avoid categorical claims here.

I did not find a separate “Russia is prohibited” clause in the public Kimi OpenPlatform Terms. The terms are broader: users must not be subject to trade restrictions, sanctions, or other applicable legal restrictions.

At the same time, the official Kimi Help Center says the API is primarily intended for supported regions, and access from other regions may depend on network availability without a stability guarantee. The platform does not publish a clear country whitelist.

Recent third-party guides from Russia report two practical issues: some users run into geographic restrictions during registration, and Russian bank cards fail international payment processing.

So it is not safe to describe Kimi API as “officially fully available in Russia.” A more accurate description is that the Terms do not publish a separate Russian ban, but registration and billing may depend on region and payment method.

If a direct account is unavailable, Moonshot itself says that K3 is also offered through inference partners. Open-weight K3 provides a separate self-deployment option as well.

Should you use a VPN?

The official documentation does not recommend VPNs as a way to bypass regional restrictions.

If the platform does not let you register or add funds from a particular jurisdiction, disguising the registration country is a poor production strategy: it creates a risk of account suspension and losing API access.

For a commercial project, it is more sensible to use a supported international account, an official inference partner, or self-hosted open weights in accordance with the K3 license.

Where is data stored?

The international Kimi OpenPlatform is operated by Moonshot AI PTE. LTD. in Singapore.

The Privacy Policy explicitly says that collected information is stored on secure servers in Singapore. When necessary, data may be transferred across borders with appropriate safeguards.

This is a notable difference from the official DeepSeek API, where the main data location is associated with China.

The Privacy Policy has a separate section for European users describing GDPR-like rights and legal bases for processing.

Does Kimi API use data to train models?

Do not automatically apply the rule that “commercial API data is never used for training.”

The Platform Terms expressly allow Moonshot to use Customer Content to provide, support, develop, and improve its Services.

The Privacy Policy also says user content may be used to improve and develop the platform, including training and refining underlying machine-learning models.

If an organization needs limits on the use of customer content for training or model improvement, the Terms say to arrange an enterprise arrangement or separate written agreement with Moonshot AI.

This should be considered before integrating confidential code, legal documents, internal company data, or other sensitive information.

This matters especially for coding agents

It is natural to connect Claude Code, Codex, or another agent to K3 and give it access to a large part of a repository.

But if a codebase contains closed-source code, customer data, secrets, or proprietary business logic, processing terms matter more than a model benchmark.

Never put an API key, password, or production credential in a prompt. Before sending a confidential repository to the managed Kimi API, a company should check its contractual limits on the use of Customer Content and, if necessary, arrange an enterprise agreement or self-host K3.

Personal data limitations

The Terms require developers to have a lawful basis for processing personal data and to provide users with the necessary privacy notices.

Moonshot separately prohibits using OpenPlatform to process Protected Health Information as defined by HIPAA.

So Kimi API should not be chosen as a medical data processor simply because the model is good at analyzing documents.

For apps that process user data, you need to consider cross-border transfers to Singapore and applicable local laws.

Kimi API in PHP

You do not need a separate official PHP SDK: the OpenAI-compatible API uses ordinary HTTP.

Example:

<?php

$apiKey = $_ENV['MOONSHOT_API_KEY'];

$payload = [
    'model' => 'kimi-k3',
    'messages' => [
        [
            'role' => 'user',
            'content' => 'Briefly explain dependency injection.',
        ],
    ],
];

$ch = curl_init('https://api.moonshot.ai/v1/chat/completions');

curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json',
    ],
    CURLOPT_POSTFIELDS => json_encode(
        $payload,
        JSON_UNESCAPED_UNICODE
    ),
]);

$response = curl_exec($ch);

if ($response === false) {
    throw new RuntimeException(curl_error($ch));
}

$data = json_decode(
    $response,
    true,
    flags: JSON_THROW_ON_ERROR
);

echo $data['choices'][0]['message']['content'] ?? '';

In Laravel, the HTTP Client is more convenient:

$response = Http::withToken(config('services.kimi.key'))
    ->post('https://api.moonshot.ai/v1/chat/completions', [
        'model' => 'kimi-k3',
        'messages' => [
            [
                'role' => 'user',
                'content' => 'Briefly explain dependency injection.',
            ],
        ],
    ]);

$text = $response->throw()->json('choices.0.message.content');

Store the API key in .env and retrieve it through config/services.php.

Do not pass the key to the frontend: the backend should control users, rate limits, tool calls, and costs.

What is currently missing from Kimi API?

In terms of platform breadth, Kimi currently trails OpenAI, Gemini, and Alibaba Model Studio.

The public developer platform currently has no comparable set of native endpoints for:

  • image generation;
  • video generation;
  • realtime speech-to-speech;
  • TTS;
  • STT;
  • embeddings;
  • fully managed vector File Search;
  • a managed agent runtime on the level of the OpenAI Agents API or Google Managed Agents.

The consumer version of Kimi may have features that are not available through the API. In particular, the Help Center explicitly says that the API for PPT generation and Deep Research is currently unavailable, even though these capabilities exist in the consumer product.

So Kimi API is best evaluated as a strong multimodal LLM backend for reasoning, coding, long context, and tool-driven agents, rather than as a universal media AI platform.

Cloud Kimi API and self-hosted K3 are different options

The official Kimi API does not offer standard on-premises deployment through the API Platform.

But K3 has open weights, so you can technically deploy it yourself or use it through a third-party inference provider.

These are two different options.

Moonshot API provides managed inference, caching, billing, Files, Formula tools, and ready-made infrastructure.

Self-hosted K3 gives you control over your data and model, but its 2.8 trillion parameters make in-house deployment extremely demanding. Even with high sparsity, this is not a model you can download to an ordinary VPS.

For most small companies, the more realistic choice is the official API or an external inference partner, not a full self-hosted K3 cluster.

When is Kimi API especially interesting?

K3 is worth testing if a product needs truly long context and complex agent behavior:

  • coding agents;
  • analysis of a large codebase;
  • work with a large batch of documents;
  • legal and patent analysis;
  • long-horizon research workflows;
  • multimodal analysis of screenshots and video;
  • complex tool-calling agents;
  • tasks where having a path to open weights matters.

For a simple chatbot or classifier, K3 is often excessive in price and reasoning budget.

Which model should you choose?

If the task is genuinely difficult and needs the strongest current capability, start with Kimi K3.

For coding workloads where speed and lower cost matter, separately evaluate Kimi K2.7 Code, especially its HighSpeed variant.

For a general-purpose backend with some simple requests, Kimi K2.6 is useful: it costs less than K3 and lets you disable thinking.

A practical production setup can use several models. For example, K2.6 without thinking handles extraction and routing, K2.7 Code handles routine coding tasks, and K3 is reserved for genuinely difficult long-context reasoning.

Conclusion

In 2026, Moonshot Kimi API is a fairly unusual alternative to OpenAI, Claude, Gemini, and other major AI APIs.

The platform's main bet is not the broadest possible service lineup, but the model itself: Kimi K3 offers a million-token context window, native multimodality, always-on reasoning, strong tool use, and long-horizon coding. The model also has open weights, leaving a path to independent inference.

Integration is technically simple thanks to OpenAI-compatible Chat Completions. The platform includes Structured Outputs, Files, Batch for K2.6, automatic context caching, and an official tool ecosystem. But there is no OpenAI-level Responses API yet, Web Search is currently being updated, and the media/audio API set is much smaller than on universal platforms.

The organizational details deserve special attention. The international OpenPlatform is operated by Singapore-based Moonshot AI PTE. LTD., data is stored in Singapore, and the standard terms allow customer content to be used to develop and improve models unless separate limits are agreed.

K3 is not trying to be another DeepSeek on price either: $3/$15 per million tokens puts it in the frontier segment. So the main question when choosing Kimi should not be “Is it cheaper than GPT?” but “Does K3 provide a clear advantage on our long and difficult task?”

If your own evaluations confirm that, the million-token context, automatic caching, strong agentic coding, and open weights make Kimi one of the more interesting platforms in its class.


Official sources

Date of publication:

Moonshot Kimi API Review: Kimi K3, Models, Tools, Pricing, and Access in 2026

Our projects

  • Anilau
    Web development and digital product launch.
  • Botmarketing
    Telegram bots, mini apps, storefronts and CRM for small businesses.
  • vietnam.anilau.com
    Listings, local services and practical guides to Vietnam.
  • bali.anilau.com
    A marketplace for goods and services in Bali.
  • ceylon.anilau.com
    Listings, services and practical information about Sri Lanka.
  • mauricetop.anilau.com
    A platform for listings and information about life in Mauritius.
  • funlab
    A platform for creating and playing AI-generated games.
  • aura
    An AI mood diary and creative space.
  • drained
    A Telegram Mini App for daily fatigue check-ins and recovery.
  • vietinfodesk
    Practical guides, services and help for life in Vietnam.
  • frau
    A cozy Nha Trang cafe profile featuring waffles, breakfast and drinks.

We work in partnership with creative agency Deep.

Try our plugins for Codex and Claude