Z.ai and GLM API Review: GLM-5.3, Flash, Tools, Pricing, and Open-Weight Models in 2026

Information current as of September 29, 2026.

Z.ai is one of the AI providers that remains much less familiar outside China than OpenAI, Anthropic, or Google, even though GLM models have been used for a long time in coding agents, OpenRouter, Hugging Face, and other AI services. In 2026, the platform has become considerably more interesting: the current GLM-5.3 and GLM-5.3-Flash are designed primarily for coding and long-horizon agentic tasks, support a million-token context window and adjustable reasoning, and Flash has also gained native multimodality.

But Z.ai is no longer just an API for LLMs. One developer platform includes web search, MCP, function calling, context caching, image and video generation, OCR, speech recognition, and several specialized agents. At the same time, key models are published with open weights, so the cloud api.z.ai is only one way to use GLM.

The platform is also interesting for developers already using the OpenAI SDK, Claude Code, or other coding agents. The main API is compatible with OpenAI Chat Completions, and a coding subscription has separate endpoints for OpenAI- and Anthropic-compatible tools.

This review explains how Z.ai API works, how GLM-5.3 and GLM-5.3-Flash differ, what inference costs, how reasoning, web search, MCP, and caching work, how the regular API differs from the GLM Coding Plan, and what to know about registration, privacy, payment, and self-hosting.

Z.ai, Zhipu AI, and GLM: what is what?

GLM is a model family historically developed by Zhipu AI. It includes language models, vision models, image and video generation, OCR, speech recognition, and specialized agentic models.

Z.ai is the international platform for using these models and related AI services. The international service and API are legally provided by Singapore-based JINGSHENG HENGXING TECHNOLOGY PTE.LTD.

Z.ai API Platform is the developer side of the service, where you can create an API key, add balance, and call models on usage-based pricing.

There is also a separate GLM Coding Plan subscription for coding tools such as Claude Code, OpenCode, Cline, and ZCode. It is not the same as the regular pay-as-you-go API: Coding Plan has its own quotas, endpoints, and billing rules.

It is important to account for this separation from the start. Do not automatically use a Coding Plan key as a regular backend for your SaaS, and do not connect a general API key to a coding endpoint if you want requests charged against your normal balance.

Current GLM-5.3 and GLM-5.3-Flash

In August 2026, Z.ai released two main models in the new generation.

GLM-5.3 is the current flagship for complex programming, agentic engineering, and extended tasks. It is built on the same base model as GLM-5.2, with most of the quality gains coming from scaling post-training.

GLM-5.3-Flash is a separate, more efficient model with a new base architecture. It has 320 billion parameters, about 18 billion of which are active, and is the first natively multimodal model in the GLM-5 family.

For practical selection, you can think of them like this:

Model Context Input Cached input Output Input types
GLM-5.3 up to 1M about $1.40 / 1M about $0.26 / 1M about $4.40 / 1M text
GLM-5.3-Flash up to 1M about $0.15 / 1M about $0.03 / 1M about $0.50 / 1M text, images, video, files

At the time this article was prepared, Z.ai's main static pricing documentation was still behind the releases and listed GLM-5.2 as the flagship at the same $1.40/$4.40 rate. Current provider listings for the direct Z.ai API already show GLM-5.3 at $1.40/$4.40 and GLM-5.3-Flash at $0.15/$0.50. Before launching a large production workload, check the price again in Z.ai Billing/API Console; the model lineup is moving faster than the documentation.

GLM-5.3

GLM-5.3 was released on August 14, 2026. Z.ai did not change the base model after GLM-5.2, but continued scaling reinforcement learning and post-training on complex executable tasks.

The focus is not on a short chat response, but on work that previously required the user to keep breaking tasks into smaller pieces: a large refactor, infrastructure analysis, research tasks, work on a codebase, using tools, and several rounds of checking the result.

The model has a million-token context window. In its official evaluation configurations, Z.ai uses context of up to one million tokens and an output budget of up to 128K tokens for long-horizon tasks.

GLM-5.3 is a text model. If an agent loop needs images, screenshots, documents as visual objects, or video, GLM-5.3-Flash or a separate vision model is a more natural fit.

GLM-5.3-Flash

GLM-5.3-Flash is not simply a smaller version of GLM-5.3. It has a different base architecture, designed from the outset for less expensive long-context inference.

The model has 320 billion parameters and activates about 18 billion. Its attention stack uses sparse and linear attention, and the IndexPool technology reduces the cost of long-context work. Z.ai reports roughly three times less attention compute and 4.4 times less KV cache than GLM-5.3 in its architectural estimate.

The main practical difference is native multimodality. Flash was trained on a multimodal corpus and understands:

  • text;
  • images;
  • video;
  • the visual structure of documents and interfaces;
  • files in supported agentic scenarios.

It can do more than analyze a screenshot: it can use visual feedback as part of a coding loop. For example, it can create a page, inspect its rendered state, find layout issues, and continue fixing them.

At about $0.15/$0.50 per million input/output tokens, Flash costs roughly an order of magnitude less than the flagship and looks like a sensible default for many production scenarios.

Reasoning is always enabled

You cannot fully turn off thinking in GLM-5.3 or GLM-5.3-Flash. Developers control its depth with:

low
high
max

If reasoning_effort is not specified, the default is max.

For example:

{
  "model": "glm-5.3",
  "messages": [
    {
      "role": "user",
      "content": "Analyze the application's architecture and find potential failure points."
    }
  ],
  "reasoning_effort": "high"
}

Z.ai recommends max for coding benchmarks, but there is no need to use it automatically for every request. Simple classification, extraction, or short data transformation may be cheaper and faster at low.

In your own open-weight deployments, there is also a clear_thinking parameter. For chat scenarios, Z.ai recommends explicitly managing whether thinking is retained between turns. In an agentic loop, do not lose the previous reasoning context.

The main API is OpenAI-compatible Chat Completions

The main endpoint for the pay-as-you-go API is:

https://api.z.ai/api/paas/v4/chat/completions

A minimal request:

curl -X POST "https://api.z.ai/api/paas/v4/chat/completions" \
  -H "Authorization: Bearer $ZAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flash",
    "messages": [
      {
        "role": "user",
        "content": "Briefly explain the difference between REST and GraphQL."
      }
    ],
    "reasoning_effort": "low"
  }'

Z.ai also officially demonstrates using the regular OpenAI SDK:

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_ZAI_API_KEY",
    base_url="https://api.z.ai/api/paas/v4/",
)

response = client.chat.completions.create(
    model="glm-5.3-flash",
    messages=[
        {
            "role": "user",
            "content": "Explain dependency injection."
        }
    ],
)

print(response.choices[0].message.content)

This makes an initial GLM test particularly simple for an app that already works with OpenAI Chat Completions.

Official SDKs

In addition to the OpenAI-compatible API, Z.ai has its own SDKs:

  • Python — zai-sdk;
  • Java — ai.z.openapi:zai-sdk.

A native SDK is more convenient when you need Z.ai-specific tools, image/video APIs, or new features that the OpenAI client does not yet expose directly.

For PHP, an official SDK is currently unnecessary: the API is ordinary HTTP and the Chat Completions format is straightforward.

Function calling

GLM supports standard function calling.

The developer describes a function:

{
  "type": "function",
  "function": {
    "name": "get_order",
    "description": "Get order details",
    "parameters": {
      "type": "object",
      "properties": {
        "order_id": {
          "type": "string"
        }
        }
      },
      "required": ["order_id"]
    }
  }
}

The model decides when to call a tool and returns the function name and arguments. The backend performs the actual action and passes the result back.

This lets you build CRM assistants, coding agents, support tools, data-analysis features, and other AI functions.

As with any LLM API, function calling must not replace backend authorization. User permissions, whether financial operations are allowed, data deletion, and other critical checks should remain in regular application code.

Structured Output: do not confuse it with OpenAI strict schema

Z.ai documentation has a Structured Output section, but the main mechanism at the moment is JSON mode:

{
  "response_format": {
    "type": "json_object"
  }
}

The developer describes the expected structure in the system prompt.

The documentation also shows JSON Schema, but validation is performed in client code, for example with the Python jsonschema library. This is not the same contract as strict server-enforced JSON Schema in some other APIs.

This distinction matters in production. If the backend expects a strict DTO, validate the GLM result after JSON mode and handle schema errors separately.

Context Caching

Context caching works automatically. If multiple requests have the same prefix — a system prompt, large document, tool definitions, or conversation history — repeated tokens may be served at a lower price.

The API returns this in usage:

prompt_tokens_details.cached_tokens

Cached input for the current flagship lineup costs much less than regular input. Current list rates for GLM-5.3 and Flash are approximately:

GLM-5.3
input:   $1.40 / 1M
cached:  $0.26 / 1M

GLM-5.3-Flash
input:   $0.15 / 1M
cached:  $0.03 / 1M

The usual practical advice applies: put stable instructions and large persistent context at the beginning of the messages, and avoid changing the format unnecessarily.

Caching is especially important for coding agents. A large share of a request may consist not of the user's new instruction, but of the same instructions, tool descriptions, and recurring project context.

What does a large request cost?

Consider 100K regular input tokens and 10K output tokens.

Model Approximate cost
GLM-5.3 $0.184
GLM-5.3-Flash $0.020

For GLM-5.3:

100,000 × $1.40 / 1M = $0.14
10,000 × $4.40 / 1M = $0.044
Total = $0.184

For Flash:

100,000 × $0.15 / 1M = $0.015
10,000 × $0.50 / 1M = $0.005
Total = $0.020

This is only the token bill. An agentic workflow may also incur web search, separate vision/media API, and other tool costs.

The difference shows why it is worth testing Flash first. If it solves the task well enough, savings over the flagship are substantial.

Web Search

Z.ai has its own Web Search API, and the same search can be connected as a tool directly to Chat Completions.

In a chat request, the tool looks roughly like this:

{
  "type": "web_search",
  "web_search": {
    "enable": true,
    "search_engine": "search-prime",
    "count": 5
  }
}

You can limit domains, result freshness, and the amount of content to retrieve.

The current price is:

$0.01 / call

You can use Web Search separately from the LLM or embed it in an agent loop so the model decides when it needs fresh information.

For production, remember that a single user request may trigger several tool calls, so the cost of a completed scenario can differ from the cost of plain text generation.

Web Reader

Z.ai also has a separate Web Reader.

It takes a specific URL, extracts the page, and returns content in a format convenient for the model. It supports different output formats, cache control, image retention, and summary options.

Separating Search and Reader is useful in an agent scenario: first the model finds relevant sources, then reads only the pages it actually needs.

MCP directly inside Chat Completions

One of Z.ai's more interesting features is server-side support for Model Context Protocol.

You can connect an MCP server in tools:

{
  "type": "mcp",
  "mcp": {
    "server_label": "my-service",
    "server_url": "https://example.com/mcp",
    "transport_type": "streamable-http",
    "allowed_tools": [
      "search",
      "get_document"
    ]
  }
}

Supported features include:

  • Streamable HTTP;
  • SSE;
  • custom headers;
  • restricting available tools;
  • third-party MCP servers;
  • official MCP services in the Z.ai ecosystem.

The model obtains the list of available MCP tools and calls them through Chat Completions. This extends the API without requiring you to implement an MCP client inside your application.

For operations that affect real business data, permissions still need to be enforced on the MCP server side.

GLM-5.3-Flash and vision

GLM-5.3-Flash has vision built into the main model, so a separate visual model is not always needed.

Flash can use images as part of its reasoning. This is especially useful for coding: the model can read an interface screenshot, analyze a visual bug, work with a chart or document, and then change code.

For a specialized visual pipeline, separate vision models such as GLM-5V-Turbo and GLM-4.6V are still available in the API.

The public pricing page lists GLM-5V-Turbo at $1.20 per million input and $4 per million output tokens, and the less expensive GLM-4.6V at $0.30/$0.90.

GLM-OCR

There is a separate GLM-OCR for bulk document recognition.

It is designed for layout parsing of images and PDFs: it extracts text and structural information about the placement of elements.

The current pricing table lists:

input:   $0.03 / 1M tokens
output:  $0.03 / 1M tokens

For a document-processing pipeline, this may make more sense than sending every scan to a large multimodal model.

GLM-Image

Z.ai offers a separate GLM-Image model for image generation.

It uses a hybrid architecture combining an autoregressive model and a diffusion decoder, and is especially suited to images where accurate text placement matters: posters, presentations, infographics, and advertising materials.

The price is simple:

$0.015 / image

Different aspect ratios and custom sizes from 512 to 2,048 pixels on each side are supported.

The API returns a URL for the finished image, which the application downloads separately.

There is also the earlier CogView-4 at $0.01 per image.

Video generation

Z.ai has several families for video generation, including CogVideoX-3 and Vidu.

CogVideoX-3 supports:

  • text-to-video;
  • image-to-video;
  • start/end frame generation;
  • audio;
  • resolutions up to 4K in supported modes;
  • 30 or 60 fps.

The current price is:

$0.20 / video

Vidu Q1 and Vidu 2 have their own options at $0.20–$0.40 per video generation, depending on the mode.

Unlike token-priced LLMs, this is a fixed price per generation request, which makes the actual cost of a media feature easier to estimate.

Speech-to-Text

GLM-ASR-2512 is available for transcription.

The pricing page lists:

$0.03 / 1M audio tokens

Z.ai estimates this at approximately:

$0.0024 / minute of audio

The API supports audio transcription, including realtime streaming.

The public Z.ai developer documentation currently covers ASR much more extensively than a full TTS/realtime voice API family comparable to OpenAI or Gemini. Consider this when comparing platforms for a complex voice pipeline.

Specialized Agents

Z.ai has several server-side Agent APIs, but for now they are more like a set of specialized solutions than one universal managed-agent runtime.

The public catalog includes:

  • GLM Slide/Poster Agent (beta) — creates presentations and posters;
  • General-Purpose Translation Agent — performs multi-step translation;
  • Video Effect Template Agent.

The current pricing page lists $0.70 per million tokens for the Slide/Poster Agent, $3 per million for the Translation Agent, and $0.20 per video for effect templates.

For a general-purpose agent with custom business functions, GLM + function calling/MCP remains simpler, while you retain control of the agent loop yourself.

GLM Coding Plan is a separate product

For developers, Z.ai sells the GLM Coding Plan subscription, which can be connected to Claude Code, Cline, OpenCode, ZCode, and other coding tools.

The current individual plan starts at about $18 per month. After the move to credit-based quotas, it includes GLM-5.3 and GLM-5.3-Flash.

Important: this is not a pay-as-you-go API for your own application.

Coding Plan has a separate OpenAI-compatible endpoint:

https://api.z.ai/api/coding/paas/v4

Claude Code uses a separate Anthropic-compatible integration.

The regular balance/resource package should use the general endpoint:

https://api.z.ai/api/paas/v4

Do not mix these setups: Plan quota and regular API balance are accounted for separately.

For your own SaaS, backend, or public API, you usually need the general API. Coding Plan is for using GLM inside a ready-made coding agent.

Registration and getting an API key

The official Quick Start recommends a simple process:

  1. Open the Z.ai Open Platform.
  2. Register or sign in.
  3. Add funds on the Billing page if necessary.
  4. Open API Keys.
  5. Create a key.
  6. Save it in secret storage.

You can then use the key with the general endpoint.

Z.ai specifically warns against putting an API key in browser/client-side code. Anyone with a public key can spend your balance, so requests from a user-facing application should go through the backend.

Payment

For the regular API, Z.ai uses balance and resource packages. The model API page offers token usage bundles, and the Quick Start links directly to the Billing page for top-ups.

The public Terms describe a Payment Method but do not publish a universal list of supported cards, PayPal, or other options for every country. For a particular jurisdiction, the actual checkout shown in your account is the source of truth.

Keep this in mind when comparing the platform with DeepSeek or Alibaba Cloud, where payment methods are documented in more detail.

For business cooperation and higher volume, Z.ai offers a separate Sales contact.

Can you use Z.ai from Russia?

The current Terms do not have a separate public ban on Russia.

At the same time, Z.ai requires compliance with applicable export-control and sanctions laws and confirmation that neither the user nor their organization is on sanctions lists.

So the absence of a separate Russian geoblock does not mean that any sanctioned person or company is permitted to use the service.

Russian reviews from 2026 report that Z.ai/API itself is accessible from Russia, but the official documentation does not confirm acceptance of Russian bank cards. Given international restrictions, Russian-issued Visa/Mastercard cards should not be treated as a reliable payment option.

If direct billing is unavailable, alternatives include third-party inference providers or self-hosting open GLM models.

For a production project, do not base payments on a VPN or concealing the account's country. It is more reliable to check the official checkout in advance and have a legitimate working payment method.

API data is not used for training without consent

It is important to distinguish regular consumer use of Z.ai from API Services.

In the Additional Terms for API Services, Z.ai states directly that End User Content is used only to provide the API, comply with law, enforce policies, and prevent abuse. API content is not used to develop or improve the service unless the customer has explicitly consented.

This is a stricter rule than the general consumer Privacy Policy.

The Data Processing Addendum further says that content the Customer or its End Users submit and generate through the API is processed in real time to provide the API Service and is not retained on the servers.

Other account, billing, and operational data may of course be stored in accordance with the Privacy Policy and applicable laws.

Where is data processed?

The international Z.ai service is operated by Singapore-based JINGSHENG HENGXING TECHNOLOGY PTE.LTD.

The Privacy Policy and API DPA indicate that the service is generally provided from Singapore and Customer Data is usually processed there.

So do not automatically treat the international api.z.ai as a China data-residency endpoint just because the GLM family was developed by Zhipu AI.

A company with strict requirements for a particular storage or processing region should still confirm the terms directly with Z.ai, because the policy allows international data transfers under applicable law.

The API Terms have unusual restrictions

Z.ai has fairly strict Additional Terms for API Services.

In particular, the API cannot be used to make decisions in several high-risk areas — healthcare, finance, investments, insurance, credit, employment, housing, legal affairs, and other significant decisions.

There is also a separate restriction on using API models, prompts, or generated content to develop, train, fine-tune, or optimize external models without special permission.

Do not confuse this with the licenses for open-weight GLM. If you download GLM-5.3 or Flash weights from Hugging Face, their use is governed by the license of the specific repository. Cloud API Terms and licenses for published weights are different contractual regimes.

Open weights: you can run GLM yourself

Both GLM-5.3 and GLM-5.3-Flash have been published with open weights.

GLM-5.3-Flash is distributed under the standard MIT License.

GLM-5.3 uses a separate GLM-5.3 License. It allows use, modification, deployment, fine-tuning, and commercial use, but adds a condition for extremely large Model-as-a-Service companies: if aggregate revenue for such a business exceeds $10 billion in any consecutive 12 months, a Z.ai security review is required before commercial use.

For most ordinary companies, this restriction has no practical effect.

SGLang, vLLM, Transformers, KTransformers, Unsloth, and other inference frameworks are supported.

Open-weight does not mean “easy to run locally”

GLM-5.3 is a very large model. A single model repository on Hugging Face takes about 756 GB.

GLM-5.3-Flash is much more efficient in active parameters, but it still has 320 billion total parameters. Even quantized, this is not a typical model for one regular VPS or home GPU.

Open weights therefore matter primarily for:

  • large private deployments;
  • inference providers;
  • organizations with their own GPU clusters;
  • data-residency scenarios;
  • research and fine-tuning tasks.

For a small SaaS, the official API will almost always be simpler and less expensive than running a full deployment yourself.

PHP and Laravel

An official PHP SDK is not required: you can call OpenAI-compatible Chat Completions over regular HTTP.

PHP example:

<?php

$apiKey = $_ENV['ZAI_API_KEY'];

$payload = [
    'model' => 'glm-5.3-flash',
    'messages' => [
        [
            'role' => 'user',
            'content' => 'Briefly explain dependency injection.',
        ],
    ],
    'reasoning_effort' => 'low',
];

$ch = curl_init(
    'https://api.z.ai/api/paas/v4/chat/completions'
);

curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json',
    ],
    CURLOPT_POSTFIELDS => json_encode(
        $payload,
        JSON_UNESCAPED_UNICODE
    ),
]);

$response = curl_exec($ch);

if ($response === false) {
    throw new RuntimeException(curl_error($ch));
}

$data = json_decode(
    $response,
    true,
    flags: JSON_THROW_ON_ERROR
);

echo $data['choices'][0]['message']['content'] ?? '';

In Laravel, the same request is easier with the HTTP Client:

$response = Http::withToken(config('services.zai.key'))
    ->post(
        'https://api.z.ai/api/paas/v4/chat/completions',
        [
            'model' => 'glm-5.3-flash',
            'messages' => [
                [
                    'role' => 'user',
                    'content' => 'Briefly explain dependency injection.',
                ],
            ],
            'reasoning_effort' => 'low',
        ]
    );

$text = $response->throw()
    ->json('choices.0.message.content');

Store the key in .env and retrieve it through config/services.php.

If you later need Web Search, MCP, or media generation, put them in separate service classes instead of turning one LLM client into a universal monolith.

Rate limits

Z.ai applies rate limits, but the current numeric limits depend on account/model permissions and are shown in API key management.

So do not hard-code an expectation for a certain RPM or TPM. A production client should handle 429, use exponential backoff, and have its own queues for non-critical requests.

For large enterprise workloads, limits and business terms can be agreed with Sales.

Which model should you choose?

For most new applications, I would start with GLM-5.3-Flash.

The reasons are practical: a million-token context window, multimodality, tool use, adjustable reasoning, and about a tenfold price difference from the flagship.

Consider GLM-5.3 as an escalation tier for tasks where your own evaluations show a clearly more reliable result: complex coding, a long agentic loop, a large refactor, or a genuinely difficult professional task.

For OCR, images, video, and speech recognition, use Z.ai's specialized models instead of trying to solve everything with GLM-5.3.

If you need a coding agent without your own backend, separately compare the GLM Coding Plan with the regular API. A subscription may be more convenient for frequent development, but its key and quota are not meant for arbitrary SaaS traffic.

Strengths of Z.ai API

Z.ai's main advantage at the moment is an unusual combination of several factors.

First, GLM-5.3-Flash is very inexpensive for its capabilities and also has native vision multimodality.

Second, OpenAI-compatible Chat Completions makes it quick to add GLM to an existing architecture.

Third, the platform already has web search and server-side MCP, so you can connect the model to external systems without inventing all the tool infrastructure from scratch.

Fourth, GLM is published with open weights. If the official API stops meeting your needs, there is a technical path to another inference provider or your own infrastructure.

Finally, the API Terms for business/developer use expressly exclude using End User Content to improve the model without consent and provide for realtime API content processing without retention.

Limitations

The platform also has some notable weaknesses.

Documentation does not update in sync with model releases: at the time of publication, the official Pricing section still listed GLM-5.2, even though Z.ai was already promoting GLM-5.3 as the current model. A fast-moving production project needs to recheck Billing and Release Notes regularly.

Structured Output currently looks more like JSON mode than a fully strict server-side JSON Schema.

The voice ecosystem is considerably narrower than OpenAI or Gemini: ASR is available, but the developer stack does not look like a universal realtime voice platform.

Also remember the difference between the general API and Coding Plan endpoints. Incorrect configuration may cause requests to use a different quota than expected.

Finally, the open-weight flagship models are very large. Self-hosting is legally and technically possible, but for a small project it remains more of a backup architecture option than a cheap replacement for a cloud API.

Conclusion

In 2026, Z.ai GLM API is worth considering as more than another little-known Chinese LLM endpoint: it is a complete international AI platform.

GLM-5.3 targets complex agentic engineering and long-horizon coding. GLM-5.3-Flash provides much of the modern agentic infrastructure at a significantly lower price and also understands images, video, and visual documents. Both models have a million-token context window, adjustable reasoning, and open weights.

Around them, the platform offers OpenAI-compatible API, function calling, automatic context caching, Web Search, MCP, OCR, image/video generation, ASR, and specialized Agents. There is a separate GLM Coding Plan for coding tools.

For a practical start, the setup is straightforward: GLM-5.3-Flash as the base model, GLM-5.3 as an escalation tier, the general /api/paas/v4 endpoint for your application, and a separate Coding Plan only for coding agents.

The combination of low Flash pricing, open weights, server-side tools, and relatively strong API data terms makes Z.ai a provider worth including in your own model evaluations — even if your primary production backend remains on OpenAI, Claude, or Gemini for now.


Official sources

Date of publication:

Z.ai and GLM API Review: GLM-5.3, Flash, Tools, Pricing, and Open-Weight Models in 2026

Our projects

  • Anilau
    Web development and digital product launch.
  • Botmarketing
    Telegram bots, mini apps, storefronts and CRM for small businesses.
  • vietnam.anilau.com
    Listings, local services and practical guides to Vietnam.
  • bali.anilau.com
    A marketplace for goods and services in Bali.
  • ceylon.anilau.com
    Listings, services and practical information about Sri Lanka.
  • mauricetop.anilau.com
    A platform for listings and information about life in Mauritius.
  • funlab
    A platform for creating and playing AI-generated games.
  • aura
    An AI mood diary and creative space.
  • drained
    A Telegram Mini App for daily fatigue check-ins and recovery.
  • vietinfodesk
    Practical guides, services and help for life in Vietnam.
  • frau
    A cozy Nha Trang cafe profile featuring waffles, breakfast and drinks.

We work in partnership with creative agency Deep.

Try our plugins for Codex and Claude