Zhipu AI
Free value-added services
AI agent Plugins and Skills

Zhipu AI

As a key player in the development of large-scale models, Zhipu AI combines these models with its own high-quality, large-scale knowledge graphs, thereby creating an artificial intelligence framework driven both by data and knowledge – a framework that helps to break through the limitations imposed by current second-generation artificial intelligence technologies...

Tags:

What is Zhipu AI?

Zhipu AI is a platform for the development of large-scale models, operated by Beijing Zhipu Huazhang Technology Co., Ltd.

Its services include GLM models, the BigModel open platform, as well as native applications and enterprise solutions such as Zhipu Qingyan.

Product and Service Matrix

ProductsMain functionSuitable for
GLM modelProvides basic capabilities for text, multimodal content, and agents.Research and application development
BigModelCall models and tools through APIsDevelopers and enterprises
Zhispu QingyanA general AI assistant designed for Chinese speakersIndividual and office users
Z.aiOffers chat, programming, and agent experiences.Global users and developers
AutoGLMPlan and execute cross-application tasksAgent users
AutoClawDeploy a long-running personal AI assistantAutomation and office users
Zread.aiMaterials to assist reading and comprehensionResearchers and knowledge workers
AMinerAcademic searching, expert and technological intelligenceResearch and industrial users

GLM model family

Zhipu offers a variety of models that transform general text into visual content, audio, images, videos, and intelligent agents.

  • GLM-5.3 emphasizes programming skills, the ability to handle long-term tasks, and network security defense capabilities.
  • GLM-5.2 is designed for project-level engineering contexts and the execution of long-term tasks.
  • The GLM-5V-Turbo supports the input of images, videos, text, and files.
  • GLM-5-Turbo and GLM-5 are designed for general-purpose agents and complex tasks.
  • The GLM-4.7 series offers high intelligence, low prices, and a free version.
  • For specific context, output limits, and tool capabilities, please refer to the corresponding model page.

GLM-5.3

GLM-5.3 was released in August 2026; it uses the same base as GLM-5.2, with improvements mainly coming from post-training.

  • The authorities place strong emphasis on complex programming skills and the ability to handle long-term tasks.
  • The model also demonstrates cybersecurity defense capabilities such as vulnerability detection.
  • The official weights have been made available on Hugging Face’s zai-org organization.
  • The standard version has around 753B parameters; there is also the GLM-5.3-Flash series.
  • The license for the model card is labeled glm-5.3, and it cannot be considered equivalent to MIT.
  • High-risk network security usage must comply with laws, authorization requirements, and model terms.

Text and Agent API

The open platform offers text generation, reasoning, tool invocation, structured output, caching, and streaming responses.

  • When selecting a model, one should consider capabilities, context, speed, and input/output costs.
  • Function Call allows the description of business tools to be provided for model selection and invocation.
  • Structured output is suitable for generating JSON and feeding it into subsequent business processes.
  • Context caching can reduce the computational cost of processing repeated long prompts.
  • Stream output is suitable for chatting and applications that require immediate feedback.
  • The tool parameters returned by the model must be verified by the server.

Vision, Images, and Videos

Multimodal models can understand images, videos, document layouts, and user interfaces, and they can also generate images or videos.

AbilityRepresentative modelTypical uses
Visual understandingGLM-5V-Turbo, GLM-4.6VAnalysis of images, videos, and documents
Image generationCogView, GLM-ImageCreative concepts and marketing materials
Video generationCogVideoX seriesShort videos and shot drafts
Interface AgentModels related to AutoGLMVisual perception and operation planning
  • Visual comprehension results may misinterpret small text, tables, occlusions, and complex charts.
  • For image and video generation, it is necessary to obtain authorization for the use of persons, materials, and trademarks.
  • Interface operations require permission control, confirmation steps, and audit logs.

Voice capabilities

  • GLM-TTS is used for text-to-speech conversion and emotion expression.
  • GLM-TTS-Clone can generate audio with a similar timbre from short audio clips.
  • The GLM-ASR series is used for speech recognition and custom hotwords.
  • GLM-Realtime supports real-time audio and video interaction as well as long-duration conversations.
  • GLM-4-Voice supports the understanding and generation of speech in both Chinese and English.
  • Before cloning any real-person voice timbre, it is necessary to obtain clear and verifiable authorization.

Online searching and knowledge bases

The open platform offers search tools, vectorization, rearrangement, parsing, image understanding, and knowledge base storage.

  • Search-Std is suitable for basic, quick searches.
  • Search-Pro emphasizes higher recall and multi-engine collaboration.
  • Knowledge vectorization converts documents into retrievable semantic representations.
  • The rearrangement model is used to optimize the order of relevance of candidate segments.
  • In-depth analysis is suitable for complex documents, but it incurs charges per page.
  • Answers from the knowledge base still need to cite their sources and be verified against the original materials.

Guide to Using Zhipu AI API

  1. Register for an account on the BigModel open platform and complete the required authentication.
  2. Create an API Key in the console, restrict access to it, and store it securely.
  3. Select a text, visual, image, video, or audio model based on the task.
  4. Read the corresponding model page to check the context, output limit, and price.
  5. Use the official cURL, Python, or Java examples to make the first call.
  6. Set thresholds for timeout, retries, concurrency, logging, and cost alerts.
  7. Perform JSON validation and business field validation on the structured output.
  8. Set up a whitelist, permissions, and manual approval for tool calls.
  9. Use a real test set to evaluate accuracy, latency, and per-request cost.
  10. After going live, monitor the model version, error rate, and abnormal consumption.

Basic calling approach

Text invocation typically involves submitting the model name and a list of messages to the chat completion interface, after which the content returned by the model is retrieved.

  • System prompts are used to explain roles, boundaries, and output format.
  • User messages should contain only the information necessary to complete the current task.
  • When historical messages are too long, they can be summarized or cached to reduce costs.
  • The API Key is stored only in the server environment and is not included in web pages or client code.
  • Do not write unverified model outputs directly into the core database.
  • Confirmation must be added when it comes to making payments, deleting messages, or sending them out.

Prices for the main model APIs

As of August 30, 2026, the open platform applies differentiated pricing based on the model, input length, and output length.

ModelEnter priceOutput price
GLM-5.28 yuan per million Tokens28 yuan per million Tokens
GLM-5.1 short input6 yuan per million Tokens24 yuan per million Tokens
GLM-5-Turbo short input5 yuan per million Tokens22 yuan per million Tokens
GLM-5 short input4 yuan per million Tokens18 yuan per million Tokens
GLM-4.7 Basic Short Output2 yuan per million Tokens8 yuan per million Tokens
GLM-4.5-Air foundation0.8 yuan per million Tokens2 yuan per million Tokens
GLM-4.7-FlashX0.5 yuan per million Tokens3 yuan per million Tokens
GLM-4.7-FlashFreeFree

Some models have higher costs for long input or long output sequences, so it is not possible to estimate them solely based on the lowest rates listed in the table.

Prices, the number of concurrent free models, and the limits related to promotions may change; the final details are subject to those indicated in the console and on the current pricing page.

Multimodal and tool prices

ProjectOfficial public priceExplanation
GLM-5V-Turbo short inputInput: 5 yuan; Output: 22 yuan per million tokensImage, video files, and text input
GLM-4.6V base1 yuan is entered, 3 yuan per million tokens is outputted.Long inputs result in a higher price range.
CogView-40.06 yuan per timeThe batch price is lower.
CogVideoX-20.5 yuan per timeMultiple resolutions
CogVideoX-31 yuan per sessionBatch processing in tables is not supported.
CogTTS4 yuan per 10,000 timesText to speech
CogTTS-Clone6 yuan per sessionTimbre cloning
Search tool0.01 to 0.05 yuan per transactionCharging is based on the type of search service.

Knowledge base price

ProjectPriceExplanation
Vectorization0.5 yuan per million TokensSuitable for various embedding models
rearrangement0.8 yuan per million TokensSome open-source reordering models are available free of charge.
In-depth analysis0.12 yuan per pageCharged based on the number of pages in the document
Storage0.04 yuan/GB/hourCharging applies beyond the free capacity.
Free storage1GBOfficial statement on the permanently free quota

If the overdue amount in the knowledge base exceeds the specified limit, data that goes beyond the free quota may be deleted; therefore, an independent backup should be kept.

Free quota and billing method

  • The pricing page currently shows that new users receive 20 million Tokens as a gift.
  • The page also shows 120 image and video resource packs.
  • Free resource packages are subject to restrictions such as registration, real-name verification, model specifications, and expiration dates.
  • When making a call, the applicable resource package is usually deducted first, followed by the cash balance.
  • When search results are used as input for a model, model token fees may also be incurred.
  • The available credit and free model policies are displayed in real time on the account console.

Batch processing, fine-tuning, and private instances

  • The Batch API is suitable for large-scale data processing that does not require real-time processing, and the cost of using some models through this API is half that of real-time calls.
  • Model fine-tuning supports LoRA or full training, and is charged based on the number of training tokens.
  • Cloud private instances are billed based on computing units and the number of days.
  • Different models require different numbers of computing units to deploy one instance.
  • When making procurement decisions, businesses need to consider throughput, latency, concurrency, as well as data and service levels.
  • Private instances do not mean that the model weights can be freely downloaded or redistributed.

The difference between open-source models and cloud services

Zhipu has made available a number of model weights and research codes, but the scope of authorization for different projects varies.

  • The official weights of GLM-5 are licensed under the MIT license.
  • The GLM-5.3 model card is currently licensed under a dedicated glm-5.3 license.
  • Early versions of GLM, ChatGLM, as well as vision and speech projects may use different licenses.
  • Before downloading the model, read the entire contents of the repository, the model card, and the license.
  • The BigModel cloud platform, commercial APIs, and Zhipu Qingyan are not fully open source.
  • Open-source software with an open license is not automatically equivalent to software that meets the OSI definition.

Which users are it suitable for

  • Regular users: Carry out Q&A tasks and office work using Zhipu Qingyan or Z.ai.
  • Developer: Invoke text, multimodal, voice, and tool APIs.
  • Companies: Build knowledge bases, agents, and business applications.
  • Content team: Creates images, videos, documents, and audio materials.
  • Researchers: Download the model weights and carry out evaluation or fine-tuning.
  • Automation team: Executes tasks by combining AutoGLM and Function Calls.

Usage restrictions and precautions

  • Models may generate incorrect, outdated, biased, or unfounded content.
  • A long context does not guarantee that all information can be accurately extracted and inferred.
  • For the generation of images, sound effects, and videos, it is necessary to ensure that the relevant rights have been obtained and consent has been given.
  • The leakage of an API Key can lead to data risks and unexpected charges.
  • Tool calls and interface operations require permissions, approval, and auditing.
  • The free quota, model names, context, and prices are updated continuously.
  • Medical, legal, financial, and security scenarios must be reviewed by professionals.

Frequently Asked Questions

What does Zhipu AI mainly offer?

It offers products such as GLM models, the BigModel open platform, ZhiPu QingYan, Z.ai, and AutoGLM.

Which modalities are supported by the Zhipu AI open platform?

It supports text, images, videos, audio, image generation, search, knowledge bases, and agent tools.

How is the charging for Zhipu AI API handled?

Text and multimodal content are typically charged based on the number of input/output tokens, while images, videos, searches, etc. are charged per use.

Does Zhipu AI offer free models?

The current pricing page indicates that GLM-4.7-Flash is available free of charge; the specific rules regarding concurrency and promotions are subject to those in the console.

Do new users of Zhipu AI have a free credit?

The current page shows 20 million Tokens as well as image and video resource packs; the eligibility criteria and validity period are specified on the account page.

Are all Zhipu AI models open source?

No, some models have their weights made available, but the licenses vary; cloud platforms, APIs, and native applications are not all open source.

Does Zhipu AI support private deployment?

The open platform offers private instances and enterprise solutions; the model specifications, computing power, and services must be determined based on the chosen solution.

©️Copyright notice: Unless otherwise specified, all articles on this site are copyrighted bySharing of AI toolsAll content on this site is original; without permission, no individual, media outlet, website, or organization may reproduce, copy, or otherwise distribute it, nor may they create mirrors of it on servers that are not owned by this site. Otherwise, we reserve the right to take legal action against such parties in accordance with the law.

Tools similar to ZhiPu AI