Tongyi Listening and Comprehension
An AI assistant that provides real-time transcription, translation, summarization, and meeting analysis.
Tags:AI conference toolsWhat is Tongyi Tingwu?
Tongyi Tingwu is a video and audio transcription and content understanding service offered by Alibaba Cloud. Leveraging the Qianwen large model sowie AI technologies for speech and video, it converts meeting recordings, interviews, courses, training sessions, roadshow presentations, and recorded content into searchable text, and generates summaries, sections, highlights of speeches, action items, keywords, as well as answers to questions.
The products are divided into the Tongyi Tingwu application designed for individuals, and the Alibaba Cloud Tingwu API intended for businesses. Individual users can use it through web pages, DingTalk mini-programs, and browser plugins;
Companies can integrate real-time transcription, document analysis, and translation into OA, IM, CRM, education systems, media repositories, and customer service systems.
Core functions
1. Real-time transcription of audio to text
It records the audio from the microphone or audio stream in real time, provides the recognized text during the meeting, and enables subsequent AI analysis to be carried out simultaneously. The official performance documentation indicates a latency of around 300 milliseconds; however, this value can be affected by factors such as the network connection, audio quality, language used, and the model employed.
The real-time interface supports up to three audio inputs, which enables the separation of different sound channels. Speaker separation can only identify different speakers; it is not possible to determine automatically who is a customer, a manager, or has a specific name – manual renaming is required after the meeting.
2. Transcription of audio and video files
After uploading or submitting recorded audio and video files, the system generates a transcript with timestamps offline. The supported audio formats include MP3, WAV, M4A, WMA, AAC, OGG, AMR, FLAC, and AIFF.
The video format support includes MP4, WMV, M4V, FLV, RMVB, DAT, MOV, MKV, WEBM, AVI, MPEG, 3GP, and OGG.
A single file must not exceed 6 GB in size and 6 hours in duration. The API reads the content using file addresses that are accessible over the public network, without directly receiving local file paths.
It is recommended that a signed address with an expiration date remain valid for at least 3 hours, to prevent it from becoming invalid while tasks are in the queue.
3. Multilingual recognition
Files and real-time services support Chinese, English, Cantonese, Japanese, Korean, as well as various modes of communication in Chinese, English, Japanese, Korean, Cantonese, German, French, and Russian. The specific combinations depend on the sampling rate and whether the interface is used in real time or offline.
Multiple people speaking at the same time, dialects, echoes from distant sources, noise, technical terms, and excessive audio volume can all reduce accuracy. For important meetings, it is necessary to use separate microphones, clear audio channels, and a glossary of terms, and to proofread the original audio before issuing the minutes.
4. Real-time and offline translation
Tongyi Listening and Understanding enables real-time bidirectional translation between Chinese, English, Japanese, Korean, German, French, and Russian; it also supports offline translation of audio and video files. The Free Speech in Chinese and English feature allows for translation into Chinese or English, or for outputting both languages simultaneously.
The translation fee is added to the transcription cost; if two versions in the target language are required, the fee will be calculated for both. Names, terms, numbers, and legal expressions must be reviewed by a native speaker or a professional.
5. Chapter Overview and Full Text Summary
The system divides long recordings into sections based on topic changes, assigns a theme and a location to each section, and generates a summary of the entire text. Users can click on specific time points to listen to the original audio for verification; this feature is ideal for quickly reviewing long meetings and lectures.
An abstract is a compressed version of the transcribed text; any recognition errors will be reflected in the summary. It is not possible to make decisions related to finance, law, human resources, or customer commitments based solely on the abstract.
6. Summary of remarks, tasks, and key terms
The summary of the speech organizes the viewpoints according to the speakers; for tasks, it identifies the person responsible, as well as the task and deadline. Keywords and key points facilitate retrieval. By making the spoken text more formal, redundant elements and modal particles can be removed, resulting in a more readable summary.
The model may mistake discussion suggestions for confirmed tasks, or it may overlook negative words. Tasks must be confirmed by the meeting participants and uploaded to the official project management system.
7. Review of Answers and Custom Prompts
Users can ask questions about the transcribed content, enabling the system to locate the relevant sections and provide answers. The enterprise API also allows for custom prompts, which make it possible to extract specific fields from conversations or generate custom summaries.
Submitting multiple prompts at once will result in separate charges. The prompts should specify the structure of the output, any missing values, and the time points to be referenced, and they should restrict the model to providing answers based solely on the audio and video content, thereby reducing unnecessary additions.
8. Mind map
Tongyi Tingwu can organize audio and video content along with their hierarchical relationships into mind maps, which are useful for reviewing courses, understanding the structure of meetings, and organizing interviews. The mind maps provide an automated summary; however, for complex arguments, it is still necessary to refer to the original text.
9. Video PPT extraction and summary
For videos that contain presentations, the system can extract the slides from the PPT files and create a summary of the presentation. Both full PPT mode and lecture mode are supported; processing one hour’s worth of video usually takes between 2 and 5 minutes, and up to 200 PPT slides can be extracted.
For this feature to work, the PPT must be in the main area of the video; people should not walk in front of the projection screen or block its view for long periods. Small text, animations, and frequent changes can affect the quality of extraction.
10. Service quality inspection and dialogue content extraction
Enterprise APIs offer services for quality inspection of communications and extraction of dialogue content, which are used for sales, customer service, interviews, and call center analysis. Enterprises must obtain permission for recording and automated analysis in advance, and they should inform their employees and customers accordingly.
Usage of the personal version
Individual users can log in to the Tongyi Tingwu application using their phone numbers; they can record audio and video via the microphone, upload files from Alibaba Cloud Disk, and view transcriptions, summaries, and other AI-generated results. The DingTalk mini-program and browser plugins are suitable for use in meetings and while browsing web pages.
The free usage limits and membership rules for individual applications may differ from those of the Alibaba Cloud Development Kit. The trial versions and pricing options listed in the catalog cannot be directly equated with the prices applicable to consumer memberships; the actual terms apply as stated on the purchase page for individual applications.
API integration
Enterprises need to activate Tingwu in the Alibaba Cloud console, create a project and obtain an Appkey, and configure RAM permissions, OSS for result storage, as well as the callback method. Offline tasks can be handled via polling, or through HTTP or RocketMQ callbacks.
By default, the QPS for creating tasks is 20 at the user level, while the QPS for querying task status and results is 100.
Offline transcription is usually completed within 3 hours, except in cases where a large volume of recordings exceeding 500 hours is submitted in a short period of time. Companies need to handle issues such as expired signature addresses, format errors, rate limits, retry attempts for callbacks, and the persistence of results.
Free trial
Users who activate the new API service can enjoy a 90-day free trial period. During this trial, a quota of 48 hours per day is provided, with support for 2 concurrent connections.
Audio and video files are recorded for 2 hours per day, with 1 stream concurrent; there is no time limit for microphone testing.
Once the daily quota is used up, a 24-hour wait is required for it to be refreshed.
Once the free trial period ends or when the commercial version is upgraded, a pay-as-you-go pricing model based on usage time comes into effect. It is necessary to set up alerts for budget limits, balance levels, and usage amounts, in order to prevent unexpected charges resulting from continuous audio streaming or repeated attempts due to errors.
Price and version comparison
| Package or version | Prices, quotas, and core benefits |
|---|---|
| Prices for the new version of the API | As of August 2026, the official prices for the new interface standards are as follows: real-time meeting transcription costs 0.6 yuan per hour, and audio/video file transcription also costs 0.6 yuan per hour; both services include speaker separation, and file transcription additionally offers automatic language detection. Features such as chapter summaries, full-text abstracts, speech summaries, question reviews, mind maps, to-do lists, keywords, key points, conversion of spoken text to written form, as well as each custom prompt, are charged at 0.064 yuan per hour each. When two of these features are used, the cost is 0.128 yuan per hour, with additional charges applied for each extra prompt. Service quality inspection or extraction of dialogue content costs 0.13 yuan per hour; video and PPT extraction along with summarization costs 0.64 yuan per hour; real-time translation costs 4 yuan per hour; and offline translation costs 0.5 yuan per hour. The services of transcription, translation, and those provided by large language models can be used together. For example, one hour of real-time transcription of a single audio stream plus real-time translation into another language costs 0.6 yuan plus 4 yuan, totaling 4.6 yuan; if simultaneous translation into both Chinese and English is required, the cost is 0.6 yuan plus 4 yuan multiplied by 2, which equals 8.6 yuan. For multiple audio inputs, the fee is calculated based on the actual duration of the audio with transcription results; silent or purely noisy streams may also incur a charge under certain circumstances. |
| Price of the old version interface | The price of the old version of the interface is significantly higher, and the official recommendation is to switch to the new version. For the old version, the standard rate for real-time recording is 10.5 yuan per hour, 9.5 yuan per hour for file recording, and 8 yuan per hour for real-time translation; these rates decrease based on the daily usage volume. For offline translation, the old version charges around 0.5 to 0.9 yuan per hour, according to a tiered pricing system. New projects should prioritize using the new version of the interface. |
Prices for the new version of the API
As of August 2026, the official prices for the new interface standards are as follows: 0.6 yuan per hour for real-time meeting transcription, and 0.6 yuan per hour for the transcription of audio and video files. Both services include speaker separation; in addition, file transcription also involves automatic language detection.
Chapter overviews, full-text summaries, speech recaps, question reviews, mind maps, to-do lists, keywords, key points, conversion of spoken language to written form, and each custom prompt are all charged separately at 0.064 yuan per hour. The cost for two of these features is 0.128 yuan per hour, with additional charges applying for each extra prompt.
The cost for service quality inspection or extraction of conversation content is 0.13 yuan per hour; the cost for extracting video and PPT content along with generating summaries is 0.64 yuan per hour.
Real-time translation costs 4 yuan per hour; offline translation costs 0.5 yuan per hour.
Transcription, translation, and large-model functions will work together.
For example, one hour of real-time transcription of a single language along with real-time translation into another target language costs 0.6 plus 4 yuan, which amounts to 4.6 yuan; when simultaneous speaking in both Chinese and English is required, the cost is 0.6 plus 4 multiplied by 2, resulting in 8.6 yuan.
Multiple input channels are billed based on the actual length of the audio for which a transcription has been generated; silent streams or streams consisting solely of noise can also incur fees under certain circumstances.
Price of the old version interface
The price of the old version of the interface is significantly higher, and the official recommendation is to switch to the new version. For the old version, the standard rate is 10.5 yuan per hour for real-time recording, 9.5 yuan per hour for file recording, and 8 yuan per hour for real-time translation; these rates decrease based on the daily usage volume.
For the old version of offline translation, the rate is approximately 0.5 to 0.9 yuan per hour on a tiered basis.
For new projects, it should be ensured as a priority that the new version of the interface is being used.
Data and Privacy
Meetings, interviews, courses, and customer calls may contain personal information and trade secrets. Companies must obtain the necessary authorization for recording, transcribing, and AI analysis, restrict access to OSS, RAM, callbacks, and result files, and establish rules regarding retention and deletion periods.
The terms of the service agreement state that when state authorities request access to data in accordance with the law, Alibaba Cloud is obliged to cooperate as required by law. Industries with high sensitivity should conduct evaluations by taking into account Alibaba Cloud’s security documents, contracts, geographical factors, as well as their own compliance requirements.
Supported platforms
- Web personal applications;
- DingTalk mini-program;
- Browser plugins;
- Importing audio and video files to Alibaba Cloud Disk;
- Alibaba Cloud Console and HTTP API;
- Integrate with enterprise systems through SDKs or APIs.
Open-source status
The Tongyi Listening and Understanding products, as well as the cloud-based transcription services, are not open-source projects. Alibaba Cloud may provide SDKs, example code, or related open-source models, but this does not mean that the listening, recognition, transcription, and content understanding platforms offered by Tongyi can be deployed entirely on one’s own.
Tongyi Listening and Comprehension Tutorial
Complete a basic task.
- Verify recordings, audio, music, and participant authorization;
- Upload or enter clear audio into Tongyi Tingwu.
- Select settings such as language, speaker, and real-time transcription of audio;
- Use audio and video files for transcription to produce transcriptions, dubbed versions, or cleaned-up content;
- Check each segment for names, numbers, pauses, volume, and mood;
- Before exporting, verify the format, loudness, copyright, and privacy requirements;
Create reusable professional workflows
- A test set is created using real noise, accents, and multi-person segments;
- Compare the differences in processing between real-time audio-to-text conversion, transcription of audio and video files, and multilingual recognition.
- Retain the original recordings and the unmodified transcripts;
- Arrange for a manual hearing before releasing it to the public;
- Statistically analyze processing time, error rate, and quota consumption;
- Regularly update the glossary, sound licensing, and deletion policies;
Which users is it suitable for?
- The person in charge of organizing meetings, interviews, assessments, and client visits;
- Teachers and students who create subtitles, chapters, and review summaries for courses;
- A team responsible for handling recorded videos, podcasts, roadshows, and media materials;
- Organizations that need real-time multilingual conference translation;
- Companies that integrate transcription and quality control into their OA, CRM, and customer service systems;
- Developers who need to perform batch analysis of audio and video files in a media repository.
Product advantages
- Integration of transcription, translation, summarization, Q&A, and PPT extraction;
- It supports both real-time streaming and offline file modes;
- Proficient in multiple languages, with fluent speaking skills in Chinese, English, Japanese, Korean, Cantonese, German, French, and Russian.
- There are entry points available for both personal applications and enterprise APIs;
- The prices of the new version of the interface are transparent, and they can be customized based on functional combinations.
Restrictions and Precautions
- The accuracy of recognition depends on sound quality, accent, noise, and multiple people speaking at the same time.
- Any important transcript must be listened to again for proofreading;
- The capabilities of large models are charged on a per-project basis, and customizing multiple prompts also incurs additional charges.
- The cost of real-time translation is significantly higher than that of transcription; it is necessary to take into account factors such as the distance, language, and duration.
- Meeting transcripts involve the privacy of participants and trade secrets; it is necessary to obtain permission first, and recording them secretly or storing them indefinitely is not allowed.
Frequently Asked Questions
Is Tongyi Listening and Understanding free?
The personal application offers a free access option; new API users can use it on a trial basis for 90 days, with recording available for 48 hours per day and 2 hours per day for files, after which pricing is based on actual usage.
What is the maximum size of a single file?
Offline audio and video files can be up to 6 GB in size and up to 6 hours in length; various common audio and video formats are supported.
Can it identify different speakers?
It is possible to distinguish between different speakers, but the true identity cannot be identified automatically; the name must be modified after the meeting.
How much does the Tongyi Listening and Understanding API cost?
The new version costs 0.6 yuan per hour for transcription; 0.064 yuan per hour for analysis using large models; 0.64 yuan per hour for extracting summaries from PPTs; 4 yuan per hour for real-time translation; and 0.5 yuan per hour for offline translation.
Does Tongyi Listening and Understanding support APIs?
Supported: Enterprises can take advantage of functions such as real-time recording, document transcription, translation, summarization, quality inspection, and PPT extraction, and can obtain the results through callbacks or polling.
Is Tongyi Listening and Understanding open-source?
It is not open source. It belongs to Alibaba Cloud’s commercial applications and API services; the availability of public SDKs or the Qianwen open-source models does not mean that the Tingwu platform is open source.
Guigong Network Security Registration No. 45132202000164