Tongyi Wanshang
A multi-modal visual creation platform under Alibaba Tongyi, covering image and video generation.
Tags:AI image illustration generationWhat is Tongyi Wanshang?
Tongyi Wanxiang is a multi-modal visual generation platform part of Alibaba’s Tongyi series; its English name is Wan. It provides online tools for creating images and videos for ordinary creators, offers APIs to developers and enterprises through Alibaba Cloud Bai Lian, and also makes available some open-source video models that can be deployed locally on its official GitHub repository.
Users can create and edit visual content using text, images, reference characters, reference videos, or existing videos; its applications include short videos, advertisements, product images for e-commerce, film previews, animations, digital avatars, and creative designs.
As of this verification, the cloud-based main line has been activated.Wan 2.7The current capabilities include text-to-video generation, image-to-video generation, reference-based video creation, video editing, as well as text-to-image generation, album creation, multiple image references, and image editing.
Terms such as “WanXiang 2.1” that were commonly used in previous descriptions, or simply “AI painting,” no longer suffice to accurately reflect the current state of this product. It is important to note that online creation platforms, Alibaba Cloud’s BaiLian API, and official open-source models represent three different ways of using this technology; the model versions, pricing models, data location requirements, and deployment needs cannot be treated as interchangeable.
Wan 2.7’s video generation capabilities
- Text-to-video:Enter descriptions of the scene, characters, actions, shots, sound, and style to create a video. Wan 2.7 supports both single-shot and multi-shot storytelling; it is also possible to use time periods in the prompts to describe different shots, making it suitable for advertising scripts, story segments, and concept previews.
- Tusheng Video:Convert portraits, product images, illustrations, or scene photos into dynamic content, controlling movements, camera movements, and environmental changes while preserving the main subject and composition. The current 2.7 image-to-video technology allows for the creation of audio-enabled videos directly.
- Reference student video:By using reference images or videos to maintain the appearance of characters, objects, clothing, and visual features, it is possible to create continuous narratives with one or multiple characters. This approach is suitable for IP-based short films, brand characters, and content that requires consistency across different shots; however, complex occlusions and sharp turns can still lead to changes in the appearance of those characters.
- Multiple lenses and long narratives:The model can automatically understand the structure of the scenes based on the given prompts, and it is also possible to specify explicitly the shot size, actions, and transitions for each time segment. Compared to generating each segment separately, this approach makes it easier to maintain the rhythm of the story as well as its semantic continuity.
- Audio video:Users can upload audio files, or they can ask the model to generate background music or sound effects based on the visuals. The supported audio formats are those that are commonly used; according to the API documentation, the length of the audio file should be between 2 and 30 seconds, and its size must not exceed 15MB. Any part of the audio that exceeds the duration specified for the video will be truncated.
- Video editing:Wan 2.7 VideoEdit allows for the modification of existing videos based on textual instructions and reference images; it is possible to replace clothing, objects, or visual elements. Editing services are charged according to the duration of the input and output video. For official projects, it is advisable to first use a short video to test the stability of the footage and the quality of edge blending.
Resolution, duration, and control parameters
The Wan 2.7 video generation API supports 720P and 1080P resolutions, with common aspect ratios including 16:9, 9:16, 1:1, 4:3, and 3:4. The duration of each generated video can be set between 2 and 15 seconds, with 5 seconds as the default value.
The 16:9 output for 720P is 1280×720, while that for 1080P is 1920×1080.
The model supports prompt texts in both Chinese and English of up to 5000 characters, reverse prompts, random seeds, intelligent prompt rewriting, and a toggle for the “AI-generated” watermark.
Video processing tasks typically take several minutes to complete, and the HTTP interfaces use an asynchronous task handling mechanism. Both the task results and their identifiers have a limited validity period; therefore, developers should periodically check for updates, download the files, and save them locally, rather than relying on the temporary result addresses as long-term storage solutions.
Fixing the random seed can improve reproducibility, but since the generation model is probabilistic, it is not possible to guarantee that the results will be exactly the same each time.
Wan 2.7 Image Generation and Editing
The main image options include Wan 2.7 Image and the higher-quality Wan 2.7 Image Pro. The flagship version supports text-to-image generation, text-to-multiple-image generation, image-to-multiple-image generation, image editing, multiple image references, and interactive editing; it also improves the rendering of Chinese and English text, ensures consistency in the elements depicted in the images, handles complex instructions better, and allows for multiple rounds of revisions.
The Standard version is cheaper and suitable for bulk product images, social media posts, concept sketches, and quick drafts; the Pro version is better suited for finished products that require high quality in terms of text, characters, or brand elements.
Multiple reference images can provide information on characters, products, clothing, poses, or styles, which is used to create a cohesive visual series. Interactive editing allows for further adjustments to the composition, background, materials, lighting, and individual elements based on existing images.
It is still not a traditional pixel-level photo editing software: small text, precise logos, hands, complex geometrical shapes, and sequences of images may exhibit errors; commercial materials require manual review and further formatting.
How to use the online version of Tongyi Wanshang
Regular users can select image or video creation tasks on online platforms, enter prompt words and upload reference materials, then set the frame size, duration, model, and resolution. It is recommended to first use a shorter duration or a standard model to determine the composition, actions, and main elements, and only then increase the resolution to produce the final version.
The prompt should specify the subject, environment, actions, camera angles, lighting, style, and sound; for videos with multiple shots, it is best to describe each shot according to the time period, so as to avoid including conflicting requirements all at once.
The online version usually offers a trial amount, which is credited in the form of points, membership benefits, or other incentives. The packages displayed may vary depending on the region, account type, as well as whether the service is accessed via web or mobile device; moreover, the free amount available can change as part of promotional activities. Therefore, the fixed numbers of points or membership prices mentioned in older tutorials should not be considered as permanent rates.
Before making a purchase, refer to the settlement page after logging in, and pay attention to whether there are charges for failed generation attempts, retries, high-definition quality, audio features, and editing functions.
Alibaba Cloud BaiLian API pricing
The following are the official original prices for the North China 2 (Beijing) region at the time of this verification, excluding any promotional discounts. Newly launched or newly released models usually come with a trial period valid for 90 days; the specific eligibility criteria can be found in the control panel.
| Package or version | Prices, quotas, and core benefits |
|---|---|
| Wan 2.7 Image Pro | 0.50 yuan per sheet, with a free quota of 50 sheets. |
| Pro | Wan2.7ImagePro: 0.50 yuan per image, with a free quota of 50 images. For the version designed for international use in Singapore, the standard version costs around 0.224826 yuan per image, while the Pro version costs around 0.562065 yuan per image. |
For the standard version of images intended for international use in Singapore, the cost is around 0.224826 yuan per image, while the Pro version costs around 0.562065 yuan per image. For 2.7 video, the cost is about 0.733924 yuan per second for 720P resolution, and around 1.100886 yuan per second for 1080P resolution.
The API keys, request addresses, and data locations vary between the Chinese region and the international region; therefore, during development it is necessary to ensure that the models, interface addresses, and keys are all in the same region.
Prices may change; for bulk orders, the console billing statements and the latest pricing table shall prevail.
Key points for API integration
Developers can connect using the HTTP interface or the DashScope SDK. Video tasks are typically carried out in the sequence of “creating a task – checking its status – downloading the results”, and checks should not be performed too frequently.
In a production environment, it is necessary to protect API Keys, and settings such as concurrency limits, budgets, timeout values, retry options for failures, and content filtering should be in place; the temporary video files generated should be saved promptly in an internal storage system.
When estimating costs, the number of trial and error attempts, the duration of input processing, high-definition resolution, storage requirements, bandwidth needs, and additional manual work are also taken into account.
Bailian also made Wan Agent Skills available; these skills enable AI agents that support such a mechanism to use Wan 2.7 for image generation and editing, and it provides examples of creating presentations by combining document models with the Wanxiang image model. Skills are essentially encapsulated API calls, and they require an Alibaba Cloud account as well as an API Key – they do not mean that the Wan 2.7 model can be used locally for free.
Open-source models and licenses
Tongyi Wanxiang is not an option where the entire product is made open source. The online platform and the Wan 2.7 cloud models are considered managed services.
What is made available on the official GitHub are specific versions of the models, inference code, and integration tools.
A representative open-source project at present is Wan2.2; its repository is licensed under the Apache 2.0 license. It offers 14B models for text-to-video, image-to-video, and hybrid generation, as well as a 5B model that can handle both text-to-video and image-to-video tasks. The 5B version supports 720P resolution at 24 frames per second, and it can run on high-end consumer-grade graphics cards.
The developers also released Wan2.2-Animate, which is used for character animation and role replacement; and Wan-Animate-2, which was released in August 2026 and includes inference scripts as well as weights. Wan-Animate-2 improves the fidelity of movements, helps maintain character identities, and enables control over the viewing perspective through text commands. It also offers a lightweight approach for real-time threshold optimization, and is licensed under the Apache 2.0 license.
Open-source deployment still requires a powerful GPU, sufficient space for downloading models, an environment for performing inference, as well as mechanisms for ensuring security; the fact that the code and weights are available does not mean that they possess the same capabilities as the latest version 2.7 available online, nor does it imply that all input materials automatically come with commercial licensing.
Tongyi Wanshang Usage Guide
Complete a basic task.
- Identify the audience, platform, format, duration, and the information that needs to be conveyed;
- Prepare scripts, shots, or reference materials that can be used in Tongyi Wanshang;
- Select the Wan 2.7 video generation feature to create a low-cost preview;
- Adjust the image, rhythm, subtitles, and sound using resolution, duration, and control parameters;
- Check each frame for characters, text, logos, lip movements, and factual accuracy;
- Export in the desired format after confirming the licensing for music, portraits, and materials;
Create reusable professional workflows
- Create scripts, shot lists, brand assets, and a list of elements that are prohibited from use;
- Unified parameters are established for the video generation capabilities, resolution, duration, and control settings of Wan 2.7, as well as for image generation, editing, and saving in Wan 2.7;
- First, use representative shots to test the model and the quota;
- Transfer the failed shots to manual editing or regenerate them;
- Uniformize subtitles, volume, colors, and end credits;
- Record the version and reviewer before publishing in batches;
Which users are it suitable for
- Short videos and content creators: producing narrative clips, animations, MVs, covers, and social media content.
- E-commerce and advertising teams: create product scene images, model presentations, advertising storyboards, and promotional videos.
- Film, television, and game teams: responsible for completing concept design, shot previews, character movements, and world-building demonstrations.
- Designers and brand teams: create a series of visual elements, multiple reference images, poster backgrounds, and creative proposals.
- Developers and enterprises: Generate images and videos in bulk via APIs or build industry-specific applications.
- Researchers and users with local setups: Testing official open-source models such as Wan2.2 and Animate.
Advantages
- Image and video capabilities are integrated within the same model family, covering generation, reference, and editing.
- Understanding Chinese prompts, rendering Chinese text, and connecting to domestic cloud services are all quite convenient.
- Wan 2.7 can create audio videos with a duration of 2 to 15 seconds, in resolution up to 1080P, and it supports multiple cameras.
- API billing is transparent: images are charged per piece, videos per second, and it can be deployed in both Chinese and international regions.
- The authorities also maintain the hosting services, SDK/Skills, as well as the Apache 2.0 open-source model ecosystem.
Usage restrictions and precautions
- AI-generated content may exhibit issues such as changes in the identity of the characters, errors in the depiction of hands, unreasonable physical movements, distorted text, discrepancies between audio and visual elements, and sudden changes in camera angle.
- Reference video footage can help improve consistency, but it cannot replace the manual review of actor, product, and brand materials.
- For medical, financial, news, educational, and advertising content, it is necessary to verify the facts; model-generated images cannot be used as genuine evidence.
- Before uploading real-person photos, voices, video clips, music, trademarks, and product images, it is necessary to obtain permission regarding the portraits, voices, copyrights, or brands involved.
- It is forbidden to create fraudulent, deceptive, unauthorized pornographic content or any other illegal material.
- When making content public, the applicable rules for labeling AI-generated content must be followed;
- When using APIs, businesses should also assess factors such as data location, data retention, access permissions, and the risk of key leakage.
Frequently Asked Questions
Is Tongyi Wanshang free?
The online version usually offers a trial quota, and Alibaba Cloud BaiLian also provides a limited-time free quota for some newly launched models; beyond that, charges are applied based on points, membership status, or API usage.
The actual rights and interests are as indicated on the page after logging in.
What is the latest version of Tongyi Wanshang?
The current main cloud platform is Wan 2.7, which covers image processing, text-to-video generation, image-to-video generation, reference-based video creation, and video editing. Wan2.2 is the official open-source series; the two have different purposes and deployment methods.
How long can Tongyi Wanshang generate videos?
The Wan 2.7 video generation API supports video lengths of 2 to 15 seconds per invocation, with options for 720P or 1080P resolution. Longer videos usually require breakdown into individual scenes, subsequent expansion, and post-production editing.
How much does the Tongyi Wanshang API cost?
In the Beijing area of China, the cost for Wan 2.7 standard images is 0.20 yuan per image, while the Pro version costs 0.50 yuan per image. For 2.7 video formats generated from text, images, or references, the cost is 0.60 yuan per second for 720P quality and 1 yuan per second for 1080P quality.
Prices may vary depending on the region and whether it is a promotional offer.
Is Tongyi Wanshang open-source?
Online platforms and the Wan 2.7 cloud model cannot be simply classified as open source. Specific versions such as Wan2.2, Wan2.2-Animate, and Wan-Animate-2 release their code and weights, and they are licensed under the Apache 2.0 license.
Guigong Network Security Registration No. 45132202000164