Tavus
Free value-added services
Comprehensive List of AI Tools AI video tools

Tavus

An AI platform that offers programmable digital avatars and personalized video capabilities

Tags:

What is Tavus?

Tavus is a real-time AI digital human platform designed for developers, product teams, and enterprises. It first gained attention for its capability to generate personalized digital avatar videos; its core product nowadays is the Conversational Video Interface, or CVI, which is used to create real-time video AI Agents that can see, hear, speak, remember, and utilize various tools.

CVI integrates digital human rendering, speech recognition, LLMs, text-to-speech, and WebRTC connections into a single API along with hosting components. Developers can embed 1080p digital human conversations in web pages, applications, Zoom, or Google Meet without having to purchase separate voice and video services.

The difference between Face and PAL

Face defines the appearance and voice on the screen, that is, the digital avatar’s face, voice, expressions, and the way it is presented. PAL defines behaviors, knowledge, prompts, goals, constraints, tools, and pipeline configurations.

A PAL can be associated with a default Face, or another Face can be selected when creating a session.

Separating the design of identity and behavior facilitates reusability: the same customer service logic can be applied to different characters, and the same digital avatar can take on various roles such as sales, training, or consulting.

Phoenix digital human rendering model

Phoenix is responsible for generating in real time facial expressions, micro-expressions, head movements, and emotional changes. The current Phoenix-4 version allows for dynamic adjustments of expressions based on context, tone, and cues from the conversation, enabling a smooth transition between speaking and listening modes.

It focuses on Face presentation and does not cover the entire range of business-related knowledge. A realistic appearance does not guarantee reliable answers; the quality of the information is still determined by the PAL configuration, the LLM, the knowledge base, and the tools used.

Raven Perception Model

Raven analyzes users’ videos, audio, body language, screen sharing content, and environmental cues to identify emotions, intentions, tones, as well as specific objects or actions. It can trigger tools when certain behaviors or audio events are detected.

Visual and emotional inference can lead to errors, and therefore should not be used alone for medical diagnosis, recruitment screening, credit decisions, or law enforcement. When sensitive inferences are involved, it is necessary to provide clear information, obtain user consent, and carry out manual verification.

Sparrow round-robin scheduling model

Sparrow handles pauses, rhythm, turns in dialogue, and interruptions. Sparrow-1 is the officially recommended version; it is faster and more natural compared to older versions as well as those that rely solely on time-based rules.

Developers can set the Turn-taking Patience level to low, medium, or high, and they can also control the sensitivity with which interruptions to PAL are detected. Responses to customer inquiries should be relatively quick, while more patience is needed in situations involving interviews and consultations.

Create custom Faces from videos

A high-quality approach involves uploading a training video of about 1 minute long, with around 30 seconds dedicated to speaking and 30 seconds to listening. This method allows for the capture of appearance, voice, expressions, and speaking habits, and its quality and level of control are generally better than those achieved with a single image.

Training videos need to be clear, stable, mainly positive in tone, and meet technical requirements. Users must possess the necessary rights to their portraits, voices, and videos.

Create a Face from an image

Users can also upload individual images of people or characters; the system will then generate training data automatically, allowing them to select from existing voices. This method enables faster deployment, but the accuracy in terms of identity, actions, and expressions is lower compared to training using videos.

The Face image is suitable for prototypes and virtual characters, and should not be used for unauthorized cloning of real people.

Over 100 stock faces

Tavus offers over 100 Stock Faces recorded by real actors, suitable for rapid prototyping, demonstrations, and teams that do not have custom assets. When using these Stock resources, it is still necessary to comply with the platform’s rules regarding permissible uses and advertising disclosures.

PAL Maker: code-free construction

PAL Maker offers a guided interface that allows non-developers to configure features such as face, voice, behavior, knowledge, goals, and tools, enabling them to start hosted conversations right away. The API, on the other hand, is suitable for integration into official products and for automated deployment.

Knowledge base and RAG

An optimized Knowledge Base can load corporate documents, providing a search context for answering queries. Documents should be organized by version, outdated information removed, and access permissions established.

RAG can reduce hallucinations, but it cannot guarantee that the correct paragraph will be retrieved every time.

Dynamic memory

Dynamic Memories enables PAL to remember users’ preferences and background information in subsequent sessions, thereby enhancing continuity. This memory constitutes personal data; it is necessary to provide options for informing users about such data, as well as for deleting, modifying it, and specifying its retention period. Passwords, payment details, or any other unnecessary sensitive information should not be stored.

Goals, safeguards, and tool invocation

Objectives define the tasks that need to be completed during a session, while Guardrails set limits on the actions that are not allowed. Function Calling and Tools are used for querying orders, scheduling appointments, or updating CRM data. The parameters of these tools must be verified by the server, and users are required to confirm high-risk operations.

Emotion control

Phoenix-4 supports both explicit and implicit emotion control. Explicit settings allow the direction of expression to be specified, while the implicit mode adjusts according to the context of the conversation.

Marketing and entertainment can make use of more pronounced emotions, while healthcare, finance, and complaint handling should avoid manipulative expressions.

Voice, pronunciation, and noise reduction

CVI includes TTS, 24kHz audio, a pronunciation dictionary, advanced noise reduction, and speaker isolation. The Custom Pronunciation Dictionary allows for the correction of brand names, personal names, and technical terms.

Text confirmation and subtitles should still be provided in noisy environments.

Video and pure audio pipelines

Developers can choose between full video or audio-only mode. Audio-only mode typically results in lower costs and less bandwidth usage, making it suitable for phone calls or background agents;

Video mode offers expressions and visual feedback, but it requires a stronger internet connection, better equipment, and greater attention to privacy.

Comparison of package options and pricing methods

PackagePrice methodFace and usage amountSuitable scenarios
FreeFree development trial available; the specific number of minutes and concurrent users are as specified in the console.Stock resources and basic CVI can be used; custom Faces cannot be created.Prototyping, API testing, and functional evaluation
StarterThe subscription or usage price is displayed on the dynamic console.Enable custom Faces and higher session limitsInitial products and small-scale official deployments
GrowthCharged at higher usage and capacity levelsCustom Faces, Higher Concurrency, Minutes and Production CapacityApplications for growth and multi-scenario deployment
EnterpriseCustom quoteCustomize Face, capacity, compliance, support, and contract termsHigh concurrency, regulated environments, and large enterprise projects

The price information on the official website does not always show specific amounts in the public text; however, it is confirmed that APIs and end-to-end dialogue stacks are available for all tiers, from Free to Enterprise. Users can customize various parameters such as faces, number of minutes, concurrent connections, and recording storage options depending on the subscription level. Before making a purchase, one should log in to the developer console to check the current pricing, including the number of minutes allowed, any additional fees, concurrent connections, and storage options for recordings.

What is included in the cost?

The official website states that the CVI price covers all the necessary components such as LLM, TTS, and WebRTC; it also includes 1080p video, 24kHz audio, perception capabilities, RAG, memory functions, security measures, various tools, as well as transcription and recording options. The total cost for businesses may also include integration services, data preparation, compliance requirements, support for deployment, and charges for any additional usage.

Generation of old videos and current CVI

Tavus still possesses Video API capabilities as well as the ability to create digital avatars; the information provided regarding the duration of a single video indicated that it could be up to 5 minutes long. The focus of this product has now shifted to real-time CVI and PAL, and it should not be described merely as a tool for creating batched, personalized recorded videos.

Zoom and Google Meet

PAL can join sessions via Zoom and Google Meet, and is suitable for interviews, training, sales, and support tasks. Before automatically joining a meeting, it is necessary to inform the participants that it is an AI, as well as to clarify whether recording, video recording, transcription, or saving of the session will take place.

Integration of LiveKit with Pipecat

Tavus offers integration options with LiveKit and Pipecat, enabling the incorporation of face rendering into existing real-time agents and media infrastructure. The team is responsible for managing the session lifecycle, network reconnections, microphone permissions, and server keys.

Fast embedding process

  1. Save TAVUS_API_KEY on the server side.
  2. Choose Stock or customize PAL and Face.
  3. The server calls the interface to create a Conversation.
  4. Embed the returned Conversation URL in an iframe or application.
  5. Users end the session voluntarily when leaving to avoid continuous billing.

Automatic testing can make use of test_mode to create sessions that do not actually involve PAL’s participation and that are terminated immediately, thereby avoiding the generation of costs associated with normal sessions for testing purposes.

Conversation API

Developers can create, terminate, and query sessions, as well as access transcripts and recordings. Once a session is created, billing begins and concurrent resources are used, so it is not sufficient to rely on closing the browser to free up those resources.

OpenAPI and server-side security

The official source provides a complete OpenAPI specification. API Keys must not be included in browser code, mobile applications, or public repositories;

The frontend can only call its own backend, which is responsible for creating controlled sessions.

The file URLs used for Face training should be short-term signed addresses; keys and permanently accessible sensitive materials must not be recorded in the logs.

GitHub and example projects

Tavus Engineering provides examples of CVI, as well as demonstration projects related to interviews and Santa, on its official GitHub repository; these resources help developers understand real-time interfaces. Among the search results, there are also third-party Python, Ruby, and C# SDKs, and it is important to distinguish between those maintained by the official team and those maintained by the community before using them.

Open-source status

Tavus’s Phoenix, Raven, Sparrow models, Face training, hosted CVI, and commercial APIs are not open-source products. The fact that the official example projects have open-source code does not mean that the core models or server-side components can be deployed privately.

Data, recording, and compliance

CVI can save transcripts and recordings of conversations, and it can handle video, audio, facial features, emotions, as well as knowledge bases. Companies must determine matters such as consent, legal basis, access rights, data retention and deletion, data location, and the rules governing data processing.

Users can request that data be deleted, but enterprise applications must also carry out synchronous deletion in their own data systems.

Tavus usage guide

Create reusable professional workflows

  1. Create scripts, shot lists, brand assets, and a list of elements that are prohibited from use;
  2. Uniform parameters are saved for the differences between Face and PAL, as well as for the Phoenix digital human rendering model and the Raven perception model.
  3. First, use representative shots to test the model and the quota;
  4. Transfer the failed shots to manual editing or regenerate them;
  5. Uniformize subtitles, volume, colors, and end credits;
  6. Record the version and reviewer before publishing in batches;

Which users is it suitable for?

  • Developers who create real-time digital human customer service and sales agents;
  • Dialogue products that require visual, auditory, and emotional perception;
  • A team responsible for creating training and interactive courses for digital avatars;
  • Companies that have integrated AI Agents into Zoom or Google Meet;
  • Applications that require RAG, memory, safeguards, and tool calls;
  • Teams that use LiveKit or Pipecat to build real-time media pipelines.

Main advantages

  • End-to-end includes LLM, TTS, WebRTC, and digital avatar rendering;
  • Phoenix, Raven, and Sparrow are responsible for appearance, perception, and rounds respectively;
  • It supports creating custom Faces from videos or individual images;
  • PAL integrates knowledge, memory, goals, safeguards, and tools;
  • Offers 1080p video, low latency, and extensive real-time control options;
  • Open API, PAL Maker, and integration with conference and real-time frameworks.

Restrictions and Precautions

  • Real-time video involves higher costs, greater bandwidth requirements, and more privacy risks compared to text or voice agents;
  • Visual emotion inference can lead to errors, and custom Faces require authorization.
  • If a session is not terminated properly, charging may continue;
  • The core platform is not open-source, and the specific price must be checked by logging in to the console.
  • LLM responses, RAG retrieval, and tool execution can still encounter errors, and high-risk operations must be manually verified;

Frequently Asked Questions

Is Tavus free?

There is a free development experience available, with access to basic CVI resources and stock assets; custom faces are only available for the Starter, Growth, and Enterprise plans.

How much material is needed to create one’s own digital avatar?

A high-quality version can be created using a video of about 1 minute, with around 30 seconds dedicated to speaking and 30 seconds to listening; it is also possible to use a single image for a quick creation, though the quality will be lower.

What is the difference between Face and PAL?

Face is responsible for appearance and voice, while PAL is responsible for behavior, knowledge, goals, memory, safeguards, and tools.

Does Tavus provide an API?

Complete API and OpenAPI specifications are provided, enabling management of Faces, PALs, Conversations, Videos, Tools, and Documents.

Can Tavus join Zoom?

Integration with Zoom and Google Meet is supported; when using these platforms, it is necessary to inform the participants about the identity of the AI and the recording policies.

Is Tavus open source?

The core model and the managed CVI are not open source; examples and demonstration projects are available on the official GitHub repository.

©️Copyright notice: Unless otherwise specified, all articles on this site are copyrighted bySharing of AI toolsAll content on this site is original; without permission, no individual, media outlet, website, or organization may reproduce, copy, or otherwise distribute it, nor may they create mirrors of it on servers that are not owned by this site. Otherwise, we reserve the right to take legal action against such parties in accordance with the law.

Tools similar to Tavus