Resemble AI
An audio platform that offers voice cloning, generation, detection, and enterprise APIs
Tags:AI audio toolsWhat is Resemble AI?
Resemble AI is a platform that covers both AI-based voice generation and the security of generative media. Its capabilities for creation and development include text-to-speech, voice cloning, text-to-audio design, speech-to-speech conversion, speech transcription, audio editing enhancements, and real-time voice agents.
Security capabilities include Deepfake detection for audio, images, and videos, explanation of detection results, real-time analysis of meetings, identity registration and matching, C2PA signing, and imperceptible watermarks.
The product’s scope has evolved from being a simple AI voice-over tool to an integrated platform that handles generation, verification, and detection. Content teams can create brand-specific voices and multilingual narration, developers can use APIs to access real-time voice functionality, and security teams can send suspicious media for analysis while keeping track of the audit results.
Resemble Ultra text-to-speech
Resemble Ultra is the default text-to-speech model used by current hosting platforms for creating and upgrading audio content; it supports simultaneous generation as well as output via HTTP streams and WebSocket streams. The API will automatically select the appropriate model based on the voice_uuid, and developers should not try to pass resemble-ultra as a regular model parameter.
The older versions of Chatterbox, Chatterbox Turbo, Chatterbox Multilingual, as well as earlier TTS versions are listed in the official documentation as services that have been discontinued; those still using the voices associated with older models need to upgrade to Ultra in order to continue generating audio. It is necessary to recheck the quality of the sound, pronunciation, latency, and compatibility with various projects before making the migration.
Voice cloning
Voice cloning allows for the creation of reusable audio files based on clear recordings of around 10 seconds in length. The API enables training by uploading the entire audio file at once, or it is possible to upload the recordings in batches after the voice has been created, with notifications of completion of the training process sent via callbacks.
The Clone API currently requires a Business or higher plan.
Short samples allow for the rapid creation of audio, but environmental noise, music, reverb, multiple voices speaking at once, and excessive compression can reduce the degree of similarity. Any use of a real person’s voice requires their explicit consent, along with agreements regarding the intended use, duration, team members involved, deletion of the model, and the scope of any derived content.
Voice Design – Sound design
Voice Design does not require actual voice recordings. Users describe the age, accent, tone, style, and intended use of the character in text form, and the system generates three candidate voices. After listening to them, the selected candidate can be converted into a voice_uuid that can be used for TTS.
Sound design is suitable for fictional characters, prototypes, and projects where it’s not possible to record real people’s voices; however, it cannot guarantee that the chosen voice will be completely unrelated to any real individuals. Brands and game projects should still undergo similarity checks, and the final candidates as well as the parameters used for generating the voices should be recorded.
Speech-to-Speech voice conversion
Voice-to-voice conversion transforms the rhythm, tone, and delivery of the input audio into the target voice; it is suitable for use in real-time character creation, gaming, voice-over revisions, and voice agents. The current document lists Core STS v2 and v1, with version v2 offering improved pitch tracking and support for 48kHz frequencies. Training data typically requires more than 10 minutes in length.
STS is better at preserving the performance quality compared to plain text generation, but it does not automatically handle issues such as input noise, misstatements, and copyright problems. Both the speaker providing the input and the target voice require authorization; in live streaming scenarios, abuse detection and the ability to terminate transmissions manually are also necessary.
Synchronous, HTTP, and WebSocket streaming APIs
The synchronous interface returns the entire audio file at once, making it suitable for short sentences and offline tasks; the HTTP streaming interface generates and returns WAV data as it is created, with timestamps for syllables and phonemes included in the file header.
WebSocket maintains a persistent connection, enabling continuous delivery of audio frames with lower latency.
WebSocket is currently intended for Business and higher-tier plans; by default it allows 20 concurrent sessions per cluster as well as 20 parallel connections per API Key. Concurrency does not refer to the number of requests per second, and developers still need to handle errors related to capacity limits, retry mechanisms, disconnected connections, and the audio_end message.
Real-time voice Agent
Real-Time Agents enable the combination of Resemble voices with prompts, knowledge bases, tools, Webhooks, and phone numbers, for use in customer service, scheduling, sales, game NPCs, and interactive training. The system supports Agent APIs, tool calls, knowledge bases, and phone connections.
The naturalness of voice in real-time proxy services cannot replace the reliability of factual information. When it comes to accounts, payments, medical matters, legal issues, or identity verification, it is necessary to restrict the permissions of such tools, arrange for human intervention, confirm sensitive operations, and keep auditable logs.
Audio editing, enhancement, and transcription
Audio Edit allows for asynchronous editing of audio files that contain custom, non-market-generated sounds; Audio Enhancement is used to reduce noise and improve the clarity of speech; Speech to Text enables the creation and management of transcription tasks. These tools are useful for cleaning up recordings, creating subtitles, and redoing specific parts of the narration.
Enhancements can alter the details of the tone, and transcription services may misinterpret names, numbers, and technical terms. For critical records, the original files should be retained and manually verified; the processed results cannot be considered accurate evidence.
Deepfake multi-modal detection
Resemble Detect is designed for audio, images, and videos; it analyzes various signs such as TTS, voice cloning, face replacement, and generative images, and can provide confidence scores along with explanatory information either when files are uploaded, through APIs, or in real-time processes. Currently, DETECT-3B Omni is regarded as the enterprise-level detection model capable of handling all three types of media.
The detection results are based on probabilistic assessments; compression, editing, re-recording, new models, and adversarial manipulation can all lead to incorrect judgments. In high-risk scenarios, a user should not be rejected automatically based solely on a score – rather, factors such as the source of the request, identity verification, manual review, and appeal processes should be taken into account.
Resemble Intelligence and Meetings
Intelligence provides more in-depth analysis of the detection results, while Meetings enables real-time monitoring of suspicious media during meetings. Team offers meeting analysis capabilities, and Business adds organization-level calendar integration as well as security notifications for meetings.
Meeting monitoring involves the audio, video, and metadata of participants. Organizations should inform attendees in advance, set limits on the duration for which recordings are kept, and check the local regulations regarding recording, privacy, and employee monitoring.
Identity search
The Identity API allows for the registration of individual or brand identities, the association of reference media, and the checking of whether uploaded content matches the registered entities; it is used to identify counterfeit or unauthorized audio or brand materials. It can serve as a mechanism for KYC processes, content moderation, and the protection of intellectual property rights.
Similar identities alone do not constitute proof of fraud, and mismatching can also affect legitimate users. Production systems require threshold calibration, secondary authentication, and manual verification; identity detection cannot be relied upon as the sole basis for biometric identification.
PerTH watermark and C2PA
Resemble Watermarker can add imperceptible markers for the encoding and decoding of media, while the PerTH project focuses on audio neural watermarks, with the goal of enabling the identification of the source even after compression and common editing processes. The platform also supports the verification of SynthID and the signing of C2PA records.
Watermarks are useful for verifying the origin of the marked content, but the absence of a watermark does not necessarily mean that the file is authentic, as third-party tools might not have included compatible markers. The best practice is to use watermarks, content certificates, Deepfake detection, and publication records together.
Flex free pay-as-you-go plan
The subscription fee for Flex is $0; no credit card is required. It includes 1 team seat, as well as access to the platform and APIs, audio, image, and video detection functions, interpretation of detection results, identity verification, real-time monitoring of meetings, SynthID checks, and C2PA signing. Users can top up Credits in advance and are billed based on the actual amount of processing done; there is no minimum commitment required, and Flex Credits do not expire.
$
Price and version comparison
| Package or version | Prices, quotas, and core benefits |
|---|---|
| Team price | Team costs $350 per month; when paid annually, the monthly cost is $280, resulting in a total of $3360 for the whole year. This amount represents a savings of $840 compared to paying on a monthly basis. Team includes 5 seats, all Flex features, lower rates for lower usage levels of certain services, Meeting Intelligence, as well as the ability to upload large files in bulk. It is suitable for medium-sized teams that want to incorporate detection capabilities into their operations, but it does not include SSO, an organization-wide meeting calendar, or enterprise-level SLAs. The annual payment discount requires a one-time commitment for the entire year. |
| Business price | The Business plan costs $1,000 per month; when paid annually, the monthly cost is $800, resulting in a total of $9,600 for the whole year. This amounts to a savings of $2,400 compared to paying on a monthly basis. It includes 20 seats, configuration options that can be adjusted according to needs, organization-level calendar integration, security notifications for meetings, and SSO. The Business plan is also the current tier that provides access to voice cloning APIs and WebSocket-based streaming audio. If the monthly expenditure exceeds $1,000, or if higher concurrency or on-premises deployment is required, it is recommended to contact sales to obtain a quote for bulk pricing. |
| Detection and secure handling costs | In Flex, the cost for audio and image detection is $0.035 per second, while video detection costs $0.07 per second; Intelligence services cost $0.025 per second. In Team and Business plans, audio and image detection are both priced at $0.015 per second, video detection is $0.03 per second, and Intelligence services are also $0.015 per second. Identity search costs $0.0005 per invocation, and Watermarker encoding costs $0.0005 per operation; decoding costs $0.0002 per operation. These three services have the same pricing in Flex, Team, and Business plans. The pricing shown per second is based on the official price list; the actual billing unit and calculation are determined according to the settings in the console. |
| How are the prices for hosted voice generation? | The current pricing page displays information on security solutions for generative media as well as the processing fees; it does not list in a single table the individual pricing rates for Resemble Ultra TTS, cloning, STS, and Agent. Therefore, the pricing structure should not rely on the per-second generation costs from older blogs or previous packages. Users who need voice generation services should register for Flex to view real-time rates, or ask sales representatives to provide quotes based on language, duration, concurrency, and deployment method. When budgeting, it is necessary to account for generation, streaming servers, cloning, Agent services, as well as security checks separately. |
Team price
The Team plan costs $350 per month; when paid annually, the cost is $280 per month, for a total of $3,360 throughout the year. This represents a savings of $840 compared to paying on a monthly basis for the whole year.
It includes 5 seats, all Flex features, lower pricing for lower usage volumes of certain services, Meeting Intelligence, as well as the ability to upload large files and batches of files.
Team is suitable for medium-sized teams that want to integrate testing into their production processes; it does not include SSO, an organization-wide meeting calendar, or enterprise SLAs. The annual payment discount requires a one-time commitment for the entire year.
Business price
The Business plan costs $1,000 per month; when paid annually, the cost is $800 per month, resulting in a total of $9,600 for the whole year. This represents a savings of $2,400 compared to paying on a monthly basis throughout the year.
It includes 20 seats, a configuration that can be adjusted according to deployment needs, organization-level calendar integration, security notifications for meetings, and SSO.
Business is also the current tier for access to voice cloning APIs and WebSocket-based streaming audio. If the monthly expenditure exceeds $1,000, or if higher concurrency levels or on-premises deployment are required, it is recommended to contact sales to obtain a quote for enterprise-use rates.
Enterprise business solution
Enterprise-customized quotes are available, offering capacity discounts, enterprise SLAs, custom model training, on-premises deployment, dedicated support, SOC 2 documentation, and custom seats. High concurrency, data isolation, as well as private cloud or air-gapped environments usually require an enterprise contract.
Local deployment is not included as part of an open subscription; when making a purchase, it is necessary to obtain written confirmation regarding the model, detection methods, hardware, update frequency, log retention period, support response times, and license scope.
Detection and secure handling costs
In Flex, the cost for audio and image detection is $0.035 per second, while video detection costs $0.07 per second; Intelligence services cost $0.025 per second. In Team and Business plans, audio and image detection are both priced at $0.015 per second, video detection is $0.03 per second, and Intelligence services cost $0.015 per second.
The Identity search costs $0.0005 per call, while Watermarker encoding also costs $0.0005 per call; decoding, on the other hand, costs $0.0002 per call. These three services have the same price in the Flex, Team, and Business plans. The pricing shown as “per second” for images is based on the official pricing table; the actual billing unit and conversion factors depend on what is indicated in the console.
How are the prices for hosted voice generation?
The current public pricing page displays the security solutions for generative media as well as the processing fees; it does not list in a single table the individual pricing rates for Resemble Ultra TTS, cloning, STS, and Agent. Therefore, the pricing should not be based on the per-second generation costs from old blogs or previous subscription plans.
Users who need voice generation should register for Flex in order to view the current rates, or ask sales representatives to provide a quote based on language, duration, concurrency levels, and deployment method. When budgeting, it is necessary to account for generation services, streaming hosts, cloning, agent phones, and security checks separately.
Chatterbox open-source model
Resemble AI has made the Chatterbox speech model series available on GitHub. The current repository includes Chatterbox Turbo, which features 350 million parameters and is designed for low latency in handling English speech as well as for various performance tasks.
Nano, which comes with 110M parameters and can run on CPUs as well as on devices with limited resources; Multilingual V3, which has 500M parameters and supports 23 languages including Chinese, along with various single-language fine-tuning packages.
The Chatterbox code and models are licensed under the MIT license; they can be used in personal, research, and commercial projects, as well as for self-hosting in closed-source products. However, it is necessary to retain the license notice and comply with applicable laws and consent requirements. Resemble Ultra, a hosting platform, and the open-source version of Chatterbox are two separate product lines, and the official documentation currently specifies that the older hosted version of Chatterbox is no longer in service.
PerTH and the Resemble Enhance open-source projects
The official GitHub also makes available the PerTH audio watermarking tool under the MIT license, as well as Resemble Enhance for noise reduction and enhancement. Developers can test watermark extraction and audio processing locally, but this does not mean that Detect, Identity, Meetings, and the hosting platform as a whole are open source.
Open-source examples and some older SDKs may no longer be maintained; for instance, the official Go SDK has been marked as deprecated. New projects should rely on the current API documentation and Client Libraries pages, and they should lock in a specific version before upgrading in order to carry out regression tests.
Resemble AI Usage Guide
Complete a basic task.
- Register for Resemble AI and create an API Key intended solely for the testing environment;
- Select a model based on input type, context, quality, speed, and price;
- First, use Resemble Ultra’s text-to-speech function to submit the minimal request and check the returned structure;
- Use voice cloning to test streaming output, parameters, and abnormal response;
- Record Tokens, number of calls, latency, error rate, and cost per call;
- Move the key to the server-side key manager before integrating it into the actual application;
Create reusable professional workflows
- Different keys and quotas are used for development, testing, and production environments;
- A representative evaluation set was created using Resemble Ultra for text-to-speech, voice cloning, and Voice Design for sound creation;
- Set timeout, concurrency, retry, throttling, and budget limits;
- Perform checks on the output regarding facts, security, format, and sensitive information;
- Monitor changes in model version, price, latency, and failure rate;
- Prepare plans for downgrading the model, implementing circuit breaking, and taking manual control;
Which users is it suitable for?
- Content teams that need a brand voice, game characters, and multilingual narration;
- Developers who create low-latency voice agents, customer service systems, and interactive applications;
- The security and risk management team responsible for verifying the authenticity of audio, images, and videos;
- Organizations that protect actors, brands, and executives from being impersonated;
- A research and engineering team dedicated to self-hosting MIT speech models.
Product advantages
- Integrate speech generation, identity verification, watermarking, and Deepfake detection in a single platform;
- Resemble Ultra supports synchronization as well as two types of streaming TTS;
- It supports fast voice cloning, Voice Design, and text-to-speech conversion;
- It checks audio, images, and videos and provides explanatory information;
- Flex does not have a monthly subscription fee, and its Credits do not expire.
- Chatterbox, PerTH, and Enhance offer self-hostable open-source options.
Restrictions and Precautions
- The focus of current public pricing has shifted to security products; the exact price for hosted voice generation needs to be checked in the console.
- VoiceCloningAPI, WebSocket, and certain production capabilities require a Business plan or higher; beginners cannot assume that all voice interfaces are available at no cost through Flex.
- Deepfake detection can lead to false positives and false negatives, and the absence of a watermark does not prove that the content is authentic.
- Cloning and STS require a sound licensing approval;
- The MIT license of the open-source Chatterbox does not exempt one from obligations related to privacy, portraits, rights to one’s voice identity, and anti-spoofing measures.
Frequently Asked Questions
Is Resemble AI free?
There is a Flex option costing 0 dollars per month; no subscription fee is required, and payment is made on a usage basis through pre-loaded Credits. These Credits do not expire, but actual generation, verification, and watermark application still incur costs.
How much does Resemble AI cost?
The Team plan costs $350 per month, or $280 per month if paid annually; the Business plan costs $1000 per month, or $800 per month if paid annually.
Custom Enterprise version. Flex comes without a monthly fee.
Is it possible to clone one’s own voice?
Yes, a clear recording of about 10 seconds is sufficient to create a quick clone; however, the API currently requires a Business or higher plan, and consent from the speaker is necessary.
Can it detect Deepfakes in videos and images?
Yes. Detect currently covers audio, images, and video, and can be combined with Intelligence, Identity, meeting monitoring, and watermarking.
The results still need to be manually reviewed based on risk.
Is Resemble AI open source?
The hosting platform and the Ultra model are not fully open source; the official Chatterbox speech model, PerTH watermarking technology, and Resemble Enhance are available as open-source projects, with Chatterbox and PerTH being licensed under the MIT license.
Guigong Network Security Registration No. 45132202000164