Cartesia
Free value-added services
Comprehensive List of AI Tools AI audio tools

Cartesia

Low-latency generative audio platform for real-time voice agents

Tags:

What is Cartesia?

Cartesia is a real-time voice AI platform designed for developers and enterprises; its main products include Sonic text-to-speech, Ink speech recognition, and Line voice bots. The platform emphasizes low-latency streaming processing, making it suitable for applications that require natural dialogue and rapid response to interruptions.

Developers can utilize voice capabilities through Playground, APIs, Python or JavaScript SDKs, as well as MCP. In addition to text-to-speech and speech transcription, Cartesia offers services such as voice cloning, voice conversion, localization, pronunciation dictionaries, and enterprise deployment solutions.

Main functions of Cartesia

  • Sonic TTS:Convert text into natural speech in real time and return the audio as a stream.
  • Ink STT:It transcribes speech for real-time conversations, handling noise, accents, phone quality, and structured data.
  • Semantic endpoint detection:Determine when a round ends based on the meaning of the spoken words, thereby reducing false detections triggered solely by silence.
  • Instant voice cloning:Sounds can be created quickly using a clear sample for about 10 seconds; this feature is available starting with the Pro package.
  • Professional voice cloning:The model can be fine-tuned through single-person recording over 30 minutes; the Startup package is available for this purpose.
  • Voice Changer:Maintain the original audio tone and rhythm, while replacing the speaking voice with another timbre.
  • Voice localization:Adapt existing voices to other languages or dialects.
  • Pronunciation dictionary:Customize the pronunciation and phoneme processing for brands, names, drug names, and technical terms.
  • Emotions and control:Adjust speed, volume, emotion, pauses, and expression according to the model’s capabilities.
  • Line Agents:A real-time intelligent agent is created by combining voice input, voice output, models, and telephone functions.
  • APIs and SDKs:It provides interfaces for Python, JavaScript, and other HTTP/WebSocket applications.
  • MCP:Allow compatible AI clients to directly execute TTS, STT, sound management, and usage queries.

Comparison of Sonic, Ink, and Line

ProductsMain inputsMain outputSuitable scenarios
SonicText, audio, and expression parametersReal-time or file-based voiceNarration, customer service, game characters, and voice interfaces
InkReal-time or recorded audioGradual transition to the final transcription, turn eventsVoice Agent, phone, meetings, and transcription
LineUser voice and Agent workflowComplete voice conversations and telephone interactionsCustomer service, booking, filtering, and automated outbound calls

Comparison of prices for Cartesia packages

The following are the monthly prices in dollars as displayed on the official pricing page; all plans offer unlimited workspace seats. Credits and Agent minutes are charged separately, while taxes, prepaid balances, and corporate discounts are determined according to the settlement page on the console.

PackagePriceMonthly creditsPrepaid Agent Credit LimitCore rights and interests
Free$20,000$TTS, STT, and basic trials
Pro$100,000$Commercial use license and instant voice cloning
Startup$1,250,000$Professional sound cloning and organization
Scale$8,000,000$Priority is given to supporting higher concurrency.
EnterpriseCustomizationCustomizationCustomizationVolume discounts, SSO, DPA, BAA, and dedicated support

Comparison of package capacity and concurrency

PackageApproximately TTS minutesApproximate STT durationTTS concurrencySTT/Agent concurrency
FreeAbout 27 minutesAbout 1 hour and 51 minutes28
ProAbout 133 minutesAbout 9 hours and 16 minutes312
StartupAbout 1667 minutesApproximately 115 hours and 44 minutes520
ScaleAbout 10,667 minutesAbout 740 hours and 44 minutes1560
EnterpriseCustomizationCustomizationCustomizationCustomization

The minutes listed above are approximate values calculated by the authorities using the current standard model; they are not fixed commitments. Text preprocessing, the model used, audio quality, silence periods, endpoints, professional cloning, and various audio tasks can all affect the actual time required.

Credits and additional billing rules

  • Standard TTS:Generally, each character costs about 1 credit, with preprocessing possibly causing slight variations.
  • Professional cloned TTS:Approximately 1.5 credits per character, which is 50% higher than the standard voice rate.
  • Instant clone TTS:Calculated at the standard TTS rate, without any additional factor for professional cloning.
  • Professional cloning training:Training successfully once costs 1,000,000 credits.
  • Retraining:Changing the base model or training data requires paying for training credits again.
  • Voice Changer:Charging is based on 15 credits per second of input audio.
  • Voice localization:225 credits are consumed at once.
  • Line Agent:The cost of a call is 0.06 dollars per minute; using a number from the platform adds an additional 0.014 dollars per minute.
  • STT:Charging is based on the model, batch or real-time endpoint, as well as the audio duration; silence is also taken into account.

Tutorial on text-to-speech integration

  1. Create key:Generate an API Key in Playground, and separate development, testing, and production environments.
  2. Select model:Give priority to using the current stable version of Sonic; do not fix old model IDs permanently.
  3. Select sound:Listen to the candidates in the sound library by language, role, scene, and speaking speed.
  4. Format text:Add punctuation, dates, times, pauses, and custom pronunciations.
  5. Use the flow method:Real-time conversations use WebSocket or streaming output to play the first audio segment as soon as possible.
  6. Handling continuity:When breaking long content into sections, use the continuous generation feature to reduce changes in the tone and rhythm of the paragraphs.
  7. Monitoring quality:Record latency, failure rate, credits, and user interruptions to gradually optimize the parameters.

Ink Real-time Transcription Integration Tutorial

  1. Verify that the audio encoding, sample rate, channels, and transmission protocol meet the requirements of a real-time interface.
  2. Connect to WebSocket before sending audio, and design a state machine for handling disconnections, reconnections, and timeouts.
  3. Distinguish between intermediate transcription and final transcription; do not write unstable text directly into business records.
  4. Listen to events such as turn.start, turn.end, and eager_end to coordinate the large model with the TTS response.
  5. Create test sets for phone numbers, dates, email addresses, currencies, product names, and industry terms.
  6. A genuine assessment is carried out under conditions of phone noise, background noise, different accents, and silence.
  7. Record the audio duration, latency, error rate, and cost per request, and set usage alerts.

Voice cloning tutorial

  1. Confirm agreement:Only the individual’s own voice can be cloned, or explicit written permission from the owner of the voice rights is required.
  2. Select type:First, use a sample for about 10 seconds to test instant cloning; if that is not sufficient, then consider professional cloning.
  3. Prepare materials:Professional cloning requires at least 30 minutes for a high-quality single-person recording; more than 2 hours is generally preferred.
  4. Uniform style:The training recordings should maintain the desired speaking pace, volume, environment, and style of expression.
  5. Submit training:Create professional sounds using the Dashboard or datasets along with the fine-tuning API.
  6. Acceptance sound:Test names, numbers, emotions, long sentences, low-frequency words, and the target language.
  7. Release control:Limit access, record the purpose of use, and indicate in publicly available content that it was generated by AI as required.

Cost control for Line voice bots

  • Set a maximum duration, a silent timeout period, and the maximum number of retry attempts for each call;
  • First, test on a web page or in a sandbox environment, and then enable paid phone numbers and outbound calls.
  • Short system prompts and tool-based responses are used to reduce the additional costs associated with large models;
  • Keep separate records for STT, LLM, TTS, and phone charges, rather than just looking at the total bill;
  • Use caching or pre-generated speech for repeated questions and answers, but avoid playing outdated information;
  • Detect robots, voice mailboxes, and invalid numbers to end unnecessary calls as early as possible;
  • Manual confirmation is added for high-risk operations, to prevent voice agents from carrying out irreversible transactions directly.

Who is Cartesia suitable for?

  • Voice Agent developer:Create a customer service system, appointment functionality, sales filtering tools, and an interactive assistant.
  • Call center:Use real-time transcription, synthesis, and phone capabilities to handle large numbers of conversations.
  • Games and character apps:Generate low-latency, controllable, and expressive character voices.
  • Education and accessible products:It offers reading aloud, dictation, voice navigation, and multilingual support.
  • Media and creators:Create voiceovers, localized dubbing, and authorized brand voices.
  • Enterprise development team:Access the voice infrastructure through APIs, SDKs, MCP, and enterprise deployments.

Product advantages

  • Sonic, Ink, and Line cover voice input, output, and full intelligent agents;
  • Optimized for real-time conversations: reduced latency of the first packet, stream-based processing, and turn detection;
  • It supports voice cloning, conversion, localization, and professional pronunciation control;
  • All packages offer unlimited work area seats, facilitating development and business collaboration;
  • It provides Python, JavaScript, MCP, and HTTP/WebSocket interfaces;
  • The Enterprise version offers SSO, compliance agreements, dedicated support, and customizable concurrency.

Usage restrictions and precautions

  • Package credits cannot be simply equated to a fixed number of minutes; the actual usage is influenced by the content and tasks involved.
  • The real-time voice quality depends on the network, audio devices, encoding, and endpoint parameters;
  • STT may still misidentify proper nouns, numbers, accents, and background noise;
  • The base model is used during professional cloning and binding training, and upgrading the model usually requires retraining;
  • The old models have an end-of-life cycle, and the production system must keep track of the official change logs;
  • Voice agents incur additional costs related to calls and large models, so it is necessary to set budgets and limits.
  • Voice cloning must not be used for impersonation, fraud, fake endorsements, or the replication of identities without consent.

Security, privacy, and audio compliance

For corporate projects, it is necessary to determine the location where data is processed, the retention period, logging practices, the sub-processors involved, and the mechanisms for obtaining consent regarding recording. In medical, financial, and other regulated sectors, DPA and BAA agreements sowie the relevant capabilities of the companies in question should be considered; however, their applicability still needs to be confirmed through contracts.

  • Voice samples, transcripts, and identity information should be collected to the minimum extent possible, with strict separation of permissions;
  • Save the rights holder, purpose, region, duration, and cancellation records for cloning sounds;
  • The API Key must not be placed in browsers, mobile application packages, or public repositories;
  • Add a synthetic identifier to the output, and provide channels for switching to a human agent and filing complaints;
  • Sensitive conversations should not be recorded, copied, or stored for an extended period without consent.

Explanation of open source, SDKs, and MCP

The Cartesia hosting platform, as well as the Sonic, Ink, and Line core models, are not open-source products; however, the developers have made available on GitHub the Python SDK, JavaScript SDK, MCP Server, and some Agent tools. The JavaScript SDK is licensed under the Apache 2.0 license, while the licenses for the other repositories need to be checked separately.

The open-source SDK is merely a client that calls the Cartesia API; it does not include the complete models or the inference infrastructure. MCP can be hosted via official OAuth connections, or the source code can be run locally, but actual generation and transcription still require an account and credits.

Frequently Asked Questions

Can Cartesia be used for free?

The Free plan provides 20,000 credits per month as well as a $1 credit for using agents, allowing for testing of TTS and STT functions. Commercial use licenses and instant voice cloning are available starting with the Pro plan.

How much is Cartesia Pro?

According to the official pricing page, the cost is 5 dollars per month, which includes 100,000 credits as well as a 5-dollar allowance for using the Agent feature. Any additional fees and taxes are charged through the control panel.

What is the difference between instant voice cloning and professional voice cloning?

Instant cloning takes about 10 seconds for a sample; it is available from the Pro plan. Professional cloning requires at least 30 minutes of high-quality audio recording; it is available starting from the Startup plan and involves 1 million credits for training purposes.

Does Cartesia support Chinese?

The current version of Sonic supports multiple languages; the official documentation indicates that it natively supports 42 languages. As for the Chinese voice options, regional settings, and model performance, these should be tested in the Playground.

Is Cartesia open source?

The core voice platforms and models are not open-source products; however, the official SDKs and other connection tools such as MCP have their source code made available. Even when using these clients, paid hosting APIs must still be utilized.

©️Copyright notice: Unless otherwise specified, all articles on this site are copyrighted bySharing of AI toolsAll content on this site is original; without permission, no individual, media outlet, website, or organization may reproduce, copy, or otherwise distribute it, nor may they create mirrors of it on servers that are not owned by this site. Otherwise, we reserve the right to take legal action against such parties in accordance with the law.

Tools similar to Cartesia