Fish Audio
Free value-added services
Comprehensive List of AI Tools AI audio tools

Fish Audio

A voice platform that offers multilingual text-to-speech, voice cloning, and API services.

Tags:

What is Fish Audio?

Fish Audio is an AI voice platform designed for content creation, voiceovers, localization, real-time voice applications, and developer integration. Users can enter scripts on a web interface to generate natural-sounding speech; they can also use short recordings to instantly create similar voices, train voice models for long-term use, produce multi-voice audiobooks, change the tone of existing recordings, and carry out tasks such as transcription, generation of music and movie sound effects, and track separation.

Currently, the platform relies on voice models such as S2.1 Pro, S2 Pro, and S1 as its core components; it also offers REST API, WebSocket streaming interfaces, as well as SDKs for Python and JavaScript. There are two separate billing systems for web members and those who use the API: members receive Credits on a monthly basis, while API usage is charged based on the number of UTF-8 bytes, the length of the audio, or the number of successful requests. It is necessary to determine which billing option to use before making a purchase.

Text to speech

Text-to-speech functionality supports over 80 languages, and it can handle text in Chinese, English, as well as combinations of multiple languages. Users can choose from pre-installed voices on the platform, their own cloned voices, or professional voices; they can also adjust the speech speed, tone, and delivery style, and export the result in formats such as MP3, WAV, PCM, or Opus.

Numbers, dates, currencies, abbreviations, brand names, and technical terms can still be mispronounced. For formal projects, it is necessary to create a set of pronunciation tests first, break down long scripts into smaller segments for listening, and then establish uniform rules for proper nouns and pauses.

S2.1 Pro voice model

S2.1 Pro is the currently recommended production model, with a focus on improving sound quality, latency in the first packet, and throughput. It supports natural language direction labels, multi-speaker conversations, language switching, real-time voice cloning, streaming audio, and word-level timestamps; it is suitable for voice avatars, game characters, interactive stories, and large-scale narration.

The delay of the first audio clip displayed on the official website is around 100 milliseconds, or within the range of 70 to 90 milliseconds; however, these figures apply under specific hardware, regional, textual, and network conditions, and they do not mean that all requests will achieve such delays. In a production environment, it is also necessary to take into account factors such as end-to-end network latency, queueing, player buffering, and concurrency limits.

S2 Pro and S1

The S2 Pro is the previous generation model; it allows for the use of natural language within square brackets to describe emotions, tone, and actions, and it can generate conversations with multiple speakers at once. The S1 uses parentheses to indicate emotions, and it is suitable for integrating into existing workflows as well as for creating specific sound effects.

Different models perform differently when it comes to labels, languages, short sentences, and complex texts. When transferring a model, it is not sufficient to merely change its name; it is necessary to reevaluate factors such as similarity in tone, rhythm, sensitive words, timestamps, and unit cost.

Instant voice cloning

Instant cloning allows the extraction of timbre and speaking style from a reference audio clip, which can then be used for voice generation in other languages. Short audio samples are suitable for quick testing; the cleaner the recording, the more it is done by a single person, with no music present and consistent pronunciation, the more reliable the results tend to be.

Similar sounding voices do not mean that identity, emotions, and accent can be perfectly replicated. Any use of a real person’s voice requires their explicit consent, along with clear specifications regarding the purpose, duration, geographical scope, users who may have access to it, and the methods for deleting it.

It must not be used to pretend to be someone else, to carry out fraud, or to make misleading statements.

Enhanced professional voice cloning

The web-based solutions also differentiate between enhanced cloning and Professional Voice. Professional Voice requires more structured recording and training processes; it is suitable for long-term brand voices, audiobooks, and projects that demand high consistency, and it is subject to limitations related to the number of available Professional Voice slots among members.

Public, private, and professional voice slots do not represent the number of minutes generated. Public voices can be accessed by members of the community, while private voices are intended for use only by a specific account or team.

When dealing with actors, clients, and unreleased brand voices, private permissions should be preferred and the sharing settings checked.

Voice Design – Sound design

Voice Design enables one to describe in text factors such as age, timbre, accent, speed, personality, and intended use, and the system then generates several possible voice options. It is suitable for quickly exploring different vocal styles for fictional characters, advertisements, and games, without the need to find existing audio recordings of real people as a reference.

Text descriptions may yield similar or inconsistent results. For important characters, reference audio and Voice ID should be kept, in order to verify whether the candidate voices are too similar to the actual individuals or protected characters.

Multi-person long audio and audiobooks

The web version allows for the organization of multiple characters and lengthy scripts, enabling the creation of podcasts, conversations, audiobooks, and narration projects. S2.1 Pro’s built-in multi-user functionality enables switching between speakers during production, thus reducing the need for post-production editing.

Long texts tend to suffer from issues such as uneven stress placement, incorrect allocation of voices, and differences in volume. It is recommended to create the text in sections, assign a specific voice to each character, and conduct a full listening test along with volume normalization before exporting the final version.

Voice Changer – voice transformation

Voice Changer converts the uploaded audio recording into another voice, while striving to maintain the original rhythm and expression. It is suitable for replacing temporary narrations with a consistent character voice, and can also be used in cross-lingual or creative audio projects.

Background music, reverb, overlapping voices, and loud noise can reduce the quality of the conversion. Voice transformation does not alter the copyright and licensing requirements associated with the original recording; both the original speaker and the target voice must have legitimate authorization.

Audio transcription

Fish Audio’s ASR capability converts audio into text with timed timestamps, which facilitates the creation of subtitles, searching, and subsequent voiceover work. The transcribe-1 API is currently charged on a per-hour basis for audio, with the cost being rounded up to the nearest second.

Multiple speakers speaking at the same time, dialects, noise, and proper nouns can affect accuracy. Legal, medical, financial, and subtitles for public use require manual proofreading; automated transcriptions cannot be used as literal evidence.

Music, sound effects, and audio tools

Web applications also offer tools for creating music and movie sound effects based on prompts, as well as for separating audio tracks and performing other audio-related tasks; these tools are suitable for use in creating materials for short videos, podcasts, and games. The amount of Credits required by each tool, as well as the rules regarding commercial use, may vary, so it is necessary to check these details on the generation page.

AI-generated content does not guarantee the absence of any copyright issues. Before publishing, it is necessary to keep a record of the creation process, and to check the platform’s terms, the degree of similarity to existing samples, as well as the requirements of the intended distribution channels regarding AI-generated music.

Free trial version

The free version costs $0, no credit card is required; 8,000 Credits are provided each month. According to the official website, the generation time is approximately 7 minutes at most. Each output can contain up to 500 characters, and it includes 3 available audio slots, standard generation speed, enhanced voice cloning, as well as the option for commercial use.

\"Up to 7 minutes\": The cost is estimated at 600 to 625 Credits per minute; the actual amount used will depend on the model, features, and settings. The monthly Credits are not a permanent balance, and the free plan is not suitable for storing audio content that requires strict confidentiality.

Price and version comparison

Package or versionPrices, quotas, and core benefits
Plus priceThe current promotional offer for the Plus annual plan is $5.5 per month, or $66 for the whole year; the original price on the page is $15 per month. 250,000 Credits are provided each month, which, based on average usage, are sufficient for around 200 minutes of use. The maximum number of characters that can be used in a single message is 15,000; the plan includes unlimited access to public voices, 10 private voice slots, Voice Design features, 1 professional voice slot, priority access to the latest models, enhanced cloning capabilities, and permission for commercial use. The annual discounts are part of time-limited promotions, and the original price may revert once these promotions end. The actual cost should be based on the currency, taxes, renewal price, and payment terms shown on the billing page.
Pro priceThe current promotional offer for the Pro annual plan is $37.5 per month, amounting to $450 for the whole year; the original price on the page is $100 per month. The plan includes 2,000,000 Credits per month, which equates to roughly 1620 minutes of usage, 3 team seats, 30,000 characters per message, unlimited access to both public and private voice channels, as well as 5 professional voice channels. A 7-day refund guarantee is provided. The refund will be granted depending on the purchase method, the amount of usage, and the current refund policies; not all orders with high levels of usage will necessarily qualify for a refund. Teams should clarify the permissions of their members, the ways in which projects can be shared, and who owns the voice channels before making the purchase.
Max priceThe annual plan for Max shows an average cost of $749 per month, or $8,988 for the whole year; the page also indicates the original price of $999 per month. 25,000,000 Credits are provided each month, which corresponds to approximately 6,250 minutes of usage, 10 team seats, and 15 professional voice slots – making it suitable for teams that carry out frequent production tasks. The Credits available with Max cannot be simply converted into minutes at a fixed ratio, as there may be differences depending on the specific model or workflow used. Customers who handle large volumes of work should conduct cost calculations based on their actual scripts, rather than purchasing based solely on the maximum number of minutes allowed.
Pay-as-you-go API pricingThe API does not require a monthly subscription or a minimum monthly spending amount. For S2.1 Pro, S2 Pro, and S1 text-to-speech services, the cost is currently 15 dollars per 1 million UTF-8 bytes. According to official estimates, 1 million characters correspond to 180,000 English words or 12 hours of audio output; however, Chinese characters typically take up more UTF-8 bytes, so the English-word conversion cannot be applied directly. The cost for ASR services is currently 0.36 dollars per audio hour, with rounding up to the nearest second; Voice Design charges 0.01 dollar per successful request. Failed requests, retries, splitting of long texts, and generating multiple versions all affect the final cost.

Plus price

The current promotional offer for Plus’s annual plan is $5.5 per month, or $66 for the whole year; the original price on the page was $15 per month. 250,000 Credits are provided each month, and based on standard usage, this amounts to about 200 minutes of usage.

Up to 15,000 characters per entry; includes unlimited public voices, 10 private voice slots, Voice Design, 1 professional voice slot, priority access to the latest models, enhanced cloning, and commercial use.

Anniversary or annual discounts are part of real-time promotions, and the original prices may revert once these promotions end. The prices should be based on the currency, taxes, renewal fees, and payment terms indicated on the settlement page.

Pro price

The current promotional offer for the Pro annual plan is $37.5 per month, or $450 for the whole year; the original price on the page is $100 per month. The plan includes 2,000,000 Credits per month, up to approximately 1620 minutes of usage, 3 team seats, 30,000 characters per message, unlimited public and private voice channels, as well as 5 professional voice channels. A 7-day refund guarantee is also provided.

The refund guarantee is determined based on the purchase channel, the amount used, and the current refund policies; not all orders that have used up a large amount of quota will necessarily receive a refund. The team should confirm the permissions of its members, as well as matters related to project sharing and ownership of audio files, before proceeding with the actual purchase.

Max price

Max’s annual plan amounts to $749 per month, or $8,988 for the whole year; the page also shows the original price of $999 per month. 25,000,000 Credits are provided each month, which translates to approximately 6,250 minutes of usage, 10 team seats, and 15 professional voice slots – making it suitable for teams that carry out frequent production tasks.

Max’s Credits are not converted from the nominal number of minutes in a simple, uniform ratio; this indicates that there may be differences between different models or workflows. Customers with large volumes of work should conduct cost tests based on the actual scripts, rather than making purchases based solely on the maximum number of minutes.

Enterprise business solution

Enterprise offers customized quotes, and provides annual or usage-based subscription plans, bulk discounts, zero data retention options, private or on-premises deployment, SOC 2 compliance features, as well as custom support; SSO may still be listed as upcoming on the current comparison page.

For zero data retention and local deployment, it is necessary to include relevant provisions in the contract and clarify the scope of what is not retained; merely because a corporate label appears on the page, it cannot be assumed that all inputs, logs, backups, and audio models will not be saved.

Pay-as-you-go API pricing

The API does not require a monthly subscription or a minimum monthly expenditure. At present, the cost for text-to-speech conversion using S2.1 Pro, S2 Pro, and S1 is $15 per 1 million UTF-8 bytes.

Official estimates suggest that 1 million Chinese characters correspond to 180,000 English words or 12 hours of audio playback; however, Chinese characters occupy more UTF-8 bytes, so the conversion using English words cannot be applied directly.

ASR currently costs $0.36 per audio hour, with the cost being rounded up to the nearest second; Voice Design charges $0.01 for each successful request.

Failure, retries, splitting of long texts, and the creation of multiple versions all affect the final cost.

S2.1 Pro Free API

S2.1-pro-free costs $0 under the fair use policy; it supports 83 languages and is intended for testing, prototyping, development, and small-scale businesses. It uses the same model as S2.1 Pro, but does not offer features such as low latency for the first data packet, data processing protocols, or enterprise-level service guarantees. The duration of the free period and related rules may also change.

Free models do not imply unlimited production as per an SLA. For critical operations, it is necessary to use paid models, implement retry mechanisms in case of delays, employ circuit breaking, use caching, and find alternative suppliers; moreover, it is important to regularly check the limits on fair usage.

API concurrency level

The maximum number of concurrent tasks is determined based on the total prepaid amount for each account: accounts in the Starter category, with a total payment of less than 100 dollars, can generally have 5 concurrent tasks; accounts in the Elevated category, with a total payment of 100 dollars or more, can have 15 concurrent tasks.

For a volume of 1,000 dollars, the value of High Volume is 50 concurrent connections; Enterprise plans are determined according to the terms of the contract.

The number of concurrent tasks is not equal to the number of requests per second, nor does it guarantee a fixed latency. Long-running requests keep the connection occupied for longer periods; therefore, developers need to use queues, rate limiting, exponential backoff, and mechanisms for monitoring.

REST, WebSocket, and SDKs

Developers can initiate standard generation via REST, receive streaming audio through WebSocket, and use the official Python and JavaScript SDKs to manage TTS, transcription, and voice models. The interfaces allow for creating, listing, updating, and deleting voices, and they can also be integrated with real-time speech frameworks such as LiveKit and Pipecat.

The API Key must be stored on the server or in a key management system; it should not be included in web pages, mobile application packages, or public repositories. Production systems should also verify the file type, limit the length of the text, record the request ID, and handle disconnections in streaming connections.

GitHub and the open-source status

The Fish Audio web platform, the hosted APIs, the account management system, and the community sound market are not fully open-source products. GitHub offers the source code for Fish Speech, various model-related resources, SDKs, and maintenance tools; however, the availability of this code does not mean it can be used for commercial purposes without any restrictions.

The code of the main branch of Fish Speech, as well as the associated model weights, are all subject to the Fish Audio Research License. Research and non-commercial use are permitted under this license.

Any commercial use, including internal corporate operations, paid products, hosting services, or APIs, requires a separate written commercial license from Fish Audio.

The information on Apache or CC-BY-NC-SA in the old pages may refer to previous versions; it is necessary to follow the current license that comes with the version being used for making decisions.

Local deployment

The official Fish Speech documentation provides options for installation on Linux or WSL, as well as via Docker, WebUI, and API servers. For inference tasks using the S2 series, it is recommended to have around 24GB of video memory; a CPU-based approach is also available. Companies can also seek guidance on deployment in VPCs, on-premises environments, sovereign clouds, or isolated networks.

The ability to download and run a model does not entail the right to use it in commercial contexts. With self-hosting, one is responsible for covering the costs associated with GPUs, ensuring model security, handling updates, preventing misuse, and complying with regulations regarding logs and data; commercial projects require a separate license before they can be used in such contexts.

Commercial use and sound licensing

Both the free and paid versions of the website indicate that commercial use is allowed, but this authorization applies only to the outputs generated by using the service under compliance with the relevant terms; it does not grant users rights to the scripts, music, voice recordings, characters, trademarks, or other materials uploaded.

The rights associated with hosted services and the license for Fish Speech research are two separate things: creating commercial content through members or APIs does not mean that the downloaded model weights can be used for commercial, self-hosted purposes. For sensitive projects, it is necessary to retain consent forms, information about the source of the audio, and records of approval processes.

Fish Audio usage guide

Complete a basic task.

  1. Verify recordings, audio, music, and participant authorization;
  2. Upload or enter clear audio into Fish Audio;
  3. Select settings such as language, speaker, and text-to-speech;
  4. Use the S2.1 Pro speech model to generate transcriptions, voiceovers, or cleaned-up outputs;
  5. Check each segment for names, numbers, pauses, volume, and mood;
  6. Before exporting, verify the format, loudness, copyright, and privacy requirements;

Create reusable professional workflows

  1. A test set is created using real noise, accents, and multi-person segments;
  2. Compare the differences in performance between text-to-speech, the S2.1 Pro speech model, and the S2 Pro and S1 models.
  3. Retain the original recordings and the unmodified transcripts;
  4. Arrange for a manual hearing before releasing it to the public;
  5. Statistically analyze processing time, error rate, and quota consumption;
  6. Regularly update the glossary, sound licensing, and deletion policies;

Which users is it suitable for?

  • Creators who produce video narrations, podcasts, audiobooks, and audio content for courses;
  • A team is needed for multilingual voiceovers, character dialogues, and localization.
  • Companies that wish to establish a brand voice or character voices quickly;
  • Developers who create Voice Agents, customer service tools, companion applications, and products designed for people with disabilities;
  • Technical teams that research and evaluate voice models capable of running locally.

Product advantages

  • It covers over 80 languages and supports seamless switching between them;
  • S2.1 Pro supports natural language performance tags and native multi-person conversations;
  • The web version offers tools for cloning, processing long audio files, changing voices, transcribing text, and handling audio.
  • REST, WebSocket, as well as Python and JavaScript SDKs are fairly comprehensive;
  • Provide a free API model under fair use restrictions;
  • There are both hosting services, as well as downloadable models for research purposes and options for enterprise deployment.

Restrictions and Precautions

  • The Credits on the website, the API balance, and self-hosted licenses are independent of one another; member minutes cannot be used directly for the API.
  • Promotional prices, free models, and fair use policies may change; the budget should be based on the settlement page and the control panel.
  • AI voices can exhibit variations in pronunciation, character portrayal, emotion, and similarity.
  • Voice cloning, voice modification, and community voices require authorization verification, and important content needs to be listened to manually;
  • The current research license for FishSpeech does not grant commercial usage rights;

Frequently Asked Questions

Is Fish Audio free?

The Free web version provides 8,000 Credits per month, which is sufficient for up to about 7 minutes of use; the API offers the s2.1-pro-free model under fair usage limits, but it lacks any production SLAs or data processing guarantees.

How much is Fish Audio?

During the current annual promotion, Plus costs 66 dollars for the whole year, Pro costs 450 dollars, and Max costs 8988 dollars; Enterprise options are available upon request.

Promotional and renewal prices may change; the prices shown on the settlement page are applicable.

How are API fees charged?

S2.1 Pro, S2 Pro, and S1 TTS currently charge $15 per 1 million UTF-8 bytes; the transcription service costs $0.36 per audio hour.

Voice Design charges $0.01 for each successful request.

Is Chinese supported?

It supports Chinese and over 80 other languages, as well as mixed text in multiple languages. When charging for Chinese characters based on UTF-8 bytes, it is not possible to use estimates derived from English words.

Is it possible to clone one’s own voice?

Instant, enhanced, or professional cloning can be used, but explicit permission from the speaker is required; moreover, a public or private voice slot must be chosen based on privacy considerations.

Is Fish Audio open source?

The hosting platform is not an open-source product. The source code and models for Fish Speech are available publicly, but it is governed by the Fish Audio Research License; research and non-commercial use are permitted under this license, while commercial use requires separate written permission.

©️Copyright notice: Unless otherwise specified, all articles on this site are copyrighted bySharing of AI toolsAll content on this site is original; without permission, no individual, media outlet, website, or organization may reproduce, copy, or otherwise distribute it, nor may they create mirrors of it on servers that are not owned by this site. Otherwise, we reserve the right to take legal action against such parties in accordance with the law.

Tools similar to Fish Audio