The subtitles say
Free value-added services
Comprehensive List of AI Tools AI audio tools

The subtitles say

Online tools for voice synthesis, subtitles, and video creation; no need to appear on camera to produce videos.

Tags:

What does the subtitle say?

The subtitles state that it is an AI tool for audio-text synchronization used in video post-production, capable of generating SRT subtitle timelines from existing audio and finalized script text.

It addresses the issue of having the text content ready but lacking accurate subtitle timing; it is not an ordinary text-to-speech tool.

Main functions

  • Upload the audio along with the final text version.
  • Automatically matches the timing of text and speech.
  • Generate a millisecond-level SRT subtitle timeline.
  • Reduce errors in proper nouns and characters with multiple pronunciations.
  • It supports the creation of video subtitles for existing scripts.
  • Alignment methods such as spaced processing are provided.
  • New users can receive a free trial period in minutes.
  • Professional alignment packages with long duration are available.
  • The account stores task and result records.
  • The text-to-speech synthesis feature is in a limited beta phase.

Which users are it suitable for

UserSuitable scenariosPremise
Short-video creatorsGenerate subtitles for voiceovers and commentary.The recording is basically read out according to the script.
Course instructorCreating subtitles for courses and knowledge-based paid contentRetain the complete lecture notes.
Corporate trainersOrganize subtitles for training videosObtain permissions to use audio and video content
Podcast editorGenerate a timeline for audio with script.There should not be too much impromptu content.
Audio content teamMatch the written text with the actual audio recordingThe monaural sound is clear.
Video post-production staffContinue formatting after exporting the SRT.Manual random checks at specific time points are still required.

The difference between phoneme alignment and speech recognition

ProjectPhoneme-text alignmentSpeech recognition
Is copy needed?A corresponding finalized copy is needed.It is usually not necessary.
Main tasksFind the time position of each text segment.Guessing the text content from speech.
proper nounsUse the text provided by the user directly.It may be identified as a homophone.
Impromptu contentAlignment may fail when the deviation is large.You can try to transcribe it directly.
Key points of the outputSubtitle timelineTranscribed text
Suitable materialsAudio and video of reading from a scriptMeetings, interviews, and free discussions

How does text-alignment work?

The system uses the text provided by the user as a reference to locate the starting and ending points of the corresponding phrases in the audio, and then generates subtitle blocks.

Therefore, the accuracy of the text depends on the final version of the script, while the accuracy of timing is influenced by the quality of the recording, the consistency of pronunciation, and sentence breaks.

  1. Prepare the final version of the spoken script.
  2. Export clear original audio.
  3. Verify that the order of the recordings matches that of the text.
  4. Upload the audio and paste the text.
  5. Select the appropriate alignment settings.
  6. Submit the task and wait for it to be processed.
  7. Preview subtitle timing and punctuation.
  8. Correct missed readings, additional readings, and pause positions.
  9. Export an SRT file for editing.

How to prepare audio

ProjectSuggestionsIt should be avoided.
SoundKeep the voice clear and the volume stable.Booming sound, too low volume, and severe reverb
BackgroundTry to use dry audio without background music.Music overpowers the speech.
Number of peopleIt is easier to achieve alignment when reading aloud individually in a continuous manner.Multiple people talking over each other and speaking simultaneously
EditingThe copy and the audio in the final video remain in the same order.Cut out phrases without changing the text.
PauseMaintain natural paragraph breaksLong periods of emptiness and repeated segments
fileUse audio that is clear and can be played properly.Corrupted files and multiple low-bitrate transcodings

How to prepare copywriting

The closer the text is to the actual spoken content, the more stable the alignment will be; punctuation and paragraphing also affect the reading rhythm of the subtitle blocks.

  • Use the version that was actually used during recording.
  • Delete unread titles and annotations.
  • Fill in the phrases that actually appear in the recording.
  • Standardize the spelling of numbers, units, and proper nouns.
  • Make natural pauses according to the meaning and speech pace.
  • Do not include time codes or irrelevant explanations.
  • Long documents are processed in batches, chapter by chapter.

What are SRT subtitles?

SRT is a common subtitle format; each subtitle block contains a sequence number, start time, end time, and the text to be displayed.

Once generated, it can be imported into most video editors and players, but the font, color, and position usually need to be set in the editing software.

Check itemsWhy is it important?Treatment method
Start timeSubtitles must not appear significantly earlier than the voice.Listen to each section and make minor adjustments.
End timePrevent subtitles from disappearing too early or overlapping.Reserve adequate time for reading.
Single line lengthBeing too long can affect reading on mobile devices.Break it into short sentences based on semantics.
Number of subtitlesExcessive fragmentation can cause the image to flicker.Merge consecutive phrases
EncodingIncorrect encoding may result in garbled text.Use the UTF encoding supported by the editor.

Interval and pause handling

Pauses, breaths, and blank sections in a speech affect the way subtitle segments are divided; it’s not sufficient to rely solely on whether the text is identical.

  • Short pauses can be retained within the same subtitle block.
  • A new subtitle block is suitable to be created at the end of a semantic unit.
  • The previous subtitle should be ended in time before a long pause.
  • The background music segments should not be forced to match the text.
  • The opening and closing sequences can be processed separately in video editing software.
  • After exporting, check the synchronization again at the final frame rate.

The subtitles state that a complete tutorial is available.

  1. Enter the subtitle to go to the official website.
  2. Register an account and obtain a trial credit.
  3. Select the audio-text alignment feature.
  4. Prepare clear audio without background music.
  5. Organize the final text to match the recording.
  6. Delete the unread titles and annotations.
  7. Upload the audio and paste the full text.
  8. Check the time limit, as well as the limits on files and characters.
  9. Select the settings related to alignment or spacing.
  10. Submit the task and wait for it to be generated.
  11. Preview the subtitle and audio synchronization.
  12. Pay special attention to the pauses and the sections where significant revisions have been made.
  13. Download the generated SRT subtitles.
  14. Import a video editing software.
  15. Set the font, position, and subtitle style.
  16. Review by playing the final video in its entirety.

Free quota and professional packages

The current homepage offers a 100-minute trial period free of charge upon registration, and it also displays a professional package that costs 125 yuan for 30,000 minutes of service.

ProjectPrice or quotaExplanation
New user experience100 minutesIt will be available upon registration; the rules of the campaign may change.
Professional Alignment Package125 yuanContains 30,000 minutes
Converted priceApproximately 0.004 yuan per minuteEstimate based on the total package price and number of minutes.
Speech synthesisNew prices have not been announced yet.It is currently in a limited beta phase or about to be launched.

The specific validity period, refund policies, invoices, and usage restrictions are subject to the purchase page and settlement rules available after logging in.

How to view the old membership rules?

The old official member page still shows that a deposit of 100 yuan grants 10,000 points and makes one a lifetime member, as well as highlighting previous benefits such as text-to-speech functionality.

Since the current homepage indicates that text-to-speech synthesis is available only as a limited beta feature, the rights and benefits of the old version cannot be considered as guarantees for the new version.

InformationContent of the old version pageCurrent usage recommendations
Membership requirementsA single top-up of 100 yuan or moreConfirm whether it is still applicable before purchasing.
PointsTop up 100 yuan to get 10,000 points.Confirm which current functions it can be used for.
Membership durationThe page states that it will be retained for life.Verify the actual equity of the account
Speech synthesisIt has offered a variety of member benefits.The current homepage is marked as in beta version.
Phoneme-text alignmentThere used to be daily limits and gift quotas.Refer to the new package page.

Current status of text-to-speech synthesis technology

On the current homepage of the official website, TTS is labeled as “coming soon” or “limited beta version”; therefore, regular users cannot assume that it is already fully available.

  • Do not commit to the current number of speakers based on the information in older articles.
  • Do not assume that TTS can be used before making a payment.
  • The eligibility for the beta test, as well as the audio quality and data allowance, may vary.
  • For commercial use of voiceovers, it is necessary to check the current licensing terms.
  • The benefits associated with the old account should be verified separately after logging in.

How should “100% accurate” be understood?

The claim of 100% accuracy on the official website is a marketing statement; it mainly serves to emphasize that the subtitles use the final text provided by the users.

It does not guarantee accuracy in terms of timeline, sentence segmentation, handling of missed parts, and across all audio conditions; manual verification is still required before publication.

SituationPossible outcomesKey points of inspection
Read aloud from the script.It is usually easier to match.Pause and subtitle block length
A few slips of the tonguePossible fault-tolerant alignmentTime points before and after the error
A great deal of improvisationPossible mismatches or segment skipsSwitch to transcribed and restructured copy
Multi-person conversationThe roles and order may be confused.Split processing or segmented handling
The background music is quite loud.Time positioning may be affected.Resubmit using dry sound.

Copyright and Privacy

Audio, text, and video content may contain confidential information; it is necessary to check the platform’s policies and the project’s confidentiality requirements before uploading.

  • Only process recordings that you own or have authorization for.
  • Permission must be obtained for the voices of the individuals and for the content of the interviews.
  • The script must not contain account or identity information.
  • Business courses should ensure that the copyright of materials is respected.
  • For confidential projects, consult the relevant regulations of your organization first.
  • Manage cloud tasks promptly after downloading the results.
  • Remove sensitive subtitle content before releasing it to the public.

API and open-source status

To date, no subtitles mentioning a public API, an official SDK, a product source code repository, or a verifiable official GitHub project have been found.

ProjectCurrent situationExplanation
Web toolsAlready providedUse audio-text alignment after registration.
Public APINot yet made publicDeveloper API documentation not found.
Official SDKNot yet made publicNo verifiable development package was found.
Product source codeNot open sourceOnline business services
Official GitHubNot confirmed yetProjects with the same name cannot be considered official.

Product advantages

  • Focus on the subtitle timeline for audio with transcripts.
  • Reduce spelling errors caused by ordinary recognition.
  • It is possible to export in the standard SRT format.
  • Registration offers free trial minutes.
  • The price per unit for long-duration packages is lower.
  • Suitable for courses and fixed script narration.

Usage restrictions

  • It is necessary to prepare text that corresponds to the recording.
  • It is not suitable for direct transcription of extensive impromptu speeches.
  • Background music and multiple people talking increase the difficulty.
  • High marketing accuracy does not mean that no manual review is needed.
  • The old membership rules may differ from those of the current product.
  • Speech synthesis is not yet a fully available feature.
  • The API, SDK, and open-source code are not yet available publicly.

Frequently Asked Questions

What does the subtitle say are its main functions?

It aligns the existing audio with the finalized script in terms of timing, and automatically generates an SRT subtitle file that can be imported into editing software.

Does the subtitle say it’s a text-to-speech tool?

No, it is primarily used for aligning audio and text, and users need to provide the corresponding text first; free-form speech without text is more suitable for speech recognition tools.

Does the subtitle say it can be used for free?

Yes, new users who register now can receive 100 minutes of trial time; the details of the promotion and the way to claim it are specified on the account page.

The subtitles ask how much the professional package costs.

The official website shows that 125 yuan covers 30,000 minutes, which translates to approximately 0.004 yuan per minute; the exact validity period is indicated on the settlement page.

Do the subtitles say that the generated subtitles are always accurate?

Text based on user-provided content can help reduce spelling mistakes, but errors may still occur regarding the timeline, sentence breaks, and wording; therefore a thorough review is necessary before publishing.

Does the subtitle say that voice synthesis is available now?

The current homepage indicates that the feature is about to be launched or is available only for a limited beta test; the text-to-speech capabilities offered on the old member page cannot be considered as being fully available at present.

Does the subtitle say that an API or open-source code is provided?

No verifiable public API, official SDK, product source code repository, or official GitHub project has been found at present.

©️Copyright notice: Unless otherwise specified, all articles on this site are copyrighted bySharing of AI toolsAll content on this site is original; without permission, no individual, media outlet, website, or organization may reproduce, copy, or otherwise distribute it, nor may they create mirrors of it on servers that are not owned by this site. Otherwise, we reserve the right to take legal action against such parties in accordance with the law.

Tools similar to those mentioned in the subtitles