The subtitles say
Online tools for voice synthesis, subtitles, and video creation; no need to appear on camera to produce videos.
Tags:AI audio toolsWhat does the subtitle say?
The subtitles state that it is an AI tool for audio-text synchronization used in video post-production, capable of generating SRT subtitle timelines from existing audio and finalized script text.
It addresses the issue of having the text content ready but lacking accurate subtitle timing; it is not an ordinary text-to-speech tool.
Main functions
- Upload the audio along with the final text version.
- Automatically matches the timing of text and speech.
- Generate a millisecond-level SRT subtitle timeline.
- Reduce errors in proper nouns and characters with multiple pronunciations.
- It supports the creation of video subtitles for existing scripts.
- Alignment methods such as spaced processing are provided.
- New users can receive a free trial period in minutes.
- Professional alignment packages with long duration are available.
- The account stores task and result records.
- The text-to-speech synthesis feature is in a limited beta phase.
Which users are it suitable for
| User | Suitable scenarios | Premise |
|---|---|---|
| Short-video creators | Generate subtitles for voiceovers and commentary. | The recording is basically read out according to the script. |
| Course instructor | Creating subtitles for courses and knowledge-based paid content | Retain the complete lecture notes. |
| Corporate trainers | Organize subtitles for training videos | Obtain permissions to use audio and video content |
| Podcast editor | Generate a timeline for audio with script. | There should not be too much impromptu content. |
| Audio content team | Match the written text with the actual audio recording | The monaural sound is clear. |
| Video post-production staff | Continue formatting after exporting the SRT. | Manual random checks at specific time points are still required. |
The difference between phoneme alignment and speech recognition
| Project | Phoneme-text alignment | Speech recognition |
|---|---|---|
| Is copy needed? | A corresponding finalized copy is needed. | It is usually not necessary. |
| Main tasks | Find the time position of each text segment. | Guessing the text content from speech. |
| proper nouns | Use the text provided by the user directly. | It may be identified as a homophone. |
| Impromptu content | Alignment may fail when the deviation is large. | You can try to transcribe it directly. |
| Key points of the output | Subtitle timeline | Transcribed text |
| Suitable materials | Audio and video of reading from a script | Meetings, interviews, and free discussions |
How does text-alignment work?
The system uses the text provided by the user as a reference to locate the starting and ending points of the corresponding phrases in the audio, and then generates subtitle blocks.
Therefore, the accuracy of the text depends on the final version of the script, while the accuracy of timing is influenced by the quality of the recording, the consistency of pronunciation, and sentence breaks.
- Prepare the final version of the spoken script.
- Export clear original audio.
- Verify that the order of the recordings matches that of the text.
- Upload the audio and paste the text.
- Select the appropriate alignment settings.
- Submit the task and wait for it to be processed.
- Preview subtitle timing and punctuation.
- Correct missed readings, additional readings, and pause positions.
- Export an SRT file for editing.
How to prepare audio
| Project | Suggestions | It should be avoided. |
|---|---|---|
| Sound | Keep the voice clear and the volume stable. | Booming sound, too low volume, and severe reverb |
| Background | Try to use dry audio without background music. | Music overpowers the speech. |
| Number of people | It is easier to achieve alignment when reading aloud individually in a continuous manner. | Multiple people talking over each other and speaking simultaneously |
| Editing | The copy and the audio in the final video remain in the same order. | Cut out phrases without changing the text. |
| Pause | Maintain natural paragraph breaks | Long periods of emptiness and repeated segments |
| file | Use audio that is clear and can be played properly. | Corrupted files and multiple low-bitrate transcodings |
How to prepare copywriting
The closer the text is to the actual spoken content, the more stable the alignment will be; punctuation and paragraphing also affect the reading rhythm of the subtitle blocks.
- Use the version that was actually used during recording.
- Delete unread titles and annotations.
- Fill in the phrases that actually appear in the recording.
- Standardize the spelling of numbers, units, and proper nouns.
- Make natural pauses according to the meaning and speech pace.
- Do not include time codes or irrelevant explanations.
- Long documents are processed in batches, chapter by chapter.
What are SRT subtitles?
SRT is a common subtitle format; each subtitle block contains a sequence number, start time, end time, and the text to be displayed.
Once generated, it can be imported into most video editors and players, but the font, color, and position usually need to be set in the editing software.
| Check items | Why is it important? | Treatment method |
|---|---|---|
| Start time | Subtitles must not appear significantly earlier than the voice. | Listen to each section and make minor adjustments. |
| End time | Prevent subtitles from disappearing too early or overlapping. | Reserve adequate time for reading. |
| Single line length | Being too long can affect reading on mobile devices. | Break it into short sentences based on semantics. |
| Number of subtitles | Excessive fragmentation can cause the image to flicker. | Merge consecutive phrases |
| Encoding | Incorrect encoding may result in garbled text. | Use the UTF encoding supported by the editor. |
Interval and pause handling
Pauses, breaths, and blank sections in a speech affect the way subtitle segments are divided; it’s not sufficient to rely solely on whether the text is identical.
- Short pauses can be retained within the same subtitle block.
- A new subtitle block is suitable to be created at the end of a semantic unit.
- The previous subtitle should be ended in time before a long pause.
- The background music segments should not be forced to match the text.
- The opening and closing sequences can be processed separately in video editing software.
- After exporting, check the synchronization again at the final frame rate.
The subtitles state that a complete tutorial is available.
- Enter the subtitle to go to the official website.
- Register an account and obtain a trial credit.
- Select the audio-text alignment feature.
- Prepare clear audio without background music.
- Organize the final text to match the recording.
- Delete the unread titles and annotations.
- Upload the audio and paste the full text.
- Check the time limit, as well as the limits on files and characters.
- Select the settings related to alignment or spacing.
- Submit the task and wait for it to be generated.
- Preview the subtitle and audio synchronization.
- Pay special attention to the pauses and the sections where significant revisions have been made.
- Download the generated SRT subtitles.
- Import a video editing software.
- Set the font, position, and subtitle style.
- Review by playing the final video in its entirety.
Free quota and professional packages
The current homepage offers a 100-minute trial period free of charge upon registration, and it also displays a professional package that costs 125 yuan for 30,000 minutes of service.
| Project | Price or quota | Explanation |
|---|---|---|
| New user experience | 100 minutes | It will be available upon registration; the rules of the campaign may change. |
| Professional Alignment Package | 125 yuan | Contains 30,000 minutes |
| Converted price | Approximately 0.004 yuan per minute | Estimate based on the total package price and number of minutes. |
| Speech synthesis | New prices have not been announced yet. | It is currently in a limited beta phase or about to be launched. |
The specific validity period, refund policies, invoices, and usage restrictions are subject to the purchase page and settlement rules available after logging in.
How to view the old membership rules?
The old official member page still shows that a deposit of 100 yuan grants 10,000 points and makes one a lifetime member, as well as highlighting previous benefits such as text-to-speech functionality.
Since the current homepage indicates that text-to-speech synthesis is available only as a limited beta feature, the rights and benefits of the old version cannot be considered as guarantees for the new version.
| Information | Content of the old version page | Current usage recommendations |
|---|---|---|
| Membership requirements | A single top-up of 100 yuan or more | Confirm whether it is still applicable before purchasing. |
| Points | Top up 100 yuan to get 10,000 points. | Confirm which current functions it can be used for. |
| Membership duration | The page states that it will be retained for life. | Verify the actual equity of the account |
| Speech synthesis | It has offered a variety of member benefits. | The current homepage is marked as in beta version. |
| Phoneme-text alignment | There used to be daily limits and gift quotas. | Refer to the new package page. |
Current status of text-to-speech synthesis technology
On the current homepage of the official website, TTS is labeled as “coming soon” or “limited beta version”; therefore, regular users cannot assume that it is already fully available.
- Do not commit to the current number of speakers based on the information in older articles.
- Do not assume that TTS can be used before making a payment.
- The eligibility for the beta test, as well as the audio quality and data allowance, may vary.
- For commercial use of voiceovers, it is necessary to check the current licensing terms.
- The benefits associated with the old account should be verified separately after logging in.
How should “100% accurate” be understood?
The claim of 100% accuracy on the official website is a marketing statement; it mainly serves to emphasize that the subtitles use the final text provided by the users.
It does not guarantee accuracy in terms of timeline, sentence segmentation, handling of missed parts, and across all audio conditions; manual verification is still required before publication.
| Situation | Possible outcomes | Key points of inspection |
|---|---|---|
| Read aloud from the script. | It is usually easier to match. | Pause and subtitle block length |
| A few slips of the tongue | Possible fault-tolerant alignment | Time points before and after the error |
| A great deal of improvisation | Possible mismatches or segment skips | Switch to transcribed and restructured copy |
| Multi-person conversation | The roles and order may be confused. | Split processing or segmented handling |
| The background music is quite loud. | Time positioning may be affected. | Resubmit using dry sound. |
Copyright and Privacy
Audio, text, and video content may contain confidential information; it is necessary to check the platform’s policies and the project’s confidentiality requirements before uploading.
- Only process recordings that you own or have authorization for.
- Permission must be obtained for the voices of the individuals and for the content of the interviews.
- The script must not contain account or identity information.
- Business courses should ensure that the copyright of materials is respected.
- For confidential projects, consult the relevant regulations of your organization first.
- Manage cloud tasks promptly after downloading the results.
- Remove sensitive subtitle content before releasing it to the public.
API and open-source status
To date, no subtitles mentioning a public API, an official SDK, a product source code repository, or a verifiable official GitHub project have been found.
| Project | Current situation | Explanation |
|---|---|---|
| Web tools | Already provided | Use audio-text alignment after registration. |
| Public API | Not yet made public | Developer API documentation not found. |
| Official SDK | Not yet made public | No verifiable development package was found. |
| Product source code | Not open source | Online business services |
| Official GitHub | Not confirmed yet | Projects with the same name cannot be considered official. |
Product advantages
- Focus on the subtitle timeline for audio with transcripts.
- Reduce spelling errors caused by ordinary recognition.
- It is possible to export in the standard SRT format.
- Registration offers free trial minutes.
- The price per unit for long-duration packages is lower.
- Suitable for courses and fixed script narration.
Usage restrictions
- It is necessary to prepare text that corresponds to the recording.
- It is not suitable for direct transcription of extensive impromptu speeches.
- Background music and multiple people talking increase the difficulty.
- High marketing accuracy does not mean that no manual review is needed.
- The old membership rules may differ from those of the current product.
- Speech synthesis is not yet a fully available feature.
- The API, SDK, and open-source code are not yet available publicly.
Frequently Asked Questions
What does the subtitle say are its main functions?
It aligns the existing audio with the finalized script in terms of timing, and automatically generates an SRT subtitle file that can be imported into editing software.
Does the subtitle say it’s a text-to-speech tool?
No, it is primarily used for aligning audio and text, and users need to provide the corresponding text first; free-form speech without text is more suitable for speech recognition tools.
Does the subtitle say it can be used for free?
Yes, new users who register now can receive 100 minutes of trial time; the details of the promotion and the way to claim it are specified on the account page.
The subtitles ask how much the professional package costs.
The official website shows that 125 yuan covers 30,000 minutes, which translates to approximately 0.004 yuan per minute; the exact validity period is indicated on the settlement page.
Do the subtitles say that the generated subtitles are always accurate?
Text based on user-provided content can help reduce spelling mistakes, but errors may still occur regarding the timeline, sentence breaks, and wording; therefore a thorough review is necessary before publishing.
Does the subtitle say that voice synthesis is available now?
The current homepage indicates that the feature is about to be launched or is available only for a limited beta test; the text-to-speech capabilities offered on the old member page cannot be considered as being fully available at present.
Does the subtitle say that an API or open-source code is provided?
No verifiable public API, official SDK, product source code repository, or official GitHub project has been found at present.
Guigong Network Security Registration No. 45132202000164