How to Use Index TTS Online

Create speech with your own voice.
Follow the steps and screenshots.

Open the tool

Before you begin

Have your script and reference file ready on your device. Choose a small first task, such as a welcome message or a two-sentence explanation. Use only speech you have permission to upload and use for your project. A single clear speaker is a better starting point than a recording with music, overlapping voices, or strong echo.

Check your account and available credits before a longer session. The signed-out interface displays Try for free and a sign-in prompt; review the balance shown after signing in rather than assuming every request is unlimited. If you need more credits, visit Pricing. Save your original text outside the tool so it is easy to revise or split into sections later.

Open the tool and prepare your account

Open the online tool on the homepage. The left side contains your inputs and the Generate Speech button. The right side initially shows Audio Examples, with playable Index TTS 2.5 and Index TTS 2 samples. On a narrow screen, these areas stack vertically, so scroll below the inputs to find the examples or your result.

Sign in before uploading audio or generating speech. If you are signed out, starting one of those actions opens the sign-in window. Complete the prompts, then return to the tool and check your inputs. Use the same account later when you want to find a recording in My Creations.

The sample players are useful for hearing example speech before you begin. Their model names describe those samples; they are not model-selection buttons. For this walkthrough, keep the task simple: prepare one short script and one authorized reference recording, then make a first test before working on a longer project.

Open the tool and prepare your account
The tool before generation. Audio Examples contains official samples, not your own output. Click the screenshot to view it full size.

Enter the words you want spoken

Click the Text field and type or paste your script. The counter beside the label shows how much of the 1,000-character allowance you have used. For a first attempt, one or two complete sentences are enough to check the voice, pronunciation, and pauses without committing a full passage.

The screenshot uses this practice text: Welcome to our guide. Today, we will create a short voice recording with Index TTS. You can use the same text or substitute a short introduction for your own project. This is an input example, not a claim about a tested generation result.

Read the field after pasting, especially the final sentence. Remove headings, editing notes, and instructions that should not be spoken. Use punctuation to separate ideas. If your script is longer than the limit, split it at natural sentence boundaries and keep the complete original in your own document for reference.

Enter the words you want spoken
Enter only the words you want to hear. The counter helps you stay within the 1,000-character limit. Click the screenshot to view it full size.

Upload the required reference voice

Under Reference voice · Required, click Upload reference audio and choose a recording from your device. The interface recommends a clean 3–10 second clip and lists WAV, MP3, and M4A. Use a voice you own or have permission to use for the intended recording.

Listen to the file before uploading it. Choose one clearly audible speaker, with little background noise, no competing music, and a steady recording level. A short, clear sample is more useful for your first comparison than a noisy conversation with several voices. Keep a copy of the reference so you can reuse it for later sections.

Wait for the upload to finish and confirm the selected filename in the upload area. If it reports a failure, retry or choose another playable file. Selecting a file does not by itself confirm that the upload succeeded. The required reference establishes the speaker input; an optional emotional recording in Advanced settings does not replace it.

Upload the required reference voice
Use Upload reference audio for the speaker reference. The separate emotional audio field is optional. Click the screenshot to view it full size.

Review optional Advanced settings

Click Advanced settings to expand the additional controls. You can leave the initial values unchanged for a first test. Collapsing this section only hides the controls; it does not turn the settings off. If you change a slider and then close the section, that selection still applies to your next request.

Emotional audio is an optional reference for delivery style. Use Upload emotional audio when you have a suitable, authorized recording and want to explore its emotional influence. Keep the required speaker reference in place. Wait for this extra upload to finish before generating, and check its displayed filename just as you would for the main reference.

Strength controls emotional style-transfer strength, on a scale from 0 to 1. The initial value is 0.8. It is not a playback-speed slider. Begin near the existing value and compare a small adjustment before making a large change. The question-mark controls provide short explanations of the fields.

Emotional strengths provides separate sliders for Happy, Angry, Sad, Afraid, Disgusted, Melancholic, Surprised, and Calm. Each runs from 0 to 1 in 0.1 steps. The screenshot shows the initial mix, including nonzero Sad, Afraid, Melancholic, Surprised, and Calm values. Record the values you use and change one at a time so that comparisons remain meaningful.

Review optional Advanced settings
The actual expanded controls. Some emotion values are already nonzero in the initial settings. Click the screenshot to view it full size.

Click Generate Speech and wait for the result

Check the text, the required reference filename, and any advanced settings. Then click Generate Speech. If you have not signed in, complete sign-in when prompted. If an input needs attention, read the error message and correct that issue before trying again.

During generation, the tool shows a synthesis status and progress display. Allow the request to finish rather than submitting the same text repeatedly. A percentage is not an exact time estimate, and reaching the end of the initial progress can still be followed by a wait for the playable audio to become available.

Watch the output area for the completed recording. If it still says the audio is being generated, wait or check the task in My Creations. When a request fails, retain your script and read the message before changing several inputs. For account or balance issues, check your signed-in account and available credits before making another attempt.

Click Generate Speech and wait for the result
Submit your text and uploaded references with Generate Speech. Click the screenshot to view it full size.

Listen, download, and compare another version

After a successful generation, the result area shows Your Generated Speech with an audio player. Press Play and follow your script while listening. Check that the entire passage is present, including the final sentence. Pay attention to names, numbers, pauses, and whether the delivery fits your purpose.

Use Download below your generated player to save the audio. Open the downloaded file on your device to verify that it is the recording you intended to keep. The screenshot here shows the output area's initial sample state, so do not mistake an official example for your completed request.

To compare a revision, save your preferred version first. Edit the text, replace a reference, or adjust an emotion setting, then use Generate Again. This submits another generation with the current inputs rather than replaying the existing recording. Change one variable per comparison and listen to the same sentence in both results.

Listen, download, and compare another version
Output-area location before generation. This screenshot shows official samples; your completed recording appears under Your Generated Speech. Click the screenshot to view it full size.

Find and organize your finished audio

Open My Creations from your account menu to revisit tasks and completed recordings. Use Refresh when you need an updated list, and check the selected status filter if you cannot find an item. A task that is still processing may not have a playable file yet.

On a completed recording, use Download audio to save a copy. Give local files useful names, such as introduction-test or introduction-final, so you can distinguish revisions. If a card shows an auto-cleanup notice, save anything you want to retain before that period expires. Delete removes a creation through a confirmation step; it is not needed when you simply want to generate a different version.

Troubleshoot the next action

The reference will not upload

Check that the file plays on your device and that your connection is working. Try a WAV, MP3, or M4A recording, then wait for upload completion. If the required reference failed, uploading an emotional clip alone will not satisfy that requirement.

The tool asks me to sign in

Finish the sign-in prompts and return to the tool. Confirm that you are using the expected account. Check the text and uploaded reference before submitting again, especially if you moved away from the page during sign-in.

Generation has not produced playable audio

A processing message means the final audio may still be pending. Give the existing task time to complete and check its status in My Creations. Avoid creating repeated copies of the same request simply because the player has not appeared yet.

The result or playback needs attention

If playback fails, try Download or open the completed item in My Creations. If pronunciation is unclear, test the difficult word in a shorter sentence. For an unsuitable voice, compare a clearer speaker reference before changing several emotional settings.

How to Use Index TTS FAQ

Yes. Upload a speaker recording in Reference voice · Required before generating. Emotional audio in Advanced settings is optional and serves a different purpose; it does not replace the required speaker reference.

Create your first short recording

Enter a short script, upload your reference voice, and start with the existing settings. Listen to the result, save a useful version, and make one deliberate change when you want to compare another take.

Try Index TTS