Add a reference voice
Upload a WAV or MP3 file, or record 3–30 seconds of clear, single-speaker speech. Once ready, you can save the reference to My Voices immediately.
Upload or record a reference voice, enter English or Chinese text, and generate natural IndexTTS2 speech directly in your browser—no local setup or GPU required.
IndexTTS2 Voice Cloning
Free Public BetaAdd a short reference voice and the text you want it to speak.
Add 3–30 seconds of clear, single-speaker audio.
Add a reference voice and text to get started.
WAV output · Signed-in generations can be viewed in History.
By generating, you confirm that you are 18 or older and that you own this reference voice or have permission to use it. Use is subject to our Terms and Acceptable Use Policy.
See IndexTTS in action
Official research samples show the model output before you try the hosted IndexTTS2 workspace yourself.
IndexTTS2 is the model currently hosted on IndexTTS Online.
Choose a demo
How it works
Add a clean reference, save it for reuse if you want, write what it should say, then generate and download WAV output.
Upload a WAV or MP3 file, or record 3–30 seconds of clear, single-speaker speech. Once ready, you can save the reference to My Voices immediately.
Enter English, Chinese, or mixed text and use punctuation to shape natural pauses.
Choose a supported emotion, generate the speech, then listen, retry, or download the WAV output.
Pricing
Free is designed for short evaluation and occasional use. Pro is the defined production plan for continuous voice generation.
Free
Try without an account
Create a free account for more
Pro
IndexTTS Models
IndexTTS2 is available here now. IndexTTS 2.5 is the newer research model and is not hosted for generation yet.
Available online
Zero-shot voice cloning with follow-reference emotion or four optional emotion overrides.
Current hosted model · English & Chinese text · WAV output
Latest research
Five-language research demos with a reported 2.28× RTF improvement.
Research page · Chinese, English, Japanese, Spanish and Arabic
VERIFIED CAPABILITIES
These capabilities are described by the official IndexTTS2 project, without unverified usage or customer numbers.
Zero-shot
Reproduce a target speaker from reference audio without per-speaker model training.
Duration
Support both controlled synthesis duration and natural-duration generation modes.
Emotion
Separate speaker identity and emotional expression through different conditioning inputs.
FAQ
Practical answers about IndexTTS2, the hosted workflow, and current Public Beta limits.
IndexTTS2 is an open-source text-to-speech model from the Index SpeechTeam with zero-shot voice cloning and emotion control. IndexTTS Online provides an independent browser-based hosted workflow around the current IndexTTS2 runtime.
Start with one clear voice
Upload or record a short sample, write the line, and generate speech without installing the model locally.