Text-to-speech (TTS) turns written text into spoken audio. Neural TTS now sounds natural enough to be hard to distinguish from a human voice.
How It Works
Neural models predict the acoustic features of speech from text, including pronunciation, rhythm and intonation, and a vocoder turns those features into audio. Some systems can adopt a style or emotion on request.
Uses
- Accessibility: reading content aloud for people with visual impairments or reading difficulties.
- Voice assistants and phone systems.
- Narration for training material and audiobooks.
- Real-time translation combined with speech recognition.
Voice Cloning
Some systems can mimic a specific person's voice from a short sample. This enables personalised assistants and restoring voices for people who have lost them — and also impersonation and fraud.
Responsible Use
- Consent: only clone a voice with the person's explicit permission.
- Disclosure: tell people when they're hearing synthetic speech.
- Fraud awareness: voice alone should never authenticate a payment or sensitive request; use verification steps.
- Licensing: check the terms for commercial use of voices and generated audio.
Quality Tips
Use punctuation to guide pacing, spell out abbreviations and numbers where pronunciation matters, and review audio for mispronounced names and technical terms.