# What Is Text to Speech > Text to speech (TTS) is technology that converts written text into spoken audio. Modern neural TTS normalizes the text, predicts the sounds, rhythm, pitch and pauses of a voice, then generates the audio waveform. It powers screen readers, voice assistants, audiobooks, voice-overs and phone menus, and today's best engines can add emotion, whispers and laughter. URL: https://kenerateai.com/what-is-text-to-speech Publisher: Kenerate AI (https://kenerateai.com) Last updated: 29 September 2026 ## Key facts - TTS stands for: Text to speech - Also called: Speech synthesis - Neural TTS: Deep-learning models that generate the waveform - Main uses: Accessibility, voice-overs, audiobooks, assistants, IVR - Opposite of: Speech to text (transcription) - Leading engines: ElevenLabs Eleven v3, MiniMax Speech 2.8 - Demo "Expressive Narrator" (English, MiniMax Speech 2.8 HD): script: Text to speech turns written words into spoken audio. Modern systems predict not just the sounds, but the rhythm, the pauses and the emotion of a real speaker. - Demo "Roger" (English, ElevenLabs Eleven v3): script: Text to speech turns written words into spoken audio. Modern systems predict not just the sounds, but the rhythm, the pauses and the emotion of a real speaker. - Demo "Aiden" (English, Qwen3 TTS): script: Text to speech turns written words into spoken audio. Modern systems predict not just the sounds, but the rhythm, the pauses and the emotion of a real speaker. - Demo "Calm Storyteller" (English, Kling V1 TTS): script: Text to speech turns written words into spoken audio. Modern systems predict not just the sounds, but the rhythm, the pauses and the emotion of a real speaker. - Demo "Upbeat Woman" (English, MiniMax Speech 2.8 HD, Emotion: happy): script: I can't believe it's finally here. After all this time, it's really happening. - Demo "Upbeat Woman" (English, MiniMax Speech 2.8 HD, Emotion: sad): script: I can't believe it's finally here. After all this time, it's really happening. ## Steps 1. Paste or type your script — Up to 10,000 characters per clip on MiniMax and ElevenLabs Multilingual v2 (5,000 on Eleven v3 and Qwen3). Punctuation shapes the pauses. 2. Pick a voice — Filter 555 preset voices by gender, language and use case, and press play on any voice to preview it. 3. Tune the read — Set speed, pitch and emotion (MiniMax) or stability (ElevenLabs), or add tags like [whispers] and (laughs). 4. Generate and download — Get an MP3 (or WAV/FLAC on MiniMax) with no watermark; every clip is saved to your history. ## In depth ### A short history of text to speech Machines have “spoken” since the 18th century, but electronic speech synthesis took off in the second half of the 20th [1]. Early systems used formant synthesis (rule-based sound generation) or concatenation (stitching together recorded sound units), which is why they sounded mechanical. In 2016 DeepMind's WaveNet showed that a neural network could generate raw audio one sample at a time and sound far more natural than previous systems [2]. Since then, neural models have become standard, and the newest engines add controllable emotion and non-speech sounds. ### Parts of a TTS system Most neural TTS systems have the same building blocks: - Text normalization — expands numbers, dates, symbols and abbreviations. - Linguistic front end — works out pronunciation, including words spelled the same but said differently. - Acoustic model — predicts prosody: pitch, duration and energy for each sound. - Vocoder / decoder — turns those predictions into an audio waveform [3]. ### Where TTS is used Screen readers use TTS so blind and low-vision users can hear interfaces and documents [4]. Beyond accessibility, it powers voice assistants, navigation, phone menus, audiobooks, e-learning and the voice-overs on millions of videos. References: [1] Wikipedia — Speech synthesis: https://en.wikipedia.org/wiki/Speech_synthesis [2] Wikipedia — WaveNet: https://en.wikipedia.org/wiki/WaveNet [3] Wikipedia — Vocoder: https://en.wikipedia.org/wiki/Vocoder [4] Wikipedia — Screen reader: https://en.wikipedia.org/wiki/Screen_reader ## FAQ Q: What does TTS mean? A: TTS stands for text to speech: software that reads written text aloud in a synthetic voice. In games, streams and chat apps, “TTS” often refers to a feature that reads messages aloud. Q: What is neural TTS? A: Neural TTS uses deep-learning models, rather than stitched recordings or hand-written rules, to predict how text should sound and to generate the audio waveform. It's what makes modern voices sound human. Q: How does text to speech work? A: The text is normalized (numbers, abbreviations), converted into a representation of sounds, then an acoustic model predicts pitch, timing and tone for the chosen voice, and a vocoder or decoder generates the audio. Q: What's the difference between text to speech and speech to text? A: Text to speech turns writing into audio. Speech to text (transcription) turns audio into writing. They're opposite directions. Q: How do I use text to speech? A: In Kenerate: open the Voice Studio, paste your text, pick a voice, adjust speed or emotion if you like, and press Generate. Download the MP3. On phones and computers, built-in screen readers also read text aloud. Q: Is text to speech free? A: Built-in system voices on phones and computers are free. For natural AI voices you can download, Kenerate gives free credits on sign-up. Q: How can I do a TTS test? A: Type a sentence into the box at the top of this page and press the button — or try the tricky test sentences above (numbers, homographs, questions) to see how well a voice handles them. Q: Is text to speech generative AI? A: Modern neural text to speech is a form of generative AI: the model generates a new audio waveform for your text instead of playing back stored recordings. Older systems stitched recorded fragments together or followed hand-written rules, which is why they sounded robotic. Engines like ElevenLabs Eleven v3 and MiniMax Speech 2.8 are generative models trained on large amounts of recorded speech. Q: What is a text to speech model? A: A text to speech model is the neural network that turns text into speech audio. Kenerate's Voice Studio runs several: MiniMax Speech 2.8 HD and Turbo, ElevenLabs Eleven v3, Multilingual v2 and Flash v2.5, Qwen3 TTS and Kling V1 TTS. Each has its own voices, languages and character limit, so the same sentence sounds different on each, as the side-by-side demos on this page show. ## Related pages - Text to Speech: https://kenerateai.com/text-to-speech - Speech Time Calculator: https://kenerateai.com/speech-time-calculator - AI Voice Generator: https://kenerateai.com/ai-voice-generator - AI Voice Clone: https://kenerateai.com/ai-voice-clone-free ## Sources - Kenerate Voice Studio (engines, voices and limits from the app's code): https://kenerateai.com/app/voice - MiniMax Speech 2.8 HD on WaveSpeed: https://wavespeed.ai/models/minimax/speech-2.8-hd - ElevenLabs Eleven v3 on WaveSpeed: https://wavespeed.ai/models/elevenlabs/eleven-v3 - ElevenLabs — text to speech documentation: https://elevenlabs.io/docs/overview/capabilities/text-to-speech - Kling V1 TTS on WaveSpeed: https://wavespeed.ai/models/kwaivgi/kling-v1-tts - ElevenLabs — Eleven v3 audio tags: https://elevenlabs.io/blog/v3-audiotags