Skip to content

Plain-English guide · live demos

What is text to speech (TTS)?

TTS turns written text into spoken audio. Here's how neural text to speech works, what it's used for, and how four of today's engines sound reading the same sentence — press play and hear the difference.

Try
Particles rise from an open book and ripple into a sound wave
  • Hear 4 engines side by side
  • Test with your own text
  • Free credits
  • No watermark

Updated by the

Quick answer

Updated

What is text to speech?

Text to speech (TTS) is technology that converts written text into spoken audio. Modern neural TTS normalizes the text, predicts the sounds, rhythm, pitch and pauses of a voice, then generates the audio waveform. It powers screen readers, voice assistants, audiobooks, voice-overs and phone menus, and today's best engines can add emotion, whispers and laughter.

Key facts

TTS stands for
Text to speech
Also called
Speech synthesis
Neural TTS
Deep-learning models that generate the waveform
Main uses
Accessibility, voice-overs, audiobooks, assistants, IVR
Opposite of
Speech to text (transcription)
Leading engines
ElevenLabs Eleven v3, MiniMax Speech 2.8

Listen first

One line, two emotions

Neural TTS can change how a line feels without changing the words. Same voice, same script — only the emotion setting differs.

Upbeat Woman

MiniMax Speech 2.8 HD · MiniMax

EnglishHappy

I can't believe it's finally here. After all this time, it's really happening.

Emotion: happy

0:05
Use this voice

Upbeat Woman

MiniMax Speech 2.8 HD · MiniMax

EnglishSad

I can't believe it's finally here. After all this time, it's really happening.

Emotion: sad

0:05
Use this voice

Same script

Four TTS engines, one sentence

The same sentence read by MiniMax, ElevenLabs, Qwen3 and Kling. Listen for pacing, emphasis and breaths.

Text to speech turns written words into spoken audio. Modern systems predict not just the sounds, but the rhythm, the pauses and the emotion of a real speaker.
  • MiniMax

    Expressive Narrator · MiniMax Speech 2.8 HD

    10.8 s
  • ElevenLabs

    Roger · ElevenLabs Eleven v3

    12.1 s
  • Qwen3

    Aiden · Qwen3 TTS

    11.3 s
  • Kling

    Calm Storyteller · Kling V1 TTS

    15.6 s

Real outputs with default settings. Bar length is the real clip length — engines choose different speaking rates for the same text.

How to

How to use Kenerate Voice Studio

Time needed
Steps
4 steps
  1. Step 1

    Paste or type your script

    Up to 10,000 characters per clip on MiniMax and ElevenLabs Multilingual v2 (5,000 on Eleven v3 and Qwen3). Punctuation shapes the pauses.

  2. Step 2

    Pick a voice

    Filter 555 preset voices by gender, language and use case, and press play on any voice to preview it.

  3. Step 3

    Tune the read

    Set speed, pitch and emotion (MiniMax) or stability (ElevenLabs), or add tags like [whispers] and (laughs).

  4. Step 4

    Generate and download

    Get an MP3 (or WAV/FLAC on MiniMax) with no watermark; every clip is saved to your history.

Kenerate AI Voice Studio text to speech script box with a narration script typed in

Open Voice Studio in Text to speech and type the words you want spoken, or start from a script template.

1 / 5 · Write or paste your script

Watch the tour · 1:41

Text to speech, emotions & tags, Hindi, voice design, cloning & chat

Chapters

How it works

How neural text to speech works

Five stages turn characters into sound. The middle one — the acoustic model — is where the leap in quality happened.

  1. 01

    Text

    Your script, with punctuation, numbers and any tags like [whispers] or (laughs).

  2. 02

    Normalization

    “4.5 kg” becomes “four point five kilograms”, “Dr.” becomes “doctor”, dates and units are spelled out.

  3. 03

    Acoustic model

    A neural network predicts sounds, rhythm, pitch and pauses for the chosen voice — the part that makes it sound human.

  4. 04

    Vocoder / decoder

    Turns that prediction into an actual waveform, thousands of samples per second.

  5. 05

    Audio

    An MP3 or WAV you can play, download and drop into a video, course or podcast.

Use cases

What text to speech is used for

From accessibility to entertainment.

  • Screen readers & accessibility

    Reading interfaces and documents aloud for blind and low-vision users.

  • YouTube & Shorts voice-overs

    Faceless channels, explainers and recaps — keep one voice across every upload.

  • Audiobooks & stories

    Long reads up to 10,000 characters per clip on MiniMax; split chapters into clips and keep the same voice.

  • E-learning & training

    Course narration that stays consistent when you edit a slide next month.

  • Ads & product demos

    Short, punchy reads; try three voices in a minute and keep the one that sells.

  • Phone menus & voicemail

    Clear greetings and IVR prompts in the languages your callers speak.

Engines

Today's TTS engines, compared

The four engine families in the Kenerate Voice Studio. Facts from the app's code.

  • MiniMax Speech 2.8 HD title card — a high-fidelity studio microphone with seven waves in seven emotion colours radiating from it

    MiniMax

    MiniMax Speech 2.8 HD

    Languages
    40 languages
    Per clip
    10,000 characters per clip

    Most natural MiniMax voice; 7 emotions, (laughs)/(sighs) interjections, pronunciation dictionary

    Open MiniMax Speech 2.8 HD
  • ElevenLabs Eleven v3 title card — a fictional adult voice actor's silhouette in a recording booth

    ElevenLabs

    ElevenLabs Eleven v3

    Languages
    70 languages
    Per clip
    5,000 characters per clip

    Most expressive; audio tags like [whispers], [laughs], [excited]

    Open ElevenLabs Eleven v3
  • ElevenLabs Multilingual v2 title card — a long

    ElevenLabs

    ElevenLabs Multilingual v2

    Languages
    29 languages
    Per clip
    10,000 characters per clip

    Lifelike and consistent over long reads

    Open ElevenLabs Multilingual v2
  • Qwen3 TTS title card — nine glowing characterful voice orbs circling a vintage ribbon microphone in a dark room

    Alibaba Qwen

    Qwen3 TTS

    Languages
    10 languages
    Per clip
    5,000 characters per clip

    9 characterful voices plus a free-text style prompt

    Open Qwen3 TTS
  • Kling V1 TTS title card — a cast of stylised character silhouettes

    Kuaishou Kling

    Kling V1 TTS

    Languages
    English & Chinese
    Per clip
    512 characters per clip

    46 character and dialect voices for short lines

    Open Kling V1 TTS

Controls

What you can control in modern TTS

Neural engines expose settings that older systems never had.

  • Speed · 0.5× – 2×

    MiniMax voices. 1× is the natural pace; 0.85–0.95× suits narration and learners, 1.1–1.2× suits ads and recaps.

  • Pitch · −12 to +12

    MiniMax voices, in semitones. A few steps down sounds bigger and older; a few up sounds lighter and younger.

  • Emotion · 7 moods

    MiniMax: neutral, happy, sad, angry, fearful, disgusted, surprised — the same words, a different read.

  • Interjections & tags · (laughs) · [whispers]

    MiniMax 2.8 reads 13 cues like (laughs), (sighs), (breath); ElevenLabs Eleven v3 reads audio tags like [whispers], [excited], [sarcastic].

  • Stability · 0 – 1

    ElevenLabs. Lower = more expressive and varied between takes; higher = steadier and more even.

  • Output · MP3 · WAV · FLAC

    MiniMax lets you pick the format and sample rate (16–44.1 kHz); every clip downloads without a watermark.

Scripts

Good sentences for a TTS test

These trip up weaker engines. Try them on any voice.

  • Homographs

    I read the book yesterday, and I'll read it again tomorrow. The wind was too strong to wind the sail.

  • Numbers & symbols

    Call 555-0142 before 9:30 a.m. — tickets are $19.99, or 2 for $35.

  • Questions vs statements

    You finished it already? You finished it already.

  • Emotion tags (Eleven v3)

    [sighs] Okay. [excited] Wait — we actually won? [laughs] No way!

Tips

How to use text to speech well

Whatever tool you use, the same habits help.

  • Write for listening: short sentences, one idea each.

  • Punctuation is the pause control — commas, full stops, ellipses.

  • Spell out anything ambiguous: abbreviations, years, units.

  • Choose the voice for the job, not just the one that sounds nicest.

  • Test with tricky lines (numbers, names) before generating a long script.

  • Label synthetic narration where your platform asks you to.

Example uses

How creators use What Is text to speech

Illustrative examples of typical workflows, not customer reviews.

  • Braille display, keyboard and headphones in front of a high-contrast laptop screen

    Accessibility tester

    An accessibility tester can hear how documents and interfaces sound when read aloud and catch numbers or abbreviations a voice reads wrong.

  • Streaming setup with a boom microphone, headphones and monitors with purple LED glow

    Streamer

    A streamer can understand how chat TTS works and pick a natural voice to read viewer messages aloud.

  • Learner seen from behind at a cafe table with earbuds and a tablet showing a waveform

    Language learner

    An adult language learner can type tricky sentences and listen to how a neural voice pronounces them, then repeat after it.

FAQ

What Is Text to Speech questions

What does TTS mean?

TTS stands for text to speech: software that reads written text aloud in a synthetic voice. In games, streams and chat apps, “TTS” often refers to a feature that reads messages aloud.

Link to this answer

What is neural TTS?

Neural TTS uses deep-learning models, rather than stitched recordings or hand-written rules, to predict how text should sound and to generate the audio waveform. It's what makes modern voices sound human.

Link to this answer

How does text to speech work?

The text is normalized (numbers, abbreviations), converted into a representation of sounds, then an acoustic model predicts pitch, timing and tone for the chosen voice, and a vocoder or decoder generates the audio.

Link to this answer

What's the difference between text to speech and speech to text?

Text to speech turns writing into audio. Speech to text (transcription) turns audio into writing. They're opposite directions.

Link to this answer

How do I use text to speech?

In Kenerate: open the Voice Studio, paste your text, pick a voice, adjust speed or emotion if you like, and press Generate. Download the MP3. On phones and computers, built-in screen readers also read text aloud.

Link to this answer

Is text to speech free?

Built-in system voices on phones and computers are free. For natural AI voices you can download, Kenerate gives free credits on sign-up.

Link to this answer

How can I do a TTS test?

Type a sentence into the box at the top of this page and press the button — or try the tricky test sentences above (numbers, homographs, questions) to see how well a voice handles them.

Link to this answer

Is text to speech generative AI?

Modern neural text to speech is a form of generative AI: the model generates a new audio waveform for your text instead of playing back stored recordings. Older systems stitched recorded fragments together or followed hand-written rules, which is why they sounded robotic. Engines like ElevenLabs Eleven v3 and MiniMax Speech 2.8 are generative models trained on large amounts of recorded speech.

Link to this answer

What is a text to speech model?

A text to speech model is the neural network that turns text into speech audio. Kenerate's Voice Studio runs several: MiniMax Speech 2.8 HD and Turbo, ElevenLabs Eleven v3, Multilingual v2 and Flash v2.5, Qwen3 TTS and Kling V1 TTS. Each has its own voices, languages and character limit, so the same sentence sounds different on each, as the side-by-side demos on this page show.

Link to this answer

A short history of text to speech

Machines have “spoken” since the 18th century, but electronic speech synthesis took off in the second half of the 20th [1]. Early systems used formant synthesis (rule-based sound generation) or concatenation (stitching together recorded sound units), which is why they sounded mechanical.

In 2016 DeepMind's WaveNet showed that a neural network could generate raw audio one sample at a time and sound far more natural than previous systems [2]. Since then, neural models have become standard, and the newest engines add controllable emotion and non-speech sounds.

By Last updated and reviewed

We build and run the Kenerate Voice Studio. Every demo on this page is a real output of the engine and voice named on it, made from the script shown; engine facts were checked against the app's code.

Sources: Kenerate Voice Studio (engines, voices and limits from the app's code) · MiniMax Speech 2.8 HD on WaveSpeed · ElevenLabs Eleven v3 on WaveSpeed · ElevenLabs — text to speech documentation · Kling V1 TTS on WaveSpeed · ElevenLabs — Eleven v3 audio tags

Try Kenerate Voice Studio

Free to start. No watermark. Download every clip.

Tried it? Tell us how it went