# Text to Speech > Text to speech turns written text into natural, spoken audio; Kenerate AI Text to Speech does it with 555 voices from four AI makers, for videos, audiobooks, lessons and ads. Paste up to 10,000 characters per clip, pick a voice, set speed, pitch and emotion, and download MP3, WAV or FLAC in up to 70 languages, depending on the engine. URL: https://kenerateai.com/text-to-speech Publisher: Kenerate AI (https://kenerateai.com) Last updated: 29 September 2026 ## Key facts - Voices: 555 presets (MiniMax 459, Kling 46, ElevenLabs 41, Qwen3 9) - Engines: 16 across MiniMax, ElevenLabs, Qwen3, Kling - Languages: Up to 70 (Eleven v3); 40 on MiniMax - Per clip: Up to 10,000 characters - Formats: MP3 · WAV · FLAC - Free start: Starter credits, no watermark - Emotions: 7 (MiniMax) - Cues: [whispers], (laughs) and more - Demo "Alicia" (English, ElevenLabs Eleven v3): script: [excited] Welcome to the show! Today we're taking a plain block of text and turning it into a voice you'd swear was recorded in a studio. [whispers] And nobody has to know. - Demo "Sage Woman" (English, MiniMax Speech 2.8 HD): script: The lighthouse had not been lit in forty years. Yet every night at nine, a single window glowed at the top of the tower, and the fishermen stopped talking as they rowed past. - Demo "George" (English, ElevenLabs Multilingual v2): script: Good evening. Tonight's forecast: light rain across the west, clearing by midnight, with a crisp, bright start tomorrow morning. - Demo "Comedian" (English, MiniMax Speech 2.8 HD): script: So I tried to build the bookshelf without the instructions. (laughs) Three hours later, I had a very nice... chair. (sighs) The bookshelf is still in the box. - Demo "Ryan" (English, Qwen3 TTS, Style: "calm late-night podcast host, relaxed, warm and unhurried"): script: You're listening to Night Shift, the show for people who do their best thinking after midnight. Grab something warm to drink. We're in no hurry. - Demo "Explanatory Man" (English, MiniMax Speech 2.8 HD): script: Here's the short version. Solar panels turn sunlight into direct current. An inverter turns that into the alternating current your home uses, and anything you don't use can flow back to the grid. - Demo "Global Male" (English, Kling V1 TTS): script: Fresh coffee, ready before you are. Order ahead, skip the line, and pick it up on your way. Morning, sorted. ## Steps 1. Paste or type your script — Up to 10,000 characters per clip on MiniMax and ElevenLabs Multilingual v2 (5,000 on Eleven v3 and Qwen3). Punctuation shapes the pauses. 2. Pick a voice — Filter 555 preset voices by gender, language and use case, and press play on any voice to preview it. 3. Tune the read — Set speed, pitch and emotion (MiniMax) or stability (ElevenLabs), or add tags like [whispers] and (laughs). 4. Generate and download — Get an MP3 (or WAV/FLAC on MiniMax) with no watermark; every clip is saved to your history. ## In depth ### How neural text to speech works Speech synthesis is the artificial production of human speech [1]. Older systems stitched together recorded fragments or used rule-based formant models, which is where the familiar “robot” sound came from. Since WaveNet showed in 2016 that a neural network could generate raw audio waveforms sample by sample, neural models have taken over [2]. Today's engines first normalize the text (numbers, abbreviations, dates), then predict rhythm, pitch and timing for the chosen voice, then decode that into audio. That's why punctuation and spelled-out numbers change the result so much — they're the model's only clues about how you want it said. ### Choosing between the engines Each maker has a different strength, and the studio lets you switch per clip: - ElevenLabs Eleven v3 — the most expressive; follows audio tags like [whispers] and [laughs] [3]. - ElevenLabs Multilingual v2 — lifelike and steady over long reads in 29 languages. - MiniMax Speech 2.8 HD — 459 voices, 40 languages, 7 emotions, interjections and a pronunciation dictionary. - Qwen3 TTS — 9 characterful voices you can steer with a free-text style prompt. - Kling V1 TTS — 46 character and dialect voices for short lines. ### Writing scripts that sound spoken Text written for reading rarely sounds right read aloud. Keep sentences short, put the key word at the end of the sentence, and read the draft out loud once before you generate. The same habits that help a human voice-over artist make AI voices sound more natural — the model can only perform the rhythm your punctuation gives it. If you need a particular character, a trailer growl or a voice that doesn't exist yet rather than a clean read of your script, the AI voice generator page covers voices by role, Voice Design and cloning. References: [1] Wikipedia — Speech synthesis: https://en.wikipedia.org/wiki/Speech_synthesis [2] Wikipedia — WaveNet: https://en.wikipedia.org/wiki/WaveNet [3] ElevenLabs — Eleven v3 audio tags: https://elevenlabs.io/blog/v3-audiotags ## FAQ Q: What is the best free text to speech? A: Kenerate gives you free credits on sign-up to try every engine — MiniMax Speech 2.8, ElevenLabs Eleven v3 and Multilingual v2, Qwen3 and Kling — on the same script, with no watermark. The best voice depends on the job: ElevenLabs v3 is the most expressive, MiniMax 2.8 HD adds emotions and a pronunciation dictionary, Qwen3 takes a free-text style prompt. Q: How many voices and languages are there? A: 555 preset voices: 459 on MiniMax, 46 on Kling, 41 on ElevenLabs and 9 on Qwen3. ElevenLabs Eleven v3 reads 70 languages, MiniMax 40, ElevenLabs Flash v2.5 32, Multilingual v2 29 and Qwen3 10. Q: How long can the text be? A: Up to 10,000 characters per clip on MiniMax and ElevenLabs Multilingual v2, 5,000 on Eleven v3 and Qwen3, and 512 on Kling. For longer pieces, generate one section per clip with the same voice and settings. Q: Can I download the audio as MP3? A: Yes. Every clip downloads as MP3; MiniMax voices can also output WAV or FLAC at 16–44.1 kHz. There is no watermark on the audio. Q: Can the voice laugh, whisper or sound emotional? A: Yes. MiniMax 2.8 has 7 emotions (happy, sad, angry and more) and reads 13 interjections like (laughs) and (sighs). ElevenLabs Eleven v3 follows audio tags like [whispers], [excited] and [sarcastic]. Q: Can I use the voices in YouTube videos? A: Yes — Kenerate's Terms give you ownership of what you generate, so voice-overs can go into videos, courses and ads. Don't use a voice to impersonate a real person or mislead listeners. Q: Can I clone my own voice? A: Yes. Instant clone copies a voice from a 3–15 second sample; you can also save a clone to reuse on MiniMax. Only clone voices you have permission to use. See the AI voice clone page for demos. Q: Is this the same as a screen reader? A: No. A screen reader reads your device's interface aloud in real time. Kenerate generates high-quality audio files from a script you write — for videos, courses, audiobooks and announcements. Q: Is there a free text to speech tool online? A: Yes. Kenerate's text to speech runs in your browser: a free account starts with free credits, enough to try every engine on your own script, and every clip downloads as MP3 with no watermark. Paste up to 10,000 characters per clip, pick one of 555 voices and generate. There's nothing to install, and credits never expire. Q: Is there a text to speech app for Android or iPhone? A: Kenerate doesn't have a separate app: the Voice Studio runs in the browser on Android phones, iPhones, tablets and computers. Open kenerateai.com, paste your text, pick a voice and download the MP3 to your phone. To have your phone's screen read aloud as you use it, turn on the screen reader built into Android or iOS instead. ## Related pages - AI Voice Generator: https://kenerateai.com/ai-voice-generator - AI Voice Clone: https://kenerateai.com/ai-voice-clone-free - What Is Text to Speech?: https://kenerateai.com/what-is-text-to-speech - Robot Voice Generator: https://kenerateai.com/robot-voice-generator ## Sources - Kenerate Voice Studio (engines, voices and limits from the app's code): https://kenerateai.com/app/voice - MiniMax Speech 2.8 HD on WaveSpeed: https://wavespeed.ai/models/minimax/speech-2.8-hd - ElevenLabs Eleven v3 on WaveSpeed: https://wavespeed.ai/models/elevenlabs/eleven-v3 - ElevenLabs — text to speech documentation: https://elevenlabs.io/docs/overview/capabilities/text-to-speech - Qwen3 TTS Voice Design on WaveSpeed: https://wavespeed.ai/models/wavespeed-ai/qwen3-tts/voice-design - Kling V1 TTS on WaveSpeed: https://wavespeed.ai/models/kwaivgi/kling-v1-tts