How speech-to-speech translation works
The classic pipeline is three steps: speech recognition, text translation, then text-to-speech. Meta's SeamlessM4T replaces that chain with one model covering about 100 input languages. Voice-preserving tools add a clone of the speaker so the translation still sounds like them.
When to use which
Live conversation with someone in front of you: a phone app's conversation mode. A recording you want to share — a podcast, lecture or voice note — in another language: translate the file and keep the voice.







