How lip sync models work
The model reads the audio — ByteDance's LatentSync uses Whisper audio features — and regenerates the mouth region frame by frame to match, keeping the rest of the face and the scene untouched.
Dub, then sync
For a translated video, dub the audio first (Kenerate's dubbing tool does 33 languages), then lip sync the video to the new track. The demo on this page is exactly that pipeline.







