Guides
Dubbing a video with AI: transcribe, translate, re-voice
Three machine steps and one human one. Skipping the human one is always visible.
Last checked 2026-08-14
In short
Upload the video or its audio; Naadly transcribes it with faster-whisper, translates the transcript into any of 35 languages, re-voices it with a catalogue voice and hands back the audio plus a timed SRT. Lip sync is not attempted - the audio is timed to the subtitle cues, which is correct for narration, interviews and screencasts and visibly wrong for close-up drama.
What each step gets wrong
- Transcription mishears names, jargon and overlapping speech. Fix the transcript before translating; every later error is downstream of this one.
- Translation is grammatical and occasionally absurd, especially with idiom and marketing copy. This is the step that needs a native reader.
- Re-voicing changes duration: German runs long, Japanese runs short. Cues move.
- Timing follows the subtitle cues, so a translated line that ran long is cut into two cues rather than sped up.
Choosing the target voice
Match the register, not the person. A documentary narrator into a bright advert voice destroys the film's tone even when the words are right. Filter the catalogue by the target language and accent, preview two, and dub one minute before you commit an hour.
For an interview, dub the interviewer and the interviewee with different voices. One voice for two people is the cheapest way to make a dub unwatchable.
Subtitles are half the deliverable
The SRT that comes out is timed from the render rather than estimated, so it can be uploaded as-is - and most platforms weight a real subtitle track heavily for reach and accessibility. Keep both the source-language and target-language SRTs; you will need them the next time the video is re-cut.
Where the limits are
Translation is metered per plan and refuses pairs it cannot do well rather than producing mush. Transcription is metered in minutes of upload with a per-job ceiling of an hour, so a feature-length file is several jobs. Indian languages route through the Indic engine and are metered separately.
Step by step
- Upload the audio or video Anything up to an hour per job.
- Read the transcript Fix names and jargon here, before the translation inherits them.
- Choose target language and voice Match the register of the original. Preview one line.
- Render the dub Audio plus a timed SRT come back together.
- Have a native speaker listen One pass. This is the step that separates a dub from an embarrassment.
Questions
- Can AI dubbing do lip sync?
- Naadly does not attempt it. The dub is timed to subtitle cues, which is right for narration, interviews, courses and screencasts, and wrong for close-up dialogue.
- How many languages can you dub into?
- 35, including 20-plus Indian languages through the Indic engine.
- Do I get subtitles?
- Yes - an SRT timed from the actual render, in the target language, beside the audio.
- Will the dubbed audio be the same length as the original?
- Close, not identical: translations change length. Cues shift, which is why you get an SRT rather than a promise.