Sign in Start free

Guides

Transcription and subtitles that are actually usable

A transcript is only worth having if the names are right.

Last checked 2026-08-14

In short

Upload audio or video and Naadly transcribes it with faster-whisper, returning text and a timed SRT. Expect good accuracy on clear single-speaker speech, worse on crosstalk, accents at speed and technical vocabulary - and plan to correct names before anything downstream, such as a translation or a dub, inherits them. Metered in minutes of upload, up to an hour per job.

What accuracy actually depends on

  • Recording quality more than anything else. A lapel mic beats a better model on a laptop mic.
  • Crosstalk. Two people at once is where every recogniser fails.
  • Vocabulary. Product names and jargon come out plausible and wrong; a glossary pass fixes them in seconds.
  • Language and accent. Indian-language audio is slower to process and is metered separately.

Making subtitles readable

A machine SRT is timed correctly and broken badly. Two lines maximum, around forty characters each, and never split a phrase across a cue if you can split at a comma instead. Read the first minute after any automatic pass - it tells you what the rest looks like.

From upload to corrected SRT

  1. Upload the audio or video. Video is accepted; the audio track is pulled out for you, and jobs are capped at an hour so one file cannot hold the queue all afternoon.
  2. Set the language if you know it. Letting the model guess costs accuracy on the first few seconds, and guesses badly on code-switched Hindi-English speech.
  3. Read the first minute. It predicts the rest: if the names are wrong at 0:30 they are wrong at 40:00.
  4. Fix proper nouns in the text, not in the SRT - the translation and any dub read the text, so a name corrected here is corrected everywhere downstream.
  5. Export. Plain text for show notes and a page, timed SRT for the player, both from the same job.

What to do with it next

The transcript is the input to the useful things: a translated SRT, a dub in another language, show notes, a searchable page, or a re-voiced version of a video whose original audio was recorded in a car park.

The ordering matters more than people expect. Correct the transcript, then translate, then re-voice: a misheard product name translated into Tamil and then spoken by a synthetic voice is three compounding errors, and only the first one is cheap to fix.

What it is metered as

Transcription is billed in minutes of audio you upload, separately from the minutes you spend synthesising, so transcribing a recording you made yourself never eats into your rendering allowance. Indian-language audio runs on a different engine and is metered on its own.

Uploaded audio is kept only as long as the job needs it, and nothing you upload is used to train a model - see the privacy policy for what is stored and for how long.

Questions

How accurate is AI transcription?
High on clean single-speaker audio, materially lower on crosstalk, heavy accents at speed and unfamiliar jargon. Always correct proper nouns before anything downstream uses the text.
Can I get SRT subtitles?
Yes, timed from the audio, and a translated SRT if you also translate.
How long a file can I upload?
An hour per job - longer recordings go in several jobs, so one upload cannot hold the queue all afternoon.