Answers
How much audio do you need to clone a voice?
Last checked 2026-08-14
In short
About 30 seconds of clean, single-speaker speech. Recording quality matters far more than length: a quiet room, a consistent distance from the microphone and normal delivery beat ten minutes of noisy audio. Naadly also requires a spoken consent statement from the speaker before it will create the clone.
What makes a reference fail
- Room echo - copied faithfully, and the usual reason a clone sounds thin.
- Music or a second speaker underneath.
- An unusual delivery: shout into it, shout out of it.
- Phone-quality audio, which limits the clone to phone quality.
How to record 30 seconds properly
- Find the least reflective room you have - soft furniture, curtains, not a kitchen or a stairwell.
- Sit a hand's width from the microphone and stay there; moving about changes the timbre mid-reference.
- Read something ordinary and continuous - a paragraph of prose, not a word list - at your normal speaking pace.
- Listen back on headphones. If you can hear the room, the clone will have the room in it.
More audio does not fix a bad recording; it averages it in. Thirty clean seconds beats ten noisy minutes, every time.
The consent step is not optional
Before a clone is created the speaker records a short spoken consent statement, which is stored with a hash of the reference clip and kept even if the workspace is later deleted. That is deliberate: it protects the person whose voice it is, and it means you can answer a client who asks whose voice they are hearing.