Guides
How to turn text into speech that does not sound like a robot
Choosing the voice is a tenth of the work. This is the other nine.
Last checked 2026-08-14
In short
Paste the script, pick a voice whose accent and pace match the footage, split it into lines so each one can be re-rendered on its own, fix names with the pronunciation dictionary, then render to MP3 or WAV. On Naadly that takes about a minute: 30 minutes a month are free, no card, and every voice in the catalogue of 1,481 is cleared for commercial use.
What text to speech actually does now
A modern text-to-speech model predicts the sound of a sentence, not the sound of each word in turn. That is why the same voice reads read two ways correctly and why it puts the stress on the right half of a question - and also why it will occasionally read a product name as a word instead of letters, which is the one thing you have to correct by hand.
The practical consequence: judge a voice on a sentence from your own script, never on the demo line. A voice that sounds perfect saying "the quick brown fox" can trip on "our Q3 EBITDA".
Pick a voice by the job, not by the name
Four things decide whether a voice fits, and only one of them is how nice it sounds:
- Accent. An American voice over British footage is the fastest way to look outsourced. The catalogue filters by accent as well as language.
- Pace. Explainers want deliberate; adverts want quick. Every Naadly voice is labelled with its measured words per second, so this is a filter rather than a guess.
- Grade. Voices are graded from the transcription error of a probe sentence. A C-grade voice is fine for a draft and wrong for a paid advert.
- Licence. Some open voices forbid commercial use. Naadly only lists ones that allow it, and prints the licence and the attribution on each voice's page.
Use the Voice Finder if you would rather answer three questions than scroll a catalogue: language, what you are making, and the delivery you want.
Write for the ear
Text-to-speech exposes bad writing. A sentence a reader skims is a sentence a listener loses. What consistently helps:
- One idea per sentence, and full stops where you would take a breath - punctuation is the only pause control that works in every engine.
- Numerals spelled out when they are meant to be heard as words ("twenty twenty-six" not "2026") and left as digits when they are meant to be heard as digits.
- Acronyms spaced or dotted ("A P I") if you want them read as letters - or, better, entered once in the pronunciation dictionary so every future script inherits the fix.
- No em-dash asides. A voice reads them as a full stop and the sentence loses its spine.
Render, listen, re-render one line
The part most tools get wrong is the second draft. If a script is one blob of text, fixing the fourth sentence means re-rendering all forty and paying for all forty. Naadly splits a script into lines, renders each one separately and lets you re-render a single line, change its voice, or nudge its speed without touching the rest, then stitches the finished lines into one file.
Formats: MP3 for video and the web, WAV for an editor that will process the audio again, FLAC to archive, OGG for the web, u-law for a phone system. Subtitles come out beside the audio as SRT, timed from the render rather than guessed from the word count.
Know what you may do with the file
Two separate questions, and most tools answer only the first. The service: on any Naadly plan, including free, the audio you render is yours to use commercially - adverts, films, courses, clients. The voice model: the open models behind the catalogue carry their own licences, mostly CC BY 4.0 and Apache-2.0, and a few of them ask for attribution. Each voice page prints its licence and the credit line to copy.
Every rendered file also carries an inaudible AudioSeal watermark and is served with a header declaring it synthetic - which is increasingly what a platform, a broadcaster or a regulator asks for.
Step by step
- Paste the script into Studio Each paragraph becomes a line you can render and re-render on its own. A CSV import does the same for hundreds of lines at once.
- Pick a voice Filter by language, accent, gender, pace and tone, or answer three questions in the Voice Finder. Preview before you spend a minute of quota.
- Teach it your names Add product names, place names and acronyms to the pronunciation dictionary once; every render on the account uses them afterwards.
- Render and listen Render the lines, listen, and re-render only the ones that are wrong. Change a line's voice or pace without touching its neighbours.
- Export Download the stitched MP3 or WAV, and the SRT if the audio is going over video.
Questions
- Is text to speech free on Naadly?
- 30 minutes of audio a month are free without a card, including commercial use of what you render. Paid plans start at $15 a month for 6,000 minutes.
- Which format should I export?
- MP3 for anything going straight into video or the web; WAV if an editor or a mastering chain will touch the audio again; u-law only for telephony.
- Can I use the audio in a YouTube video that earns money?
- Yes. The audio is licensed for commercial use on every plan. Check the voice's page for an attribution line if its model asks for one.
- Why does one voice mispronounce a name that another gets right?
- The voices come from different models with different pronunciation front-ends. Fix it once in the pronunciation dictionary and it holds for every voice on the account.