Use cases
For accessible audio versions
An audio version and a caption file, out of the same render.
Last checked 2026-08-14
In short
If you publish text or video, an audio version and a real caption track are the two accessibility artefacts people actually use. Naadly produces both from one pass - audio in five formats plus an SRT timed from the render - in 35 languages, licensed for distribution on every plan.
What to produce
- An audio version of long-form text, one line per paragraph so it can be corrected.
- Captions for every video, timed from a real render rather than estimated.
- Translated captions where your audience is not monolingual.
- A clear note that the narration is synthetic - which is both honest and increasingly expected.
What this does not replace
A screen reader. Users who rely on assistive technology have their own voice, speed and shortcuts, and they do not want your audio player. Fix your headings, labels and contrast first; an audio version is an addition, not an excuse.
Who actually uses an audio version
- People with low vision or dyslexia who prefer listening to a page over fighting it, but who are not screen-reader users.
- People doing something else - commuting, cooking, walking - which is most of the audience for a long article.
- Second-language readers, for whom hearing and reading together is materially easier than either alone.
- Compliance reviewers, who want to see that captions exist and are accurate rather than auto-generated guesses.
That mix matters because it decides the production choices: a clear voice at a slightly slow pace, paragraph-level lines so a correction is cheap, and captions timed from the real render rather than estimated from the text.
Doing it without creating a maintenance problem
The failure mode is not the audio, it is the drift: the article is edited, the audio is not, and six months later the two disagree. Keep the script as one line per paragraph in a project, so an edit re-renders one line rather than the whole piece, and note the render date next to the player.
- Paste the text, one paragraph per line.
- Put names, acronyms and any product vocabulary into the workspace pronunciation dictionary once.
- Render to WAV, export the SRT from the same job, then encode MP3 for delivery.
- Publish with a line saying the narration is synthetic, and the date it was made.
- Re-render only the changed lines when the text changes.
What it costs to cover a site
Audio runs at roughly 150 words a minute, so a 1,200-word article is about eight minutes: the 30 free minutes a month cover a couple of pieces, and the 6,000-minute Pro plan at $15 covers a weekly publishing schedule with translated versions on top. Captions are produced from the render, so they cost nothing extra.
There is no separate accessibility licence to buy: audio may be published and distributed on every plan, including free, and each voice's own licence and credit line are printed on its page for the ones that ask for attribution.
Questions
- Does an audio version make my site accessible?
- No. Semantic markup, labels, contrast and keyboard support come first; an audio version is a genuine addition on top of that.
- Can I publish the audio?
- Yes - commercial and public distribution is licensed on every plan.