Guides
AI voice for podcasts: intros, ads, corrections and full episodes
Use it for the bits nobody wants to record twice.
Last checked 2026-08-14
In short
AI voice earns its place in a podcast at the edges: intros and outros that change every week, dynamically inserted ad reads, corrections you cannot get the guest back for, and translated versions of finished episodes. A fully synthetic conversation is still audibly synthetic, because two machine voices do not listen to each other.
The four jobs worth automating
- Intro and outro. The episode number and title change every week; the read does not.
- Ad reads. One script, one voice, many variants - and re-rendered when the offer changes.
- Corrections. One sentence, in the same voice as the show's narrator, dropped into a published episode.
- Translations. Transcribe, translate, re-voice - a second-language version of a finished episode without re-recording it.
Where listeners notice
Interviews, banter, and anything with a laugh in it. A podcast is sold on the sense that people are actually talking; synthetic conversation reads as a corporate explainer within two exchanges. If you need feeling in a narrated segment, act it and use the voice changer instead.
Transcripts are free reach
Transcription produces a searchable transcript and an SRT from the episode audio, which is show notes, chapter markers and a web page with words on it. Metered in minutes of upload, an hour per job.
A weekly intro in four minutes
- Write the intro as one line per element - show name, episode number, title, guest - in the studio's long-form editor, so next week you retype three words rather than the whole read.
- Pick one voice and keep it. The voice id is stable, so the intro you render in March matches the one from January; favourite it so it is one click away.
- Put the show name in the pronunciation dictionary. Do it once and every render on the workspace inherits it, including the ad reads.
- Render to WAV, not mp3. You are about to master it with music and compression; encode once at the end.
- Keep the project. Next week you duplicate it, change the episode line and re-render only that line.
Per-line rendering is the point: changing the title does not re-render the eleven seconds either side of it, so a weekly intro costs seconds of metered audio rather than a minute.
Ad reads and variants
A sponsor read is the case where synthetic voice is plainly better than a human one: the same script has to exist in four lengths, with three offer codes, re-cut whenever the sponsor changes the URL. Write the base read once, keep each variable on its own line, and render the set as a batch from a CSV - one row per variant, one file out per row.
Two things to hold to. Disclose the synthetic read if your ad network or jurisdiction requires it, because every render carries an AudioSeal watermark and is detectable either way. And do not clone the sponsor's own spokesperson without their written consent - that is a right-of-publicity problem, not a technical one.
What it costs at podcast volume
A weekly show with an intro, an outro, two ad reads and the occasional correction spends roughly two to three minutes of rendered audio per episode - ten to twelve minutes a month, inside the 30 free minutes. Add translated versions of full episodes and you are metering the episode length itself: four 40-minute episodes is 160 minutes, which is why the 6,000-minute Pro plan at $15 a month is the one that fits a translated back catalogue.
Transcription is metered separately from synthesis, in minutes of audio uploaded, so a transcript of an episode you recorded yourself does not spend your synthesis minutes.
Questions
- Should I make a whole podcast with AI voices?
- Narrated and documentary formats work. Conversational ones do not - and the audience will say so.
- Can I fix one sentence in a published episode?
- Yes - render the corrected line in the same voice and splice it. That is much cheaper than getting a guest back.
- Can I get a transcript?
- Yes: upload the episode and get text plus a timed SRT.