Hear it yourself
One script, every way this service can say it, and the real person's own recording underneath it as the ceiling. Every file downloads. Nothing here is rendered when you press play — these are the files, rendered once on 2026-08-25, so they cannot be quietly re-rolled until they flatter us.
The script
Oh - [laugh] you actually did it. I told them you would, and nobody believed me. [sigh] Right. Let's see what it sounds like.
Chosen because it has a laugh in it, a sigh, and two sentences that need different speeds. A line that only needs to be pronounced correctly makes every engine sound the same.
Kaleb Danforth
Word error 0.0% · likeness to the real person 0.70 expressive against 0.70 for the read — close, not closer, which is why neither is sold as the same person.
-
The real person 25.1s Download0:00···
The speaker's own recording from the public corpus - not synthesis at all. It is here as the ceiling, and it is a different passage, because it is the audio no engine was allowed to hear.
This clip is a passage of their own from the corpus, unheard by any engine here.
-
Flat read 4.6s Download0:00···
The engine with the realism pass switched off: one tempo, no rests, no mastering. This is what the catalogue sounded like before, and what most open voices still sound like.
-
Naadly read 4.7s Download0:00···
The same engine and the same voice with the realism pass on - sentence tempo, punctuation rests, room tone, mastering. Included on every plan, and about a fiftieth of realtime.
-
Naadly expressive 8.0s Download0:00···
The second engine, prompted with the same person's real recordings, so it acts the line instead of reading it. The qualified voices only - listed at /voices?expressive=1 - and about the length of the line to render.
Iris Thornton
Word error 0.0% · likeness to the real person 0.79 expressive against 0.68 for the read — close, not closer, which is why neither is sold as the same person.
-
The real person 21.0s Download0:00···
The speaker's own recording from the public corpus - not synthesis at all. It is here as the ceiling, and it is a different passage, because it is the audio no engine was allowed to hear.
This clip is a passage of their own from the corpus, unheard by any engine here.
-
Flat read 4.7s Download0:00···
The engine with the realism pass switched off: one tempo, no rests, no mastering. This is what the catalogue sounded like before, and what most open voices still sound like.
-
Naadly read 4.8s Download0:00···
The same engine and the same voice with the realism pass on - sentence tempo, punctuation rests, room tone, mastering. Included on every plan, and about a fiftieth of realtime.
-
Naadly expressive 8.0s Download0:00···
The second engine, prompted with the same person's real recordings, so it acts the line instead of reading it. The qualified voices only - listed at /voices?expressive=1 - and about the length of the line to render.
Daphne Hayward
Word error 0.0% · likeness to the real person 0.78 expressive against 0.78 for the read — close, not closer, which is why neither is sold as the same person.
-
The real person 25.0s Download0:00···
The speaker's own recording from the public corpus - not synthesis at all. It is here as the ceiling, and it is a different passage, because it is the audio no engine was allowed to hear.
This clip is a passage of their own from the corpus, unheard by any engine here.
-
Flat read 4.7s Download0:00···
The engine with the realism pass switched off: one tempo, no rests, no mastering. This is what the catalogue sounded like before, and what most open voices still sound like.
-
Naadly read 5.0s Download0:00···
The same engine and the same voice with the realism pass on - sentence tempo, punctuation rests, room tone, mastering. Included on every plan, and about a fiftieth of realtime.
-
Naadly expressive 8.9s Download0:00···
The second engine, prompted with the same person's real recordings, so it acts the line instead of reading it. The qualified voices only - listed at /voices?expressive=1 - and about the length of the line to render.
Wes Whitaker
Word error 0.0% · likeness to the real person 0.73 expressive against 0.61 for the read — close, not closer, which is why neither is sold as the same person.
-
The real person 23.0s Download0:00···
The speaker's own recording from the public corpus - not synthesis at all. It is here as the ceiling, and it is a different passage, because it is the audio no engine was allowed to hear.
This clip is a passage of their own from the corpus, unheard by any engine here.
-
Flat read 4.8s Download0:00···
The engine with the realism pass switched off: one tempo, no rests, no mastering. This is what the catalogue sounded like before, and what most open voices still sound like.
-
Naadly read 5.2s Download0:00···
The same engine and the same voice with the realism pass on - sentence tempo, punctuation rests, room tone, mastering. Included on every plan, and about a fiftieth of realtime.
-
Naadly expressive 8.9s Download0:00···
The second engine, prompted with the same person's real recordings, so it acts the line instead of reading it. The qualified voices only - listed at /voices?expressive=1 - and about the length of the line to render.
The blind test
Numbers cannot tell you whether a voice is pleasant to listen to, and we cannot mark our own homework. So: two unlabelled clips, one question, and the running tally published below however it turns out.
- Which one sounds more like a person, rather than a machine reading? 2 answers so far — not enough to publish a percentage (20 needed)
- Which one is the same person as the reference? 3 answers so far — not enough to publish a percentage (20 needed)
Why there is no rival's audio on this page
Because we would be hosting their output, which their terms cover and ours cannot, and because their models change without notice — a clip we posted this month would be a straw man by the next. What we compare instead is what they publish: the price per finished hour and the feature-by-feature reading, with the date we read it. Their own demos are a click away and you should listen to them. On acted emotion, they are still ahead of us; we say so on that page too.
Where the human recordings come from
The human recordings are held-out clips of the same speakers from LibriTTS-R, a public research corpus of read audiobooks, used under CC BY 4.0: openslr.org/141. The same licences are listed in full on the licences page.
