Guides
Text to speech with emotion: what actually works, and what it costs
A reading engine cannot act, however good it sounds. Here is the honest version of what can.
Last checked 2026-08-14
In short
Most text-to-speech models, including the fast ones behind large catalogues, have no input for emotion at all: they predict one reading of a sentence, so the only variety you can add is pacing and pauses. Acting needs a different kind of model, and Naadly runs one on 18 voices - the same catalogue people, conditioned on their own recordings - where you can write [laugh] or [sigh] into the script. It takes about three times the length of the line instead of a fiftieth of it, so it is metered in minutes (20 a month on Pro) and used per line rather than per audiobook.
Why the flat read is a model limit, not a settings problem
A fast reading model turns text into one predicted delivery. There is no emotion vector to set, no direction to give, no take two: feed it an exclamation mark and you get slightly different intonation, not excitement. That is the whole explanation for the complaint people make about big voice catalogues, including Naadly's own 1,486, and no amount of speed, pitch or pause tuning changes it.
What tuning does fix is the mechanical evenness: sentences that all run at the same tempo, pauses that are all the same length, a file with no room tone. Naadly's natural delivery does that on every voice, and it is the difference between a robot and a good reader. It is not the difference between a good reader and an actor.
What the expressive lane is
A second, slower model that takes an audio prompt and performs the text against it. For each expressive voice the prompt is 20-25 seconds of the original speaker's own recordings from the dataset the catalogue voice was trained on - not a re-synthesis of them, and never a customer's audio. So it is the same person acting, rather than a new voice that happens to be emotional.
- Sounds you can write in.
[laugh],[chuckle],[sigh],[gasp],[groan],[cough],[sniff],[inhale],[exhale],[clear_throat]. They are performed, not read out. - Its own timing. Speed, gaps and word timings do not apply - the model decides the pacing, which is the point.
- Per line. Up to a paragraph a call, no streaming, no queued audiobooks: it spends a second of a container per second of audio.
- Same rules as everything else. Watermarked with AudioSeal, disclosed, licensed for commercial use, and logged in the same provenance record.
The trade-off nobody else publishes
Making a voice act moves it away from the person it was cloned from. That is measurable, and we measure it: every expressive voice's page shows its speaker similarity to the original recordings next to the reading engine's similarity to the same recordings, so you can see the cost rather than take our word for it. A voice only becomes an expressive edition if it reads accurately, moves measurably more than the flat read, and stays as close to the real speaker as the reading engine does.
That gate is why there are 18 of them and not 1,486. Anyone offering emotion across a whole catalogue of thousands is either running a much more expensive model than they are charging for, or not measuring.
How to use it
- In the voice shop, open the Can act collection, or add
?expressive=1to the catalogue URL. - Play the voice's Hear it act sample before you spend a minute on it - the acted sample and the read sample are the same person.
- In the studio, set Delivery to Expressive and write the sounds into the script where you want them.
- Render a line at a time. If a take is not the one you wanted, render it again - the model is not deterministic, so takes differ, exactly like a session with a human reader.
On the API it is delivery: "expressive" on POST /v1/speak, and GET /v1/expressive lists which voices have it with their measured numbers. Assistants get the same thing through the list_expressive_voices MCP tool.
When to use the acted lane instead
There is a second way to get a real performance into a catalogue voice: record the delivery yourself - or have an actor do it - and convert it onto the voice. The emotion is then a human's, which is still the best emotion available, and it works on far more voices. It costs you a recording and it moves the voice further from the original speaker than the expressive lane does. Both numbers are published; pick by which cost you would rather pay.
Questions
- Can every Naadly voice show emotion?
- No. 18 of 1,486 have an expressive edition, and the voice pages say which. The rest read naturally but do not act.
- Does it sound better than ElevenLabs?
- On acted delivery, a hosted model trained for it is still ahead, and we publish our measured numbers rather than claim otherwise. What is different here is that the expressive voice is the same catalogue person, its likeness is measured on the page, and the file is yours commercially with a licence you can read.
- Why is it metered in minutes when standard speech is not?
- Because it is autoregressive: a minute of audio costs about a minute of a container, where the reading engines cost about a second. Metering the expensive lane is what keeps the cheap one uncapped.
- Can I use it for an audiobook?
- Not in one go: a script goes through /v1/jobs a few thousand characters at a time, because a performance runs on one slot and is metered in minutes. A whole book still belongs on the reading engines with natural delivery.
- Are the laughs real?
- They are generated by the model, not spliced from a sample library, and they are in the voice you picked. Like any take, some are better than others.
