Notes
Why we meter minutes instead of characters
You should be able to render a line eleven times to get it right.
Published 2026-08-10 · Last checked 2026-08-14
In short
Services that resell a third-party model meter characters because that is what they are billed. Naadly runs open models on its own containers, so its marginal cost is machine time - which makes minutes the honest unit and makes retries cheap enough to stop counting. That is also why Indian-language minutes are metered separately: that model really is twenty times heavier.
What a render actually costs us
Piper renders many times faster than realtime on a small container. Kokoro needs about 700 MB per render, so a 1 GB container fits exactly one at a time. The Indic model runs at roughly 1.25x realtime on a container three sizes larger. Those three facts are the entire pricing model: standard speech can be uncapped, Indic minutes cannot.
What per-character pricing does to a workflow
It is not that per-character is dishonest - for a reseller it is the only unit that matches the bill. The problem is what it teaches you to do. Every take is charged, so the fourth attempt at a line feels expensive, and the rational move is to accept a read you do not like. The pricing quietly sets the quality ceiling of the work.
Per-character also prices things you did not hear: a script submitted with a typo, a batch you cancelled, a chapter rendered in the wrong voice. Minutes of output cost nothing for the attempt and everything for the result, which is the way round that matches how people actually produce audio.
What that means for how you work
- Re-render one line as many times as it takes.
- Audition six voices on your real script rather than the demo line.
- Render a whole book twice - the second pass is where it gets good.
- Never do arithmetic to decide whether to fix a sentence.