Sign in Start free

Run it on your own hardware

Every model this service speaks with is open-weight and commercially licensed, and there is no third-party AI key anywhere in it. So the whole thing — catalogue, studio, API, queue, watermarking, consent records — can run inside your network, on machines you can point at. A hosted competitor cannot sell you that, whatever the contract says.

Who this is for

Three sorts of buyer, in our experience: the one whose scripts are not allowed to leave the building, the one whose per-character bill has stopped being funny, and the one who needs speech to keep working when the internet does not. Nothing about the product changes — a script written against the hosted API runs against your deployment by changing the base URL.

What a machine needs

No GPU. None of the engines need one, and none of the numbers on this page assume one — which is the part procurement usually does not believe.

The services, and the memory each one measured

Sizing a deployment by guess is how it fails in week two, so these are the figures from our own production measurements (status shows the same services running).

ServiceEngineMemoryNeeded
Web application and queue
The shop, the studio, the API, the job queue and the provenance records. Object storage can be S3, R2, MinIO or a mounted disk.
FastAPI, PostgreSQL, an object store 1-2 GB always
Reading voices
1,480+ voices at 21-43x realtime depending on how many slots the container is given; ~100 ms to first audio byte. This is the lane almost all volume goes through.
Piper (MIT) 512 MB-1 GB always
Natural voices
Heavier and more natural: ~700 MB per render, 8.4x aggregate realtime with two slots.
Kokoro (Apache-2.0) 1 GB recommended
Japanese
A separate service because the Japanese front-end - the G2P that makes Japanese worth shipping - lives in it.
piper-plus (MIT) 512 MB if Japanese
Indian languages
22 Indian languages plus Indian English at 1.25x realtime, so Indic minutes are the expensive ones: about 0.8 CPU-hours per audio-hour.
Indic-Mio Q8 (Apache-2.0) 4 GB if Indic
Expressive delivery
Autoregressive, so about the length of the line on four threads, and 86 s to load. The qualified voices only.
Chatterbox-Nano (MIT) 8 GB if expressive
Cloning and performed delivery
torch floors this container at ~820 MB before a request arrives; peak 1.44 GB under concurrency. Consent records are enforced by the application, not by policy.
kNN-VC, OpenVoice V2 4 GB if cloning
Transcription and translation
3.6x realtime on English, 1.8x on Hindi, ~1 GB peak beside the two translation models. Hosted, this container is started per job and stopped again; on your hardware it can simply stay up.
faster-whisper small, Marian/M2M 2 GB if dubbing

What a licence includes

What it does not

Why there is no price on this page

Because the honest number depends on things we do not know yet, and a headline price would have to assume the worst of them. It is quoted per deployment, on:

Hosted pricing is published in full on pricing, and most buyers should start there — a self-hosted deployment is worth it when the data boundary or the volume makes it worth it, not by default.

Ask for a quote

Kept to answer you and nothing else — see the privacy policy. Or write to [email protected] if a form is not your thing.