Voice is free. Here is how to use it like it matters.
Narration costs nothing per take, which changes how you should write it. Pacing, takes, and laying voice against picture.
- Written by
- The Cetus team
- Seeds a claim
- 5
- Models
- 3

On this page4
Voice generation in Cetus costs nothing. Not a promotion — it runs on a public synthesis endpoint and we do not pay per call, so charging for it would be charging for something that costs us nothing. The interesting consequence is not the saving. It is that a free thing should be used differently from a metered one.
Generate six takes, not one
With a metered model you write carefully and generate once. With a free one the right move is the opposite: write three versions of the line, generate two takes of each, listen to all six, keep one. It takes about ninety seconds and it is the single biggest quality difference available in this product for zero credits.
Most people do not do this because they are carrying a habit over from tools where every generation had a price. Drop the habit here.
Write for the ear, then for the clock
- Short sentences. Synthesis handles a full stop far better than a comma, and a long subordinate clause is where the prosody falls apart.
- No parentheses, no dashes mid-sentence. Break them into separate sentences instead; you will get the pause you wanted and it will land in the right place.
- Write the numbers as words. 2026 read aloud is not reliably twenty twenty-six, and you cannot correct it after the fact.
- Around 150 words per minute. Count the words, divide by two and a half, and you have your seconds — which is what you actually need when the picture is a fixed number of eight-second clips.
Picture: 5 shots x 6s = 30 seconds
Budget: 30s at ~150 wpm = about 75 words
Written: 71 words, three sentences
Leave the last two seconds silent. A narration
that ends exactly on the last frame sounds rushed
even when it is not.Lay it against the picture, not under it
The mistake is writing narration that describes what is on screen. If the shot shows a chef lifting a lid, the line should not say a chef lifts a lid — the audience can see that, and hearing it makes both halves feel redundant.
Picture says what. Voice says why. If both say what, one of them is wasted.
Practically: write the shot list first, then write the narration against it as the thing the pictures do not show — the context, the stake, the turn. This is also the reason Director generates the treatment before the shot list and the voice after it.
Where free stops being enough
The synthesis is good and it is not a performance. It will not do sarcasm, it will not land a joke, and it will not carry an emotional turn on its own. For anything that depends on delivery, use the free voice as a scratch track to lock timing, and record a human against it.
That workflow — synth for timing, human for the take — is what a lot of professional work does anyway, and it is free to try here.



