Start frames beat adjectives
One reference image does more than a paragraph of description, and it costs less. Why, with the same scene made both ways.
- Written by
- The Cetus team
- Seeds a claim
- 5
- Models
- 3

On this page5
There is a moment in everyone's second week where they write a four-line prompt describing a person in enormous detail — the jacket, the scar, the particular grey of the light — run it twice, and get two different people. The instinct is to describe harder. The fix is to stop describing.
Why description does not carry identity
A text prompt is a distribution, not a specification. Every sentence you add narrows the distribution, but it never collapses it to a point, and the things it fails to narrow are exactly the things a face is made of. You can pin the jacket. You cannot pin the distance between the eyes.
A start frame collapses it completely. The model is not asked to imagine a person from a description; it is handed a person and asked what happens next. Everything that was ambiguous in the paragraph is now a pixel value, and the only remaining freedom is motion — which is what you actually wanted the model to decide.
Describe what should change. Show what should stay the same.
The same scene, made both ways
A woman in a workshop, turning to look at the camera. Three clips of each, same seed budget.
Text only
A woman in her forties in a canvas apron turns from a
workbench to look at camera. Warm workshop light,
sawdust in the air, brushed steel tools behind her.
-> three different women. Two aprons, one denim jacket.
The workshop changed layout in every clip.
Start frame + motion
[reference still: the workshop, her at the bench]
She turns from the bench to look at camera. Slow.
-> the same woman, the same room, three times. The only
difference between the clips is the speed of the turn.The second prompt is fourteen words. The first is thirty-one and does less. This is the general shape of the trade: the reference does the describing, and the prompt is free to be about movement.
It is also cheaper
This is the part people do not expect. A still costs two credits. A video clip spends one of your video allowance, and the allowance is the scarce meter, not the credits.
So the arithmetic runs like this: iterate on the still until the frame is exactly right — four or five attempts is ten credits, which is nothing — and then spend one video unit animating the frame you already know is correct. The text-only route spends video units to discover things a still could have told you for a fraction of the cost.
Every failed clip in the first approach was a video unit finding out what the person looked like.
Making a start frame worth animating
- Frame it as the first frame, not as a poster. If the shot is a push in, the still should be the wide end of that push. A beautifully composed final frame is the wrong place to start from.
- Leave room in the direction of motion. If she turns left, there should be space on the left. Models will not invent room that is not in the frame; they will crop or they will refuse to move.
- Keep the subject away from the edges. Anything touching a border tends to deform first when motion starts.
- Get the light right in the still. Lighting is the thing video models are least willing to change mid-clip, which is a feature — the still is where you settle it.
- One subject, one action. A start frame with three people in it is three identities to hold, and the clip will drop one.
When to describe instead
Start frames are not free of cost — they lock composition as well as identity. If what you want is variety, a text prompt run five times is the right tool, and that is genuinely most of the work on a batch. Reach for the reference when you need the fifth clip to contain the same person as the first.
The rule of thumb that has survived contact: if the thing you are struggling to get is a noun, use a start frame. If it is a verb, write a better sentence.



