Keeping a character the same person across twelve shots
Identity does not survive a text description. What does survive is a reference still, a fixed vocabulary and a shorter clip.
- Written by
- The Cetus team
- Seeds a claim
- 5
- Models
- 3

On this page5
One clip of a person is easy. Twelve clips of the same person is the problem that separates a demo from a piece of work, and it is not solved by writing a better description. It is solved by changing what you hand the model.
Three things that hold identity, in order
- A reference still, generated once and reused. This does almost all of the work. Make the face at image prices, then animate that exact frame for every shot the character appears in.
- A fixed vocabulary. Whatever words you use for the character in shot one, use the same words, in the same order, in shot twelve. Not synonyms. A model reading dark-haired woman in a canvas apron and then woman with brown hair wearing a work smock is being asked about two people.
- A shorter clip. Identity degrades with time, not with complexity. The same face is far more likely to survive six seconds than ten, and cutting is free.
Build a character sheet before the first shot
Half a page, written once, pasted into every prompt. Ours looks like this, and the discipline is that it never changes mid-project — if it needs to change, it changes for every shot and the earlier ones get re-rendered.
CHARACTER: MARA
woman, mid forties, short grey-streaked hair
canvas work apron over a navy shirt
no jewellery, no glasses
reference still: mara-bench-01
USE VERBATIM. Do not paraphrase.
Every shot: start from the reference still.The list is short on purpose. Every extra clause is another thing the model can weight differently between shots, and beyond about five attributes you are adding variance rather than removing it.
Vary the shot, not the person
Once the character is fixed, all your variation has to come from somewhere else, and the good news is that this is exactly where variation belongs: shot size, angle, camera move, light, what she is doing.
A sequence of twelve shots that are all medium and all static reads as a slideshow no matter how consistent the face is. A sequence that goes wide, medium, close, insert, wide reads as a scene. Coverage is the interesting variable and identity is the constant, which is the opposite of how most people write their first shot list.
Fix the person. Vary the camera. Doing it the other way round is why the set does not cut together.
Where it still breaks
- Profile to front. A reference still shot front-on will hold front-on and three-quarter views well and will invent a profile. If you need a profile, make a second reference still of the profile.
- Distance changes. A reference still framed as a medium does not reliably produce the same face at wide, because at wide there are barely enough pixels for a face to be anything. Wides are where you should be least precious about identity — and audiences are too.
- Hands holding objects. Not an identity failure, but it is where clips fail, and it fails more often on longer durations. Veo 3.1 Lite is the model to reach for here.
- Two characters in one frame. Two identities, one clip, and the model will average them if they are similar. Cast them to look different on purpose.
The workflow that survived
Generate the reference still and iterate until it is genuinely right, because everything downstream is a copy of it. Write the character sheet. Write the shot list with coverage in it. Render every shot from the reference still with the sheet pasted in. Cut before the clip drifts.
That is exactly the order Director works in, which is not a coincidence — the ordering came out of doing it by hand badly first.



