CetusLoading
All posts
Comparison11 min · 7 Aug 2026

Choosing a model: speed, cost and what each one is good at

Every model in the catalogue on one page, with the same three prompts run through all of them.

Written by
The Cetus team
Seeds a claim
5
Models
3
On this page6
  1. Omni Flash — the default for video
  2. Veo 3.1 Lite — the one that holds together
  3. Nano Banana Pro — stills
  4. Voice and script
  5. The three prompts, through everything
  6. A decision rule that fits in a sentence

The catalogue is deliberately small. Three models do all the generation work, and each of them is the obvious answer to a different question. This is what they are, what they cost, and the three prompts we ran through all of them.

Omni Flash — the default for video

Renders 4, 6, 8 or 10 seconds, and it is the only model in the catalogue where duration is a choice. It takes text or a start frame, it follows camera instructions more literally than anything else we run, and it will hold a genuinely static frame when you ask for one, which sounds trivial until you have watched three models slowly push in on a locked-off shot.

  • Best at: camera moves you specified, short beats, anything where you need a 4 or 6 second clip rather than a padded 8.
  • Weaker at: faces over long durations. At 10 seconds a close-up will drift. Cut before it does.
  • Reach for it when: you are building coverage, iterating fast, or the shot is about movement rather than about a person.

Veo 3.1 Lite — the one that holds together

Eight seconds, always. Not four, not ten — asking for another duration is a refused generation, not a rounded one, so it is worth knowing before you plan a shot list around it.

What you get for that rigidity is coherence. Faces stay the same face, hands behave more often than they have any right to, and physical interactions — pouring, opening, handing something over — survive the full eight seconds far more reliably. It is also more opinionated: give it a framing it thinks is wrong and it will quietly give the shot more room.

  • Best at: people, hands, contact between objects, anything that has to still be recognisable at the end of the clip.
  • Weaker at: obedience. If you have written a precise camera move, it may improve on it.
  • Reach for it when: the shot has a person in it and the person matters.

Nano Banana Pro — stills

One image model, on purpose. There are cheaper and faster tiers available upstream and we do not sell them, because a tier we cannot price is a tier that fails at generation time for reasons a user cannot see.

Two credits an image. It responds to material and light far more than to composition adjectives, and it is the correct first stop for anything that will eventually be animated — see the post on start frames for why that is also the cheap path.

Voice and script

Both are free, and both are free for the same honest reason: they cost us nothing per generation. Voice runs on a public synthesis endpoint. Script generation is text. Charging for either would be charging for something we do not pay for.

Practically, this means there is no reason to ration them. Generate six versions of the narration and pick one.

The three prompts, through everything

1  Wide, static. A ferry crosses a grey harbour, rain on
   the water, low cloud. Slow.

     Omni Flash    static held, rain convincing         best
     Veo 3.1 Lite  drifted right, added a second boat

2  Close-up. A man lights a match, cups the flame, looks
   up. Warm light on his face, dark behind.

     Omni Flash    face shifted at ~5s, flame good
     Veo 3.1 Lite  same face throughout, hands correct   best

3  Medium. Hands fold a paper crane on a linen table,
   window light camera left.

     Omni Flash    paper deformed mid-fold
     Veo 3.1 Lite  fold completed, plausible             best

Which reads as a clean win for Veo 3.1 Lite until you notice that two of the three prompts are about a person or a pair of hands, and that the one Omni Flash won is the one with an explicit camera instruction. That is the split, and it holds across everything else we tested.

The question is never which model is better. It is whether the shot is about a person or about a camera.

A decision rule that fits in a sentence

  • Is there a face or a pair of hands doing something? Veo 3.1 Lite.
  • Did you write a specific camera move, or do you need a clip that is not eight seconds long? Omni Flash.
  • Are you still working out what the shot looks like? Nano Banana Pro, then animate the frame you liked.
  • Is it narration or a script? Take as many attempts as you want; neither costs anything.

The catalogue will grow, and when it does this page gets rewritten rather than appended to. A comparison that only ever adds rows stops being a comparison.

Share

The loop every post is about

Whatever the subject, the method underneath it is the same three moves.

  1. Step 1

    Settle it in a still

    Composition, light and identity are decided at two credits an image, not at a video unit a clip.

  2. Step 2

    Change one thing

    One clause per re-run. A rewrite that touches everything teaches you nothing you can use next week.

  3. Step 3

    Cut before it drifts

    Six seconds of a shot that holds beats ten of one that does not. The limit is on the clip, not on the film.

More from the workshop

All posts

Stop reading, start running

Say it once. Get a film back.

Video, stills, voice and script from one line. Start from a template or write your own. The first ones are free.

Start creating