← The Brand Playbook

People who ask 'which program is this?' after watching a film · 4 min read

The AI Tools Behind Every Film We Make

We stitch together an image model for character sheets, a video model for motion, a voice model for lines, and a human editor for the parts none of them get right.

What actually draws the character?

An image model, not a video model. We use Midjourney to build the character sheet before anything moves: a front view, a 3/4 turn, a profile, an expression sheet, sometimes a wardrobe sheet if the character changes outfits. That sheet is the reference every later shot gets locked to.

Getting a face that repeats isn't one click. We lock a seed and feed the same reference image back into every new pose or scene, and we throw out most of what comes back. A usable four-pose character sheet is often the survivor of 20 to 30 generations, not the first try.

  • Front view and profile for the model to match against
  • 3/4 turnaround for angled shots
  • Expression sheet: neutral, smiling, surprised, talking
  • Wardrobe or prop sheet if the story needs one

What turns a still image into motion?

Video models: mainly Kling and Runway, sometimes Luma, depending on which one is handling motion and camera moves well that week. You feed in the still and a motion prompt and it renders a short clip, usually 4 to 10 seconds.

Nobody's video model holds a face perfectly steady for 15 straight seconds - the eyes drift, the proportions creep, especially past about 5 seconds. So a '15-second film' is almost never one continuous render. It's 6 to 10 short clips, each re-anchored to the character sheet, cut together so the drift doesn't show.

Where does the voice and music come from?

ElevenLabs does the voice - either a stock voice from their library or a cloned one if the brand already has a spokesperson we're matching. For music we mostly pull from a stock library rather than generate it; AI music tools are good at a vibe but bad at hitting a specific cue on a specific frame, and a 15-second spot lives or dies on that timing.

Sound design - footsteps, a door, room tone, the little clink of a coffee cup - is still done by a human editor by hand. We tried AI-generated sound effects and they're close enough to notice something's off and not close enough to use.

What did we try and drop?

We dropped one-click 'text to full video' tools early. They're fast, but they regenerate the character from scratch on every shot, so the face and outfit drift scene to scene - fine for a mood reel, unusable for a brand that needs to look like the same character twice.

We also dropped dedicated lip-sync tools. Overlaying a separate mouth-movement pass on top of a rendered face reliably looked wrong around the jaw. We get better results letting the video model animate the mouth as part of the original render, even though that means more re-rolls to get a clean take.

What does this actually cost per finished film?

Every generation - image or video - burns a credit whether we keep it or not, and most don't get kept. A finished 15-second, 9:16 film usually represents somewhere in the range of 150 to 300 individual generations once you count the discarded character-sheet attempts and the re-rolled clips.

That's the gap between the $97 tier and the $449 tier: it's not a fancier model, it's more generation budget and more human editing time to chase down a clean character across more shots and a longer script. The tools are the same across every tier. What changes is how many attempts we're willing to pay for to get it right.

Common questions

Is this all one app, like a text-to-video button?

No. It's an image model, a video model, a voice model, and a human editor stitching the output together in a normal video editor. No single tool does the whole job well.

Which tool matters most?

The character sheet stage. If the reference images are wrong or inconsistent, every shot built from them inherits the problem, and no later tool fixes it.

Do you write the scripts with AI too?

We use text models to draft options, but a person picks the line, tightens it, and matches it to the brand's actual voice before it's recorded.

Does the toolkit change often?

Yes - we re-test the leading image and video models every few months and swap when one clearly gets better results, so this list will look different in a year.