How to plan camera moves before you touch a single prompt

Most AI filmmakers write prompts. Few plan shots. Here is how borrowing a single habit from traditional cinematography gives your footage shape, intention, and a reason to cut.

By Leeby Shmeeby

There is a specific frustration that arrives about halfway through an AI film project. You have clips. Some of them are genuinely good. But when you drop them into a timeline they feel arbitrary, like a collection of images that happen to move, rather than a sequence that builds toward anything. The camera is always at roughly the same height. The framing is always roughly the same width. Nothing motivated the lens to be where it is, and the viewer can sense that absence even if they could never name it.

The cause is almost never the model. It is the absence of shot planning before prompting. Traditional cinematographers work from a shot list that specifies not just what is in the frame but why the camera is where it is, what it is doing, and what emotional work that is supposed to do. That habit translates almost perfectly into AI production, and it costs nothing except ten minutes with a piece of paper before you open any tool.

How to plan camera moves before you touch a single prompt

Start with coverage, not inspiration

Before you write a single prompt, decide how many distinct angles you need for each scene. A useful minimum for any scene with dramatic weight is three: a wide that establishes space and relationship, a medium that shows body language and expression, and a close-up that delivers the emotional detail. That is not a creative rule, it is a structural one. Without coverage you cannot cut, and if you cannot cut you are making a slideshow.

Write this down per scene, on paper or in a plain text file, before you open any generation tool:

```
Scene 3 — Mara reads the letter

1. Wide (establishing): camera static, low angle, Mara small in frame, kitchen behind her. Purpose: isolation.
2. Medium (over-shoulder): slight push in, letter visible but unreadable. Purpose: complicity.
3. Close-up (insert): hands holding letter, paper trembling. Purpose: visceral detail.
4. Reaction close (face): static, eyes tracking down page, then stillness. Purpose: the decision landing.
```

This takes three minutes. It means every prompt you write has a job, and you will notice immediately when a generated clip fails at that job, which is far more useful than noticing it looks slightly wrong.

Translate coverage into prompt structure

Each shot on your list maps to a prompt fragment that should appear early in your text input, before any description of action or mood. Camera position, lens character, and movement all belong in the opening clause. Here is the structure:

```
[MOVEMENT or STATIC] [LENS CHARACTER] [FRAMING] — [SUBJECT + ACTION] — [ENVIRONMENT] — [MOOD or LIGHT]
```

Applied to the wide above:

```
Static low-angle wide shot, slight upward tilt — a woman in her late thirties stands alone reading a letter, shoulders drawn in — large farmhouse kitchen, morning light through a single window behind her — quiet dread, muted tones
```

And for the close insert:

```
Extreme close-up, static, shallow depth — two hands grip a folded letter, paper edges slightly trembling — warm ambient light from below frame — tense stillness
```

The movement instruction comes first because current video models weight the opening of a prompt heavily when deciding how to initialise the frame and whether to add camera motion at all. If camera intention is buried after three sentences of mood writing, you will get the model's default behaviour, which tends toward a slow, unanchored drift.

Use lens language as a constraint, not decoration

Words like "telephoto compression", "wide-angle foreground emphasis", and "shallow focus rack" do observable work in current video generation. They are not decoration. A telephoto framing prompt tends to flatten the space between subject and background, which creates a hemmed-in quality useful for tension. A wide-angle foreground prompt tends to produce more spatial depth and a sense of the environment pressing in on the subject. Use these deliberately, as you would choose a lens on set.

For close-ups, adding "very shallow depth of field, background fully dissolved" consistently reduces background noise and artefacts in generated footage, which is a practical editing benefit as well as a compositional one. Less background information means less background to go wrong.

The before and after

Without a shot list, a typical AI film scene ends up with three clips at roughly the same focal length, all with similar slow push-in motions, because that is the path of least resistance. With even the minimal four-shot coverage plan above, you have wide-to-tight progression, a motivated reason to cut, and footage that can be assembled into something that reads as intentional. Viewers will not know you planned it. They will simply feel that the film has a point of view.

This week, take one scene you are already working on and write the coverage list before generating anything new. See how many of your existing clips actually have a defined job, and how many were generated on instinct alone.