PTENES
INEMA.CLUB
0 of 5 read

Module 4.4 · Track 4 · Camera Language

Cinematic Video in Seedance 2: direct, don't generate

The practical synthesis of the track. All the camera language now goes into a real generator— with a storyboard, locked identity, and YAML prompts. AI doesn’t think: it interprets structure. And the goal is never to generate—it’s to control.

Reading now ~18 min · full module 5 sections

What you will understand

  • Why a six-shot storyboard — scale, speed, control, emotion, action, relief — keeps the model from improvising and breaking continuity.
  • How lock the identity of the character with a reference image, eliminating the drift face, proportions, and design.
  • Why write the prompt in YAML — AI doesn't think like a human; it interprets structure—and the essential rules of this format.
  • How the transitions control pacing, emotion, and attention; and why, in Seedance, they break—making direction and iteration inseparable.
Section 1 of 5·Principle

1.AI interprets structure, not desire

All the language from the previous three lessons — the pillars, the movement, the shots — eventually has to make its way into the generator. And here is the principle that organizes this entire lesson: AI doesn’t think—it interprets structure.1 Seedance doesn’t “understand” your scene the way a human collaborator would; it mechanically analyzes the text and looks for clear blocks. The more structured your description, the more predictable the result.

This changes what makes a good prompt. A long, beautiful paragraph full of adjectives is exactly the kind of input that confuses the model: it can't tell what's camera, what's action, what's lighting, and it mixes everything together. The solution isn't to write more—it's to write organized. That’s why the lesson uses YAML: not for a programmer’s aesthetic, but because it’s a control system. Everything in its place, with no ambiguity.

The goal is control, not generation

Repeat this as the module’s mantra: the goal isn’t to generate video—it’s to control it like a filmmaker. Generating is easy; anyone can type a sentence and get a clip. Control is the director’s job: deciding the sequence, locking the identity, structuring each block of time, anticipating the transitions. The difference between an amateur and a director with AI isn’t in the tool — it’s in how much of the scene they governs on purpose, instead of accepting what the machine improvised.

Stop and predict

You have a clear scene in your head and describe it in a single, rich, flowing paragraph, with camera, light, action, and emotion mixed together in the same sentence. The model returns something confused that ignores half of what you asked for. Why did such a carefully written text fail?

See one possible answer

Because AI interprets structure, not prose. In a single paragraph, it can’t distinguish camera movement from the subject’s action or the lighting—everything becomes an ambiguous soup, and it “guesses.” The same content, broken into blocks (camera, action, lighting, transitions), removes the ambiguity and gives it a clear plan to follow.

Fig. 1 · The same content: ambiguous paragraph versus clear blocks
SINGLE PARAGRAPH camera + action + light mixed together AI takes a guess BLOCKS (YAML) camera: action: lighting: transitions: everything in its own field AI follows the plan
Section 2 of 5·Storyboard

2.Six-shot storyboard

Before writing a line of prompt, draw the sequence. The storyboard — here, six shots — isn’t a formality: without it, the model starts to improvise breaks visual continuity between shots.2 With it, you give the scene a skeleton the AI can follow instead of inventing one.

Six shots that already tell a story

The example sequence follows a motorcyclist—stylized, with a strong silhouette and consistent proportions—and its six shots form a complete narrative arc: a drone from above the road (scale), a establishing shot with the pilot picking up speed (speed), a medium shot during a curve (control), a close-up (emotion), a action with the jump (action) and a establishing shot of a closing sunset (relief). Notice: each shot uses a different size on purpose, exactly as you learned in the previous lesson.

This order is not random. It is cinematic structure in its purest form: scale → speed → emotion → action → relief. First we understand the world, then feel the speed, then connect with the character, then see skill and risk, and finally feel the release. A good sequence flows in this progression—and the storyboard is where you ensure that flow before AI has a chance to scramble it.

Fig. 2 · Six shots, one arc: scale → speed → emotion → action → relief
DRONEWIDEMEDIUM CLOSE-UPACTIONWIDE scalespeedcontrol emotionactionrelief
▸ Going deeper: why six shots, not one long take optional

Optional layer—on granularity and control.

It might seem simpler to ask for a single fifteen-second long shot. But long shots are where the generator makes the most mistakes: the longer the continuous take, the greater the chance of drift, distortion, and loss of identity. Breaking the scene into six short shots gives you control points: each shot is generated, evaluated, and, if needed, redone on its own, without compromising the others. It also enforces shot grammar — each one with its own size and function. Six shots isn’t just storytelling; it’s a risk management strategy against AI’s weaknesses.

Section 3 of 5·Identity

3.Lock the character’s identity

The second practical pillar is identity lock. You’ve already seen in Module 4.1 that Seedance changes faces; the solution here is even more direct. Upload the storyboard image (or character sheet) as a reference and, in the prompt, lock it explicitly the identity — instruct the model to match that reference, with a strict lock.3

The instruction has a recognizable form: base: "matching reference @image1 with strict identity lock". This tells the generator to keep exactly the same face, helmet, suit, proportions, and colors across all shots. It’s also worth locking the environment — the road, the mountains, the light — so the setting doesn’t change from shot to shot. This single step eliminates the three most common ghosts of AI video: the drift of identity, style changes, and random redraws.

Why this comes before the YAML

Order matters: lock first, structure second. A flawless YAML is no use if, in shot 4, the character turns into someone else. The identity lock is the foundation of continuity—it ensures who this is in the scene; the YAML governs what happens to this person. Without the first, the second builds on sand. With both, you have a stable character moving through a structured sequence—which is, essentially, what a film set delivers.

# Identity lock — before anything else # upload the storyboard/sheet image as @image1 base: matching reference @image1 with strict identity lock lock_character: exact same face, helmet, suit, proportions and colors across all shots lock_environment: same road, mountains and lighting throughout production_notes: no deformation, no identity drift, no redesigns, no extra limbs
Fig. 3 · No lockup: the face drift; with a lock: the identity is maintained
NO LOCK changes with each shot — drift WITH LOCK the same, from shot 1 to 6 @image1 IDENTITY LOCK
Section 4 of 5·YAML

4.Writing in YAML — the rules

O

YAMLYAML is a way to write instructions in labeled blocks — each piece of information on its own line chave: valor. In the context of prompts, it organizes the scene into fields (camera, action, lighting, transitions) so the model knows exactly what belongs where.
is the way to turn the storyboard into instructions the AI can execute. Instead of a paragraph, you break the scene into blocks: character, camera, action, light, sound, transitions. Since Seedance interprets structure, the YAML remove the ambiguity and helps the model understand exactly what belongs where. But YAML only works if you follow a few rules — and they sum up the entire camera language of the learning path.

The six essential rules

One function per field. Camera isn’t action; don’t mix them. Describe physical movement, not an abstract idea. Don’t write “she runs fast”; write “the feet compress, the body leans, the scenery speeds up”— the model understands physics, not adjectives. Keep the camera simple and controlled — one movement per shot, as you learned in 4.2.

Always include light in every shot; even a short phrase like “warm side light” greatly improves the result. Define blocks of time clear — 00_02, 02_04, 04_06 — so every second is described, with no gaps. Write the transitions explicitly, with a moment, direction, and purpose. The master rule behind it all: the structure breaks, the result breaks.4

# YAML shot — one function per field, physical movement, always include lighting # one of the six blocks in the motorcyclist sequence base: matching reference @image1 with strict identity lock shot_03_medium: time: 04_06 camera: medium shot, single slow tracking move alongside the rider — no zoom, no orbit action: rider leans into a curve, feet compress on the pegs, gravel sprays from the rear wheel lighting: warm low side light, long shadows, dust catching the sun sound: engine growl (diegetic), low synth pulse rising (non-diegetic) transition_out: match-on-action into the close-up at 06s — the lean continues across the cut production_notes: no deformation, no identity drift, physical motion only
Fig. 4 · Anatomy of a block: each field, a function
time: 04_06 camera: one slow track action: leans, feet compress lighting: warm side light transition_out: match cut when—with no gaps one movement, only physics, not an adjective always present connects to the next shot
Section 5 of 5·Transitions

5.Transitions, limits, and iteration

The transitions are the glue that holds the entire sequence together. In cinema, they are not effects — they are storytelling tools that control three things: rhythm, emotion, and attention.5 Four of them structure the example scene, and each has a distinct purpose.

Four transitions, four purposes

O hard cut (cut) is the structural transition: it starts the rhythm, connecting one shot to the next without disguise. The whip pan is the movement transition: the camera spins quickly, creates blur, and hides the cut—it serves to connect fast actions and maintain the energy. The smash cut is the contrast transition: a sudden change in scale, sound, or emotion, used for impact and surprise. It is the match on action (or passage) is the continuity transition: the movement carries across the cut, so the viewer doesn’t feel the edit—it keeps the action flowing. Together, in the example: the cut starts the rhythm, the whip pan builds speed, the smash cut creates impact, and the match on action sustains the movement until the final relief.

The frustrating limitation: transitions break in Seedance

And here comes the truth the course doesn’t hide: in Seedance, transitions frequently break. Even with a correct prompt, movement can drift, timing can fail, and shots can collapse. This is normal. AI doesn't executes — it brings closer. Expecting it to get the transition right on the first try is the mistake of someone who still thinks like a user, not a director. The right response isn't frustration: it's iteration. Generate, observe the precise point of the cut, adjust the block, and generate again. Control comes from directed repetition.

Close out the path with the sentence that sums it up: a good sequence flows — scale, speed, emotion, action, relief — and the goal was never to generate, but to control. You entered this track knowing how to operate tools; you leave it speaking the language they obey: the pillars, motion, shots, continuity, structure. The next track — VFX and Action — builds on all of this and goes further, into slow motion, particle systems, and action scenes. You no longer prompt: you direct.

Fig. 5 · Four transitions—and the loop that fixes what AI only approximates
CUTWHIP PAN SMASHMATCH rhythmspeed impactflow AI MOVES CLOSER — YOU ITERATE generateobserve the cutadjust

Before moving on: four quick checks

No grades, no score. Answer from memory, then reveal the answer to compare.

01Why YAML instead of a well-written paragraph?Reveal

Because AI doesn't think — it interprets structure. One paragraph mixes camera, action, and lighting, and the model “guesses.” YAML breaks the scene into labeled blocks (camera, action, lighting, transitions), removing ambiguity. It isn’t formatting—it’s a control system.

02What is the narrative arc of the six-shot storyboard?Reveal

Scale → speed → emotion → action → relief. Drone (scale), wide (speed), medium (control), close-up (emotion), action (impact), sunset wide (relief). Each shot uses a different size for a reason—and the storyboard ensures this flow before the AI improvises.

03What does the identity lock solve, and why does it come before the YAML?Reveal

It keeps the face, outfit, proportions, and colors consistent across all shots (and the environment, if it’s locked too), eliminating drift, style changes, and redraws. It comes first because it ensures who this is in the scene; the YAML governs what happens. Without the constraint, the best structure builds on sand.

04In Seedance, transitions break. What’s the right response to that?Reveal

Iteration, not frustration. AI gets close; it doesn’t execute: movement drifts, timing fails, shots collapse—and that’s normal. A director’s response is to generate, observe the exact point where it breaks down, adjust the block, and generate again. The goal was never to get it right on the first try; it was to control.

❧ My journey

In this module
0 of 5 sections read
In track 4
0 of 19 topics
In the course
0 of 138 covered
Continue to the next track
Module 5.1 — Visual Effects
Next →