PTENES
INEMA.CLUBPROThink like a director

AI Filmmaking Program · Course 3

Think like a director

Eight lessons to train the eye that makes decisions: intent, anchor, composition, light, and depth—what separates an ordinary image from a scene that holds your attention.

Course 3 · Lesson 1

The questions every director does

By the end of this lesson, you’ll evaluate any scene—your own or someone else’s—by answering the four questions behind a directing decision.

AI executes direction; it doesn’t create intent. If the scene has no internal logic, clear focal point, or reason for the camera, the result is inconsistent — no matter how advanced the tool is. This entire course trains the part the machine can’t do for you.

watch this lesson on video (English · optional)

↓ role to study

01 AI executes; the intent is yours

There’s a division of labor this course won’t let you forget: the tool generates pixels; the person who generates meaning is you. A scene without internal logic turns out beautiful and empty — and no technology update can fix emptiness.

Intent has a practical definition here: being able to justify every choice in the scene. If the answer to “why this angle?” is “that’s how it came out,” the scene is weaker.

The photographer sees this in graduation shoots: when she positions the family against the window light, there’s a reason—highlight the graduate, soften the faces. The AI needs to receive that same reason in writing.

The tool has no opinion. Every good scene is your opinion, carried out.

02 The five laws behind every scene that works

The whole course is built on five laws — memorize their names now, because the next lessons explore them one by one:

  • Visual hierarchy — what the eye sees first, second, and third. Without hierarchy, the audience won’t know where to look.
  • Camera logic — the camera’s position and movement have a narrative purpose, not just an aesthetic one.
  • Motion control — what moves in the scene, and why. Movement without a purpose becomes noise.
  • Emotional anchor — the element that holds the gaze and the emotion (lesson 2 is entirely about this).
  • Readability — can you understand the scene in one second? If it needs a caption, it failed.

The bistro owner tested the five laws in a video for the new location’s grand opening: he decided that viewers would see the door opening first, then the dining room, and finally the counter—and the video stopped looking like “cell phone footage.”

03 The four questions to ask before generating

Before any generation, the director questions the scene itself: Why this angle? Why this movement? Why this duration? Why this light? Four questions, four answers—or the scene goes back to the drawing board.

The side effect is liberating: the questions turn “it didn’t look good” (useless) into a precise diagnosis (“the duration isn’t justified—cut 2 seconds”). You stop redoing things in the dark.

The social media manager adopted the questions during approval with the inn owner: instead of debating taste, they answer the four questions together—and the conversation ends with a decision, not an opinion.

Before

"Just make me a video of the venue." Angle, movement, duration, and light left to chance — and redoing it becomes a lottery.

After

"Low angle to make it feel imposing, camera moving forward slowly because it's an invitation, 8 seconds because that's how long it takes to take in the setting, late-afternoon light because the house lives on dinner parties."

The payoff: four answers written before generating—and each new attempt improves something specific, not the odds.

04 Each tool has a role — none of them is yours

In the next courses, you’ll come across various tools — image generators, video generators, camera tools. Keep this rule of thumb handy: each tool has a defined role in the workflow; none of them decides for you. The director chooses the right tool for the right task.

It's the same reasoning a photographer applies to their own equipment: the long lens has a purpose, the reflector has a purpose — and neither one chooses the photo. The hierarchy is always: intention first, tool second.

Test yourself

A generated scene came out technically perfect, but “doesn’t work.” Based on this lesson’s framework, what is the most likely diagnosis?

Practice now 0/3 done

Examine a real scene

Apply the four questions to a specific case and compare your answers with the answer key — in ~10 minutes.

It’s an exercise in reading and decision-making: nothing to install or generate, and nothing can go permanently wrong. Disagreeing with the answer key is part of it — the value lies in justifying your answer.

The case: Regina, a social media manager, received an AI-generated video from the bakery for the launch of its display cake: the camera spins quickly around the cake for 20 seconds, with uniform white store lighting and a high angle. The owner thought it looked “modern”; Regina felt that “something isn’t right.” The cake is handmade and sold as a token of family affection.

Annotated answer key

A top-down angle flattens the cake and suggests surveillance — for warmth, use table height, like someone sitting down to eat. A quick spin sells technology, not coziness — a slow move toward the cake invites you in. Twenty seconds for just one product gets tiring — 8 to 10 is enough. White store lighting sterilizes the scene — warm side lighting, like light from a kitchen window, tells the right story. The scene's problem wasn't technical: it had no intention — all four answers described a jewelry store, not a family bakery.

You diagnosed a scene with four questions and redirected it in writing—exactly the judgment this lesson’s promise called for.

Summary

  • AI executes direction; intent — the reason behind each choice — remains your job.
  • A scene that works follows five rules: visual hierarchy, camera logic, motion control, emotional anchor, and legibility.
  • Before generating, ask four questions: why this angle, this movement, this duration, this lighting? If you don’t have an answer, go back to the drawing board.
  • Tools have defined roles in the workflow; the decision-making role can’t be outsourced.

Your next step

You just gained the filter that separates direction from luck—four questions that fit on a napkin.

In the next 15 minutes: take the latest video or photo you published of your work and answer the four questions about it. Wherever you can’t answer, you’ve found what to rework first.

In the next lesson, the most decisive of the five laws gets its own lesson: the anchor—the element that holds the viewer’s gaze, and that changes depending on the type of work.

Course 3 · Lesson 2

The anchor of the scene

By the end of this lesson, you’ll identify the right anchor for any scene—a person, product, or rhythm—and name it before writing your first AI prompt.

In the previous lesson, you learned the five laws behind every scene that works. The anchor is the most decisive of the five—and also the easiest to skip. Without naming it, the viewer’s eye wanders across the entire scene and never settles anywhere, even when the lighting, angle, and movement are technically correct.

↓ role to study

01 The anchor is what catches the eye and holds it

In this program, anchor is the single element the eye finds first in a scene and returns to whenever it gets lost. Every scene that works has one; every confusing scene, when you look closely, has none — only competing candidates.

Without naming the anchor before generating, the AI chooses on its own—usually the most technically striking element, not the most narratively important one. Strong color, brightness, movement: any of these can win by accident if you don’t decide first.

A maternity photographer decides this before pressing the button: focus on the hands clasped over the belly, not on the dress or the studio wall behind them. That decision — made beforehand, not corrected afterward — is what today’s lesson teaches you to repeat.

02 In a portrait, the anchor is almost always a person

When the subject is a person, the anchor is usually in a human detail: the eyes, a hand gesture, the curve of a smile. That detail carries the emotion of the entire scene—the rest (clothing, background, ambient light) just needs not to compete with it.

Back to the maternity shoot: if the prompt for the AI only says "pregnant woman in a studio, soft light," the machine decides on its own what to center — sometimes the face, sometimes the whole body, with no clear reason. Naming the anchor takes that decision out of chance.

Before

"Maternity photoshoot in a studio, soft light, neutral background." AI chooses the framing on its own — sometimes a headshot, sometimes full body, with no consistency between attempts.

After

"Hands clasped over the belly in the foreground; soft side lighting highlighting the fingers; face and background understated, outside the center of attention."

The payoff: the named anchor—the hands—becomes the standard every new attempt must meet, instead of starting from scratch with each generation.

03 For a product or business, the anchor is the transformation

When the subject is a product or service, the anchor is rarely the object sitting still—it’s the transformation it brings about: the before and after the video exists to show. An entire store, a counter, a display window: none of these is the anchor; they’re just the setting around it.

The ice cream shop owner launching a new flavor put this into practice: instead of showing the shop, the freezer, and a customer paying, the anchor became a single gesture—the spoon sinking into the container as ice cream runs down the side. Everything else in the video began serving that moment.

It's the same reasoning that separates a corporate video (shows everything, shows nothing) from a launch video (shows one thing clearly).

04 Without a fixed object, the anchor can be rhythm

Not every scene has an object or a face to hold the viewer’s gaze. In a dance, workout, or any activity with continuous movement, the anchor can be the rhythm itself—the pulse, cadence, or repetition that creates an expectation of continuity.

The social media manager who handles a dance studio’s accounts learned this while recording reels of the class: without a fixed object or face in each cut, she started cutting on the strong beat of the movement—and the videos stopped looking choppy because the anchor (the rhythm) stayed constant even as the angle changed.

The anchor isn’t the most beautiful element in the scene. It’s the element the scene can’t exist without.

Test yourself

A reel shows a sequence of jumps from a dance workout, with no object or face consistently in focus. What is the most likely anchor in this scene?

Practice now 0/3 done

Choose the right anchor for a real scene

Read the case, choose the anchor from four candidates, and compare it with the answer key — in ~8 minutes.

It’s an exercise in reading and choosing: nothing to generate or download, and no cost. There’s no answer that breaks anything — just one that teaches you more than the other.

The case: Otávio, owner of a bicycle shop, asked for a launch video for the store’s new electric bike. The first result had four elements competing for attention: the storefront in neon light, the shiny chrome handlebars, the smiling face of the hired model, and the store sign in the background. In the test group, no one could say what the main subject was after watching.

Annotated answer key

The electric bike is the right anchor: it’s the product Otávio needs to sell, and the transformation that matters — the new mobility it offers. The model’s face competes by being human; the sign and neon compete by being eye-catching — but neither is the reason the video exists. Once you name the anchor, the camera and lighting organize around it: close-ups of the handlebars and spinning wheel, the model appearing subtly only as a human scale reference, the store reduced to a discreet background.

You applied this lesson’s reasoning to a case that isn’t yours—naming the anchor before judging the scene. That’s exactly the skill today’s promise called for.

Summary

  • The anchor is the single element the eye finds first and always returns to — without it being named, the scene scatters even with perfect technique.
  • In a portrait, the anchor is almost always a person: their gaze, gesture, and expression.
  • For a product or business, the anchor is the transformation — the before and after the video exists to show.
  • Without a fixed object, the anchor can be rhythm: the pulse of the movement carries attention on its own.

Your next step

You just learned to name, in one word, why any scene exists.

Over the next 10 minutes: look at a recent photo or video from your work and write, in one sentence, what its anchor is. If you can't write the sentence, that's probably the scene that most needs to be redone.

In the next lesson, the anchor gets company: you put together the entire pre-production pipeline—script, storyboard, character, wardrobe, and setting—as a small studio would before filming anything.

Course 3 · Lesson 3

From script to scene, like a studio

By the end of this lesson, you’ll put together, for a real scene from your work, the five elements a professional studio decides on before filming: target emotion, storyboard, character, setting, and shot list.

You already know that AI needs intention, and you know how to name the anchor. What’s missing is the link that turns those two decisions into a complete, coherent scene from beginning to end—the same pre-production process a film studio follows before turning on any camera, adapted so you can work on your own.

watch this lesson on video (English · optional)

↓ role to study

01 It all starts with the emotion you want to evoke

Before creating any image, every professional director asks a simple question: what should the audience feel in this moment? The answer determines the angle, lighting, and composition afterward—never the other way around. Jumping straight into generation without answering that question is like cooking without knowing what dish you’re making.

The pizzeria owner planning a video to launch a new pizza for Father’s Day decided this before writing any prompts: the emotion was “the warmth of family gathered together,” not “an appetizing pizza by itself on a plate.” That one choice changed the angle (a full table, not an isolated close-up), the light (warm, late-afternoon), and even the duration (slower, to give viewers time to recognize the faces).

Emotion doesn’t decorate the scene. Emotion decides the scene — before any image exists.

02 The storyboard is the story’s first visual blueprint

One storyboard is the sequence of simple frames — rough sketches, with no color or detail — that shows the order in which the story will appear on screen. It defines the framing, camera viewpoint, and pace at which tension builds, all before you spend a single costly generation.

The pizzeria owner sketched four simple frames by hand before opening any image tool: the pizzeria’s storefront at night, the oven opening with the pizza inside, the pizza coming out steaming, and the family’s table receiving the dish. Only after seeing that sequence make sense on paper did he move on to generating actual images.

Getting the order wrong in a draft costs a pencil; getting it wrong in the final generation costs time and tool credits. That’s why the storyboard always comes first.

Test yourself

You already know the target emotion for the scene and are ready to begin. What’s the next step before generating any final images?

03 Create the character once, and reuse it every time

In professional production, the person or product that carries the scene is described once, in a neutral context, then reused in every new scene. Separating “who it is” from “where it is” avoids the most common problem for people generating images on their own: the same product looking like a different product each time.

For the new pizza campaign, the “character” wasn’t a person—it was the pizza itself, described once and reused in every following scene.

Before

Scene 1: “margherita pizza, thin crust, golden edge.” Scene 4: “pizza with basil and melting cheese.” Two unrelated descriptions — the pizzas come out visibly different.

After

One description, written once: “margherita pizza, thin crust with a golden edge, fresh basil, cheese melting at the edges.” The same sentence goes into every new scene without rewriting it.

The payoff: one reusable description replaces four separate ones—the pizza in frame 1 is visibly the same pizza in frame 4.

04 The environment sets the scene's scale and tension

The environment isn't decorative scenery — it determines how much the scene tightens or expands the viewer's emotions. A cramped space brings you closer and creates intimacy or urgency; a wide-open space expands and creates a sense of generosity or grandeur. The same scene in two different environments tells two different stories.

In the pizzeria campaign, the cramped, steamy kitchen, filmed up close, created a sense of urgency and dedication in the preparation—while the abundant outdoor table, filmed from a distance, created a sense of plenty and celebration. Same pizza, two settings, two emotions—and both scenes served different parts of the same video.

05 The shot list: every shot with a stated reason

With the emotion, storyboard, character, and environment decided, the final step is to translate all of that into a shot list: for each shot, the type (wide, medium, close-up), camera position, movement, and narrative purpose. It’s the same logic as the four questions from lesson 1, now applied shot by shot instead of to the entire scene at once.

A shot with no stated purpose is dead weight — cut it. A shot with a purpose is irreplaceable, even if it lasts two seconds.

Apply the same logic

  • Pizzeria owner: the campaign shot list has four lines — storefront (wide shot, establishes the place), oven opening (medium shot, builds anticipation), pizza coming out (close-up, it’s the anchor), table with everyone gathered (wide shot, closes with the target emotion).
  • Photographer: before an outdoor couples portrait session, she builds the whole sequence in one go — target emotion (“connection, not posing”), a three-moment storyboard, coordinated outfits chosen once, setting (a garden at dusk), and a shot list with every frame justified.

Practice now 0/5 done

Build the pre-production plan for a scene from your work

Write, on paper or in a simple document, the five elements in the pre-production line for one of your real scenes — in ~15 minutes.

It’s a planning exercise, with no need to open any generation tool: you just write. Nothing here is final — the goal is to come away with a draft you can adjust later, not a perfect version on the first try.

You’re done when you’ve written down the five elements and can justify each one out loud—the same pre-production process that supports a professional scene, adapted to your pace.

Summary

  • Before creating any image, decide what emotion the scene needs to evoke—it determines the angle, lighting, and composition, not the other way around.
  • A storyboard is the inexpensive draft that tests the story’s order before you spend a final generation.
  • Describe the character — person or product — once, in a neutral context, and reuse it in every scene that follows.
  • The setting determines scale and tension: a cramped space intensifies the emotion, while a wide-open space expands it.
  • A shot list only works when every line has a stated purpose — type, position, movement, and reason.

Your next step

You just put together the entire pre-production pipeline for the first time, the way a small studio would before filming.

In the next 15 minutes: save the draft you wrote in today’s exercise somewhere permanent (a note on your phone, a simple document). You can reuse it — the next similar scene starts there, not from scratch.

In the next lesson, this same pre-production pipeline gets moving: you turn the static storyboard into an animated sequence, with a camera move chosen for each shot.

Course 3 · Lesson 4

From the frame to the movement

By the end of this lesson, you’ll write a camera-movement prompt by breaking it down into subject, setting, movement, and style—instead of using a vague adjective—and turn a shot from your storyboard into video.

You already have the still storyboard from the previous lesson. AI video tools understand physics—position, direction, speed—not feelings. Asking for “a dramatic movement” is like asking a driver to “drive with emotion”: they don’t know what to do with that, and the result comes out generic or unstable.

watch this lesson on video (English · optional)

↓ role to study

01 One clear movement per shot — three are enough

Every AI video tool understands physics, not adjectives. “Dramatic movement” means nothing to it; “the camera slowly moves toward the subject” means everything. Choosing one concrete movement per shot instead of a vague feeling is what separates a prompt that works from one that comes out broken.

Three movements cover most scenes: the progress (the camera moves toward the subject, creating tension), the accompaniment (the camera follows the moving subject, creating fluidity) and the outline (the camera rotates around the subject, revealing depth and volume).

The photographer who turned a couple’s shoot storyboard into an animated teaser chose one movement for each shot: a slow push-in on the opening shot (builds anticipation), tracking the couple walking hand in hand (flow), and an arc around the final kiss (reveals them, creates intimacy).

The three movements, and when to use each

  1. Progress (the camera moves closer) — to build anticipation or tension.
  2. Follow-up (the camera follows the subject) — to make a movement feel fluid.
  3. Outline (the camera rotates around) — to reveal depth in a still moment.

02 Distance determines emotional closeness

Beyond movement, the camera’s distance from the subject changes how viewers feel: a close-up delivers emotion and detail; a medium shot focuses on the person and the action; a wide shot conveys setting and scale—“where we are,” before anything else.

In the couple's photo shoot teaser, the wide shot showed the whole garden (establishing the location), the medium shot showed the couple walking side by side (the main action), and the close-up was reserved just for their clasped hands and the ring — the moment carrying the strongest emotion in the entire video.

03 The movement style sets the scene’s mood

The same movement can have different styles, and each style carries its own mood: a smooth, controlled, stabilized camera conveys calm; a raw, shaky “handheld” camera conveys urgency or realism; slow motion conveys weight—the moment matters more than the others.

The photographer kept the camera stabilized throughout the teaser, with one exception: slow motion on the final kiss, the only moment in the video that needed to carry more weight than the others.

04 The prompt structure: subject, setting, movement, style

After choosing the first three—movement, distance, and style—the prompt for the video tool follows a simple structure, always in the same order: first the subject (who or what is in the scene), then the setting (where it takes place, with one or two mood details), then the camera movement you chose, and finally the style. This fixed order prevents you from forgetting anything and keeps the prompt short.

For the kissing shot: subject (the couple, in profile, moving closer), setting (garden at dusk, low golden light), movement (slow arc), style (slow motion, soft). Four sentences, one fully planned scene.

05 The first result is rarely the final one

One point of honesty that can save you frustration: the first generated video is rarely perfect, even when the whole structure is right. Regenerating, adjusting one word in the prompt, and trying again are all part of the process — the costliest mistake isn’t needing a second try; it’s not knowing how to diagnose why the first one didn’t work.

Common mistake

Asking for several camera movements in the same shot — “circle around, move in, and still do a pan.” The result usually comes out shaky and unfocused because the tool tries to obey everything at once, and nothing comes out clean. The photographer tried this once in the kissing shot; splitting it into two shots—outline in one, movement in the other—made both come out clean.

Practice now 0/4 done

Animate a shot from your storyboard

Write a motion prompt with the structure’s four parts and generate at least one shot — in ~12 minutes.

AI video generators with a free plan often limit how many generations you can make per day, and that quota changes frequently. So write and review the entire prompt before generating—you can test the structure on paper first, without spending one attempt after another on screen.

Cinematic video shot, single clear camera movement only, 5-8 seconds.

SUBJECT: <describe your subject — person, product, or scene>

ENVIRONMENT: <where it happens, 1-2 details that set the mood>

CAMERA MOVEMENT (choose exactly one): slow dolly in toward the subject /
tracking shot following the subject / slow orbit around the subject

STYLE (choose exactly one): smooth steadycam / raw handheld / dramatic
slow motion

Photorealistic, cinematic lighting, no text, no watermark.

You turned a still storyboard frame into a moving shot, with a named decision behind each choice — exactly what this lesson promised.

Summary

  • Video AI understands physics, not adjectives: one clear movement per shot—push-in, tracking, or orbit—works; “dramatic” doesn’t.
  • Camera distance determines emotional closeness: a close-up brings you closer, a wide shot establishes the scene, and a medium shot balances the two.
  • The motion style sets the mood: stabilized feels calm, handheld feels urgent, slow motion feels weighty.
  • Every prompt follows the same structure, always in the same order: subject, setting, movement, style.
  • Stacking several movements in the same shot is the costliest mistake—split them into separate shots, one movement at a time.

Your next step

You’ve just animated your first shot by naming a movement choice — no more “just do something dramatic.”

Over the next 15 minutes: animate two more shots from the storyboard you put together in the previous lesson, each with its own movement and style, and watch all three in sequence.

In the next lesson, you learn to plant a question in the viewer’s mind—using contradiction, setting, and reveal, the three ingredients that keep someone watching.

Course 3 · Lesson 5

The scene that makes the audience ask

By the end of this lesson, you’ll plan a scene that plants a question in the viewer’s mind—using setting, contradiction, and a sequence that reveals things gradually—instead of giving everything away in the first second.

You already know how to handle pre-production and animate a shot with deliberate movement. But a technically correct video can still be abandoned by the third second if it gives viewers no reason to keep watching. That reason has a name, and this lesson teaches you how to build it.

↓ role to study

01 You’re not creating a scene—you’re creating a question

So far, you’ve learned to choose the anchor, plan preproduction, and animate a shot. One ingredient separates a competent scene from one that holds your attention: a good scene doesn’t reveal the answer all at once — it plants a question in the viewer’s mind, then answers it.

The owner of a neighborhood burger shop decided this before filming the launch of a secret burger that wasn’t on the menu: instead of showing the whole product in the first scene, he hid the result until the final cut—and the question “What will it look like assembled?” held viewers’ attention until then.

A scene that reveals everything at once doesn’t need to be watched to the end.

02 The environment tells the story before any words.

A well-chosen setting already says a lot on its own, without needing a caption or narration to explain the obvious. Clutter, steam, movement, texture: every detail in the setting is a clue the viewer reads without realizing they’re reading it.

The same burger shop owner tested this in practice: a clean, empty kitchen, filmed in silence, said nothing about the business. The kitchen in the middle of the Friday night rush—with smoke rising, pans stacked, hands crossing—told the story of “people really working hard here” all by itself, without a line of text.

03 Contradiction is what holds the viewer’s gaze

An element that’s out of place creates curiosity faster than any “beautiful and correct” image. When something in the scene contradicts expectations — delicacy where you expect strength, care where you expect haste — viewers stop to understand why.

In the secret burger video, instead of showing the same finished product as always, the owner filmed the heavily tattooed hands of a burly cook placing a single lettuce leaf with extreme care. The contrast between the rough look of the hands and the delicacy of the gesture held more attention than the finished burger alone ever could.

04 The sequence reveals things in stages, not all at once

A compelling scene follows an order: show the world (the setting, without rushing), move closer (the camera approaches what matters), reveal (the character or product appears, but not entirely yet), break expectations (something changes or surprises), and only then show the final intention (the emotion or answer that closes the question raised at the start).

The secret hamburger video followed exactly this order: the burger joint’s exterior at night (the world) → the camera moving closer to the kitchen (moves closer) → hands assembling the burger layer by layer, without showing the top (partially reveals) → the whole burger suddenly appearing, off the menu (breaks expectations) → a surprised customer’s face as they take a bite (shows the intent: to create desire).

Before

One static shot, fifteen seconds long, showing the finished burger from above — the whole answer revealed in the first second.

After

Nighttime storefront → kitchen coming closer → layered assembly, without showing the top → complete burger only in the final cut, with the bite.

The payoff: the same information, reorganized into five parts, turned an ignored video into one people watched to the end.

Test yourself

A promotional video shows the entire product, ready and still in close-up, in the first second. Based on what you learned in this lesson, what is the most likely problem?

Practice now 0/3 done

Rewrite a scene that reveals everything at once

Read the case, rewrite the sequence in five parts, and compare it with the answer key — in ~10 minutes.

It’s an exercise in reading and rewriting: nothing to generate or publish. Disagreeing with the answer key is part of it — the value lies in explaining why each part of your sequence comes in that order.

The case: Juliana runs a neighborhood burger shop and recorded a video to launch the monthly combo: a single fixed shot showing the whole burger from above for fifteen seconds, with the caption “our new combo.” In tests, 90% of viewers left the video before the third second.

Annotated answer key

The original video answered everything in the first image: it showed the entire product, straight on, with no open questions—so viewers had no reason to keep watching. A sequence that works: the kitchen in motion during the rush (world) → the camera moving toward the grill (approach) → hands assembling the combo layer by layer, without showing the top (partial reveal) → the whole combo appearing only in the final cut, alongside a customer taking a bite (rupture + intent). The question “what will this combo look like when it’s assembled?” holds attention until the end.

You applied this lesson’s logic to a scene that isn’t yours—planting a question before answering it. That’s exactly the skill today’s promise called for.

Summary

  • A scene that works plants a question in the viewer’s mind, then answers it — it doesn’t reveal everything at once.
  • The environment already tells part of the story on its own, without needing a caption or narration to explain the obvious.
  • An unexpected element — a contradiction — holds the viewer’s attention more than any “beautiful and correct” image.
  • A sequence reveals things in stages: world, approach, reveal, rupture, intention — always in that order.

Your next step

You just learned to plant a question before answering it—the trick that keeps someone watching past the third second.

Over the next 10 minutes: take the next video you're going to shoot and write the five parts of the sequence before filming anything. If any part is missing, that's the weak point in the script.

In the next lesson, you learn to direct the viewer’s eye within each shot—thirds, lines, negative space, and scale, and why the composition changes from shot to shot instead of repeating.

Course 3 · Lesson 6

Composition in time

By the end of this lesson, you’ll create a sequence of shots using a different composition tool for each one—thirds, lines, negative space, scale—instead of repeating the same composition from beginning to end.

You already know how to plant a question in a scene. What’s left is deciding where the eye goes first within each shot. Without that decision, even a well-planned sequence gets tiring—because every photo or shot uses the same visual arrangement, without variation.

↓ role to study

01 The rule of thirds creates structure without centering

A rule of thirds is the most commonly used composition tool: instead of centering the subject, you place it at one of the points where the imaginary lines intersect. The frame gets breathing room, and the eye has somewhere to move.

Diagonals do the opposite: instead of calmly organizing the frame, they create energy and a sense of movement—used when the scene needs to feel alive, not still.

A photographer shooting a corporate portrait at a coworking space applied the rule of thirds: she positioned the face in the upper-right third, leaving the rest of the frame open so the workspace could appear in the background without competing with the face.

Centering is the easiest compositional choice. It’s almost never the right one.

02 Lines that guide the eye, and the center that stabilizes

Straight lines — a long table, a hallway, a window edge — guide the eye to the main subject without anyone noticing they’re being guided. Centering with symmetry does the opposite: it conveys control and stability, useful when the scene calls for calm authority, not movement.

For the shot of the team gathered around the conference table, the photographer used the table itself as a leading line to the leader's face, seated at the head. For the founder's solo portrait, they centered the face with symmetry — because that scene called for still authority, not energy.

03 Empty space isolates; scale expands or brings you closer

Empty space around a small subject isolates it and draws attention to it — the eye has no choice but to settle on what remains. The shot scale, from a tight detail to an entire setting, determines whether the scene speaks of intimacy or grandeur.

Before

Hands typing, medium shot, desk full of objects competing with the main action for attention.

After

Hands typing, tight framing, almost nothing else on the desk—only the keyboard and fingers visible in the frame.

The payoff: the same gesture, isolated by the empty space, became the only possible point of focus.

After that isolated close-up, the photographer wrapped up the portfolio with a very wide shot of the entire office, showing the company’s scale—the exact opposite of the close-up of the hands.

04 The composition changes from shot to shot; it never repeats

Each shot handles a different composition task — no good sequence repeats the same composition from beginning to end. If every shot uses the same third, the same centering, and the same negative space, the whole thing gets tiring before it ends, even if each shot is technically correct on its own.

In the complete lookbook, the photographer alternated among five different compositions: the rule of thirds in the individual portrait, leading lines at the conference table, symmetry in the founder’s portrait, negative space around the hands, and a wide shot of the office—none repeated twice.

Test yourself

You’re putting together a sequence of four photos for the same client. Which choice makes the sequence feel more lively and professional?

Practice now 0/4 done

Audit the composition of your latest photos

Label the composition tool used in four or five of your recent photos and fix one repeated pattern—in ~10 minutes.

It’s an observation exercise using photos that already exist: you don’t need to take any new photos or download any software. If you can’t find any repetition, the exercise is still complete — the goal is to train your eye, not force yourself to find a flaw.

You’re done when each of the 4 or 5 photos has a named compositing tool and you’ve identified (or ruled out) any repetition among them.

Summary

  • The rule of thirds creates structure without centering; diagonals add energy when a scene needs to feel alive.
  • Straight lines guide the eye to the subject without anyone noticing; symmetry conveys control and authority.
  • Empty space isolates and draws attention to a small subject; shot scale determines intimacy or grandeur.
  • A good sequence never repeats the same shot composition from one shot to the next — each one serves a different purpose.

Your next step

You just learned to vary composition on purpose instead of repeating the same visual arrangement out of habit.

Over the next 10 minutes: choose three photos or shots you still plan to generate or shoot this week, and decide before clicking anything which composition tool each one will use — without repeating any.

In the next lesson, you learn that lighting isn’t for illumination—it’s for controlling what the audience feels, including the art of taking light away instead of only adding it.

Course 3 · Lesson 7

Light is emotion

By the end of this lesson, you’ll choose among five lighting techniques for a real scene—remove light, move the source closer, hide 80%, contrast colors, contrast exposure—with emotion in mind, instead of lighting everything.

You already decide on the anchor, pre-production, movement, and composition. One variable remains—the one that most quickly changes how a scene feels: lighting. Most beginners get this backward—they light everything, when the professional effect almost always comes from taking light away.

↓ role to study

01 Taking light away creates mystery; adding it doesn’t

Professional cinematography is rarely about adding light — it’s about removing it on purpose. When you take light away from part of the scene, the shadows deepen, and the frame starts to feel mysterious and alive instead of flat and uniform.

A simple way to apply this: light the subject from only one side, leaving a rim light thin at the edge, against a darker background. The eye follows contrast, not brightness—you don't need to light everything for something to stand out.

The owner of a beauty salon filming a reel of the last haircut of the day, at night, turned off the overhead lights and left only a side light on the client’s face, with the rest of the salon dark. The result looked like a magazine video, not a cell phone recording in the dark.

Light isn’t there so you can see the scene. It’s there so the audience can feel something.

02 Softness comes from distance; every light needs a real source

Soft light doesn’t depend on the size of the source — it depends on the distance. A small light placed very close to the subject spreads out and wraps around it as if it were large; the same light far from the subject becomes hard and creates sharp shadows along the edges.

And every light in the scene needs to look like it comes from somewhere real — the sun, a lamp, a screen. If viewers believe in the source, they believe in the whole scene; light without a clear source always looks artificial, even when it’s technically perfect.

At the beauty salon, instead of asking for generic "studio lighting," the owner decided the video's light would come from the illuminated mirror at the station itself — a source that was really there, so no one questioned whether the light was "fake."

03 The 80/20 rule: almost all hidden, a little visible

Only 20% of a scene’s light needs to be visible; the other 80% stays hidden in the shadows. That proportion creates depth and makes the image feel cinematic instead of flat and uniform.

Depth also comes from layers—foreground, midground, background—and haze or vapor between them makes the light itself visible in the air, not just on the surfaces it touches.

Before

A room with all the ceiling lights on, equally bright, with no shadows anywhere—the video looks like surveillance footage, not a scene.

After

Only the workbench light is on, with the rest of the room in shadow; steam from a hot towel becomes a visible layer of light in the air.

The payoff: removing 80% of the ambient light created more depth than any new light could have.

04 Color contrast creates tension; exposure creates focus

Warm and cool tones side by side create visual tension—and tension here is another word for life in the frame. A scene with a single color temperature feels comfortable, but it’s also forgettable.

Showing only what matters and letting the rest fall into shadow or blow out creates absolute focus on a single point in the scene, without needing to crop or draw an arrow.

In the salon, the warm light from the station against the cool light coming in from the street through the window at night gave the video an "end of the workday" feel without needing a single word on screen.

Test yourself

A scene has warm light on the counter and cool light coming through the storefront at night, side by side. What effect does this choice create, according to this lesson?

05 Moving light feels alive

A candle, a trembling reflection, light passing through a leaf in the wind: moving light brings a frame to life. Completely still light, no matter how well positioned, tends to look dead after a few seconds on screen.

Apply the same logic

  • Beauty salon owner: in the day's closing video, show only the final snip under the workbench light, letting the rest of the studio fall into darkness—the final cut becomes the only possible point of focus.
  • Photographer: in an outdoor night shoot, use the wavering reflection of a streetlight on a puddle to give the light movement, instead of a still, even studio light.

Practice now 0/4 done

Write a decomposed lighting prompt

Write an image or video prompt by breaking down the light source, distance, contrast, and movement — in ~12 minutes.

AI image and video generators with a free plan often limit how many generations you can make per day, and that quota changes frequently. Write and review the entire prompt before generating—you can test the lighting decision on paper first.

Cinematic lighting, one clear light source only.

SUBJECT: <describe your subject>

LIGHT SOURCE: <a real-world source — window, lamp, screen, candle>,
positioned <side / behind / close>

CONTRAST: mostly in shadow, only 20% of the frame lit; <warm/cool>
tones against <the opposite tone> in the background

MOVEMENT: <steady light / flickering light / reflection moving>

Photorealistic, cinematic, no text, no watermark.

You chose lighting based on emotion and a real-world source, not on “lighting everything”—exactly the mindset shift this lesson’s promise called for.

Summary

  • Taking light away creates mystery: a thin rim light against a dark background separates the subject without needing to light everything.
  • Softness comes from distance, not font size; every light needs to look like it comes from somewhere real.
  • Only 20% of the light needs to be visible; layers of fog or steam make the light itself part of the depth.
  • Color contrast creates tension, exposure contrast creates focus, and moving light feels alive.

Your next step

You just replaced the instinct to “light everything” with the professional instinct to decide where to take light away.

Over the next 10 minutes: look at a space where you work right now and identify a real light source already there — a window, lamp, or screen — and imagine the scene with only that source, without any extra lighting.

In the final lesson of this program, you’ll bring it all together in one skill: arranging foreground, middle ground, and background in every frame, and telling stories about human relationships through position and gaze alone.

Course 3 · Lesson 8

Depth: the scene in three layers

By the end of this lesson, you’ll create a scene with a defined foreground, middle ground, and background, then place two people or elements on different layers to tell a relationship through position and gaze alone.

You already decide on the anchor, pre-production, movement, composition, and lighting. One final layer remains—literally: depth is what makes a frame feel like a real place you can step into, instead of a flat surface with things stuck on top.

↓ role to study

01 Every professional scene has three layers

Depth is the illusion of three-dimensional space within a flat frame. It comes from organizing the scene into three layers: the foreground—where the viewer “is”—the midground—where the main subject lives—and the background, where the context breathes. Without all three, the image looks flat: beautiful, but with no space to step into.

The layers only work when they’re visually distinct from one another — through focus, light, or color. If the three blend together, the eye can’t tell them apart, and the sense of depth disappears even if all three are technically there.

A photographer planning a family photo shoot in a park decided this before taking a single shot: slightly blurred branches in the foreground, the family in the midground, the lake in the background — three clear layers, each with its own role in the composition.

A flat image shows what exists. A layered image shows where each thing is in relation to the others.

02 Depth of field determines what stays sharp

A depth of field controls how much of the scene is in focus at once—and that’s what guides where the eye goes first. Shallow depth of field isolates the subject, blurring the foreground and background; deep depth of field keeps everything legible, useful when the setting matters as much as the person.

The lens choice reinforces this effect: a wide-angle lens expands the space, exaggerating the distance between layers; a telephoto lens compresses the space, visually bringing together layers that are far apart in real life.

For the main portrait in the family photo shoot, the photographer used a shallow depth of field to isolate the faces; for the wide shot of the entire park, they switched to a deep depth of field, because the setting mattered as much as the family there.

03 Layers tell human stories without dialogue

Placing people on different layers of the same frame tells a story without a single line of dialogue. A short distance between two people speaks of intimacy; a long distance speaks of isolation. Movement speaks of change; complete stillness speaks of tension.

In the family photo shoot, the photographer placed the grandparents in the foreground, calmly looking at the camera, with the grandchildren running in the background, out of focus. The composition alone told the story of "one generation watching, another living in motion" — without anyone having to say a word.

Test yourself

In a family portrait, the grandparents appear sharp and close to the camera, looking into the lens; the grandchildren appear blurry and far away, running in the background. What relationship does this composition suggest, without a single written word?

04 The five decisions become one: you direct, you don’t just generate

Composition, camera and lens, lighting, movement, and character position—the five decisions in this course—don’t work in isolation. In a professional scene, they work as a single system: lighting separates the layers arranged by the composition; movement reveals the depth the lens compressed or expanded; each person’s position conveys the relationship the script called for during preproduction.

It's this integration — not mastery of a single technique — that separates people who generate standalone images from those who direct an entire scene.

Before

A video with good lighting, but no layers — everyone on the same plane, with the same sharpness and distance. Technically correct, forgettable.

After

The same video with defined foreground, middle ground, and background, deliberate depth of field, and each person placed in a layer that says something about them.

The payoff: the difference between the two versions isn’t technical—it’s a decision. You’ve spent this entire course learning to make that decision intentionally instead of accepting the AI’s first generation.

Practice now 0/4 done

Build a scene with three layers and one relationship

Put together, on paper or in a real scene, the three layers of depth with a human relationship told only through positioning—in ~15 minutes.

It’s an exercise in planning and observation: you can do it on paper, with an existing photo, or by arranging a real scene — nothing here depends on generating a new image or using any specific tool.

You completed the course by building, in practice, the same scene that sums up all eight lessons: depth with human connection built in, decided by you—not by chance.

Summary

  • Depth comes from organizing the scene into three layers—front, middle, and back—that are visually distinct through focus, light, or color.
  • Depth of field determines how much of the scene stays sharp; the lens reinforces it by expanding or compressing the space between layers.
  • Placing people on different layers tells a story without dialogue: distance speaks of intimacy or isolation, movement speaks of change.
  • The course’s five decisions — composition, camera, light, movement, and position — work as one system, never in isolation.

Your next step

You just completed the director’s way of seeing: from intention and anchor to depth and relationships, all eight lessons followed one line of reasoning, from why things are there to where they are in the frame.

Over the next 15 minutes: choose a real scene from your work and map out on paper the five decisions from the entire course — anchor, pre-production, movement, composition, and lighting — applied to this lesson's three layers of depth.

Course 4, “The Production Workflow and the Language of the Camera,” takes that trained eye into the mechanics: purposeful sequences, the physics of movement, the right pace for each emotion, and the camera vocabulary a production uses from storyboard to final scene.