AI Filmmaking Program · Course 2 · Track C
Three lessons in fine-tuning: sequences planned as preproduction, light that tells a story, and a character whose face never changes — in any world.
Lessons
Build a short sequence with cinematic language—each generation has a purpose, nothing is left to chance.
Master the five qualities of light and color temperature — and change the emotion of a scene without touching anything else.
Visual reference, text anchor, and locked wardrobe: the same believable character in four different settings.
Course 2C · Lesson 1
By the end of this lesson, you’ll create a sequence of 3 to 5 shots in Midjourney, deliberately choosing the lens, shot size, and lighting—and know how to justify each choice.
In film, clarity comes before beauty. A beautiful image that doesn’t connect with the previous one is a dead end—and that’s exactly what you get when you generate “one more pretty image.” This lesson turns the generator into a pre-production tool: every frame is created to serve a purpose in the sequence.
watch this lesson on video (English · optional)
↓ role to study
Each shot size does a job: the wide shot shows scale and place; the open places the character in the space; the medium makes the action clear; the close delivers emotion; the detail points out what matters. A good sequence alternates these functions — that’s what creates rhythm.
The photographer has assembled wedding portfolios this way for years, even without naming it: the whole church (wide), the couple at the altar (medium), the ring sliding onto the finger (detail), the mother’s tear (close-up). Four sizes, four purposes—and the album tells the story on its own.
Clarity comes before beauty.
A simple lens rule, enough to direct the shot: 24mm dramatizes the space (environments swallow the character), 35mm balances naturally, 85mm isolates and moves you. You don't need to know optics—you need to know what feeling you want.
And the golden rule of lighting in this track: every light has a named source. The moon is the key light, the lamp is the fill light, and the phone screen is the accent. Brightness without a source is the amateur’s tell—the eye can tell it wouldn’t exist in the real world.
The photographer working with the chef client applied both rules to his portrait: an 85mm lens to isolate his face, and the only light coming from the lit stove—real source, real drama.
Test yourself
Your night scene came out with a beautiful but “fake” glow. What’s the fix in this lesson?
Before generating the first frame, make four decisions on paper: one logline (the story in one sentence), a shot list (3 to 5, each with a size and purpose), the fixed character descriptions, e o locked format — 16:9, the wide rectangle used in movies, in every frame.
It's the difference between ordering and fishing. Someone who fishes generates twenty images and hopes three of them match; someone who orders generates five, each with a place in the sequence.
For the jewelry product shoot, the photographer wrote the logline — "the jewel crosses the city night until it finds its owner" — and the list of five shots before opening the generator. Result: the first batch was already 80% usable.
Before
Twenty disconnected generations, “just to see what comes out” — three usable ones, none of them connected.
After
Logline + list of 5 shots + fixed characters + locked 16:9 — five generations, each with a role in the sequence.
The payoff: 4× fewer generations, and the result is a sequence—not a folder full of one-offs.
Practice now 0/4 done
Walk away with a short storyboard—3 to 5 frames in 16:9, each with a justified lens, shot size, and lighting—in ~12 minutes.
Everything happens on paper and in the generator: none of your files are changed. A plan that doesn’t work regenerates on its own without affecting the other frames.
<tamanho do plano: extreme wide / wide / medium / close-up / insert>, <lente: 24mm / 35mm / 85mm lens>, <personagem: a MESMA descrição fixa em todos os planos>, <ação deste plano, em uma frase>, <luz COM fonte nomeada: moonlight as key light, a desk lamp as practical light...>, cinematic realism, 16:9 --ar 16:9
You assembled a sequence where every frame has a chosen purpose, lens, and light—control over chance, just as this lesson promised.
Summary
Course 2C · Lesson 2
By the end of this lesson, you’ll change only the lighting in a scene — keeping everything else identical — and name the emotion that light will evoke in the viewer before generating it.
"Cinematic lighting" is the most common request and the one that fails most often: it's vague, and a vague request becomes a lottery. Light isn't a finishing touch — it's the strongest lever for creating emotion in a still image. Mastering the five languages of light and color temperature is what separates a pretty scene from one that tells its own story without needing a caption.
watch this lesson on video (English · optional)
↓ role to study
Cinema has settled on five lighting families, each with its own emotional purpose:
The photographer landed a contract to shoot an indie band’s album cover. They asked for “something dark, kind of threatening” over the phone. She knew exactly what to do: low-key lighting, almost the entire frame in shadow, hard light coming from a single point, isolating the lead singer. She didn’t write “dark” anywhere in the prompt—the lighting combination conveyed the mood on its own.
Light doesn’t decorate the scene—it decides what the audience feels.
Every light has a background color, called color temperature, measured in degrees Kelvin. You don’t need to memorize the number—think of three drawers: the warm (memory of a fireplace and home), the neutral (midday light, no drama) and the cool (computer screen, hospital, future).
Each drawer carries a ready-made feeling: warm means safety and memory; neutral means raw, documentary reality; cool means isolation, technology, fear. The strongest scenes in cinema often mix two temperatures in the same image — a warm character in a cool room, for example, creates tension before anything happens.
The photographer did a newborn shoot for a family who wanted, in the mother’s words, “cozy, not a studio.” Replacing the room’s neutral light with warm 3200K light, as if it came from the lamp in the corner, was enough to make the whole scene feel like a memory—even though it was made in the moment.
Test yourself
A client asks for a scene that feels “cold, isolated, almost science fiction” for a tech product launch. What light temperature do you choose?
"Beautiful lighting, dramatic shadows, mysterious mood" seems like a complete request — but it's just one adjective repeated three times. The generator doesn't know where the light is coming from, so it makes that up, and the result varies with each attempt.
What follows may look like a technical recipe in English, but it’s just the answer to four simple questions, always in the same order: where the light comes from, which direction it points, whether it’s warm or cool, and where the shadow falls.
Before (vague)
"cinematic lighting, beautiful shadows, moody atmosphere" — three adjectives, no light source, a result left to chance with each generation.
After (named)
"single overhead tungsten bulb, hard light from above, warm 3200K, deep shadow across the left side of the face, sharp falloff into black" — source, direction, temperature, and shadow, all four named.
The payoff: the same scene stops being a lottery and becomes a commission—you choose the emotion before you hit generate, not after you see the result.
The photographer landed a contract with a high-end real estate agency to photograph an apartment for sale. Her old prompt—“cozy light at sunset”—produced three different results every time she tried. By naming the source (west-facing window), direction (side-lit, low), temperature (warm, 3200K), and shadow (long, cast across the floor), the second attempt was already the listing’s cover image.
The most revealing test in this lesson is simple: take exactly the same scene—the same character, framing, and pose—and change only the lighting. The subject stays identical. The whole story changes.
Warm light coming in from the side creates a sense of safety; the same scene with harsh light from above creates tension; switch to cool moonlight, and the mood becomes dangerous. Nothing changed except the light — which is exactly why it’s the strongest lever you have.
The photographer relies on this every month: she reuses the same photo of a cake from an artisan bakery for two campaigns for the same client. With warm window light, it becomes a "cozy breakfast" ad; with cool light and a hint of neon in the background, the same photo becomes a "holiday party" ad. One photographed scene, two campaigns sold—without rescheduling the shoot.
Practice now 0/4 done
Leave with three versions of the same scene (warm, cool, dramatic) and one sentence per version naming the light source, direction, color temperature, and shadow—in ~12 minutes.
You’re generating new images, never deleting anything: each attempt adds another file instead of replacing one. If the lighting didn’t come out as expected, the problem is almost always one of the four questions left unanswered—review them and generate again.
<a mesma descrição de cena e personagem em todas as versões>, <fonte da luz: janela, abajur, lua, letreiro...>, <direção: de cima, lateral, de trás, de baixo>, <temperatura: warm 3200K / neutral 5600K / cold 7000K>, <sombra: onde cai e como se comporta — soft falloff / hard falloff into black>, cinematic realism, 16:9 --ar 16:9
You changed the lighting while keeping the scene identical and named the emotion before generating — exactly what this lesson promised.
Summary
Course 2C · Lesson 3
By the end of this lesson, you’ll create a character anchor and wardrobe lock that keep the same face and wardrobe recognizable across at least four different settings.
A beautiful frame is worth nothing if the character’s face changes in the next one. It’s the most common flaw beginners run into when generating sequences — and the most common reason a client rejects an entire project. This lesson gives you the defense professionals use: it’s not about cloning a face perfectly, but about maintaining a stable, believable visual identity.
watch this lesson on video (English · optional)
↓ role to study
Before any technique, understand what makes a sequence of images “recognizable.” There are three layers, and they don’t carry equal weight:
Marina, a freelance photographer in Recife, nearly lost a six-month contract with a cosmetics brand because of this. The first batch of images of the campaign "ambassador" looked beautiful — but each one seemed to be a different woman: lighter eyes here, a thinner nose there. The client rejected everything in five minutes: "this isn’t the same person." That’s when Marina learned the hard way that generating a beautiful face once is easy — keeping it consistent is the real work.
Midjourney has a feature called Visual Reference: you upload an image of your character, and the next generations use that face as their starting point. It’s not magic — the adjustment that determines whether it works well is called Reference Strength, and it has three positions.
In strength low, the AI has more creative freedom—and the character starts to drift from the original with each new scene. At strength high, the lock is strong, but the result can look stiff, with poses and expressions that repeat too much. The professional starting point is strength balanced: a stable identity with natural variation from scene to scene.
The photographer tested all three on the same workday: at low strength, the client complained that "she didn’t look like the same model anymore" after the third scene; at high strength, they complained that "it always looks like the same stiff pose"; with the balanced setting, they approved the entire batch without asking for adjustments.
The second line of defense is textual, and it works in any generator: the character anchor — a fixed block of text that you copy, word for word, into every prompt for that character. The AI remembers repeated information better than a description that changes every time. Repetition creates stability.
What follows is just this: a descriptive paragraph in English, written once, saved, and pasted unchanged into each new scene—like an identity sheet you don’t rewrite, just copy.
extremely handsome wealthy older man around 68 years old, silver slicked-back hair, deep blue eyes, elegant aged face, wearing a beige cashmere overcoat, cream turtleneck, tailored trousers, luxury watch, old money aesthetic
Common mistake
Rewrite the description with synonyms for every new prompt — “gray hair” in one scene and “silver hair” in the next. To the AI, those are two different details, not the same person described in different words. Copy the same block every time, without swapping one term for another.
The photographer builds the character anchor before the first generation, together with the client—like filling out a casting sheet. After that, she just pastes the same paragraph into each new scene and changes only the setting around it.
Even with a visual reference and a text anchor, the face can still vary a little from scene to scene—that’s a real limitation of the technology, not a mistake on your part. The third safeguard addresses exactly that gap: the wardrobe lock.
Faces change; silhouettes endure. You recognize a strong character by their color, shape, clothing, and accessories—a specific coat, a pair of glasses, a necklace, a pair of shoes. When you lock in these elements in the text anchor, the audience can still recognize the character even if a fine detail of the face changes.
The photographer created a “traveler” character for a travel agency campaign: a red scarf around the neck and a specific leather backpack, locked into the anchor in every scene. Even though the face varied slightly from city to city in the sequence, no one in the client review had any doubt it was the same person—the scarf and backpack did the work.
None of the three safeguards solves everything on its own—they work together. A visual reference + a character anchor + a wardrobe lock is the combination that keeps a character believable even when you change the entire setting: streets in another city, a hotel lobby, a café, a beach.
Building a stable identity
The photographer applied all three safeguards together for a clothing boutique that wanted a “regular customer” wearing pieces from the collection in four different posts: breakfast, travel, office, evening event. Reference saved, anchor pasted, wardrobe locked—four scenes, four posts, the same recognizable person in all of them, and the entire collection featured without hiring four models.
Practice now 0/4 done
Walk away with a recognizable character in 4 different settings, using a visual reference, a text anchor, and locked wardrobe—in ~13 minutes.
Each generation is a new file — the original reference stays saved and intact, and you can repeat any scene as many times as you like. If the character “drifts” in a scene, the problem is almost always that the anchor was rewritten with different words: go back to the original text and paste it again without changing it.
<a MESMA âncora de personagem, copiada sem alterar: idade, rosto, e ao menos 2 peças fixas de figurino>, <cenário deste plano: cidade, ambiente, hora do dia>, <ação simples do personagem nesta cena>, cinematic realism, consistent character --ar 16:9
You placed the same character in four different worlds, and they remained recognizable—the promise of this lesson, fulfilled.
Summary