PTENES
INEMA.CLUBPROFundamentals—where it all begins

AI Filmmaking Program · Course 1

Fundamentals—where everything begins

In eight lessons, you’ll go from square one: create your first realistic image, bring it to life, and put together your first mini-movie — with no technical background needed.

Course 1 · Lesson 1

The anatomy of realism

By the end of this lesson, you’ll be able to recreate the mood and lighting of a professional photo you admire in an AI-generated image—using only what you have today.

Have you ever asked an AI for an image and gotten something that looked like a plastic doll? Most people give up there, thinking the tool is weak. The problem isn’t the tool: no one taught the AI—or you—to look at light. That’s what this lesson fixes.

watch this lesson on video (English · optional)

↓ role to study

01 Light matters more than the tool

Nine out of ten AI-generated images look like cheap plastic. The cause is almost never the AI you chose: it’s the light no one described. In cinema, there’s a rule of thumb — lighting accounts for about 70% of what makes a scene look expensive. The tool accounts for the other 30%.

The photographer knows this by heart: the same face, photographed at noon in a parking lot and near a window in the late afternoon, looks like two different people. The camera is the same. The light is what changed.

It’s the same with AI. When you write a prompt—the so-called prompt — without mentioning lighting, the AI chooses for you. And it picks the dullest lighting possible.

If you describe the light, you direct. If you don’t, you hope.

02 Too perfect looks fake

Here’s the paradox that separates amateurs from professionals: skin with no pores, fabric with no folds, a floor with no marks — all of it screams “computer.” The human eye trusts imperfection because the real world is imperfect.

That’s why people who work seriously with images do the opposite of what instinct suggests: they ask for texture, grain, visible pores, and small signs of wear. They deliberately make the image “worse”—and it starts to look like an expensive photograph.

A social media manager sees this in practice: a post that’s too perfect looks like a bank ad, and the audience scrolls right past. A photo with a real-life texture makes them stop.

03 The professional formula: observe, describe, reference

Every cinematic result comes from the same three-part recipe: observation (you look at a photo you admire and notice the lighting, lens, and mood), technical description (you turn what you saw into precise words in the prompt) and reference (you provide your own photo for the AI to copy the color palette and style).

A reference image is the most powerful shortcut of the three: instead of trying to explain a color tone with words, you show it.

A restaurant owner can use this tomorrow: take the cover photo of an award-winning restaurant he admires, provide it as a reference, and ask for his dish in that same window light with soft shadows.

Before

"A beautiful photo of a plate of pasta." AI decides everything — and returns a generic stock photo.

After

"Handmade pasta on a deep plate, late-afternoon window light from the side, long shadows, subtle steam, rumpled linen tablecloth." Same tool, magazine-cover result.

Result: zero extra dollars spent—the entire difference came from the description.

04 Let your chat AI be your cinematographer

"But I don't know how to talk about lenses and shadows" — you don't need to. You'll use a chat AI (ChatGPT works, even in the free version) as a translator: it looks at the reference photo and writes the technical description for you.

The process is simple: you send the photo you admire to ChatGPT along with a ready-made request, and it returns two things — an analysis of the photo in a photographer's language and a technical prompt ready to paste into any image generator.

For the photographer, it’s like having an assistant who can describe any light she shows them. For the social media manager, it means turning any client reference into a technical request in a minute.

The ready-to-use prompt is in the exercise just below. It’s a recipe: you don’t need to understand every word—just swap in the photo.

Test yourself

Your image came out looking plastic. According to this lesson’s rule, what’s the first thing to suspect?

Practice now 0/4 done

Recreate the mood of a photo you admire

Walk away with an AI-generated image of yours in the mood of a reference you chose—in ~10 minutes.

Nothing here changes your files or costs anything: you’re only talking with the AI. If the result looks strange, just generate it again—each try is free and independent.

Analise esta imagem como um fotógrafo profissional e diretor de fotografia.
Descreva com precisão: pose, ângulo, expressão, estilo e direção da luz,
sombras, lente, profundidade de campo, cores, texturas, roupa, composição
e clima geral. Seja fiel ao que está visível na imagem.

Depois da análise, escreva um prompt técnico completo, EM INGLÊS, em um
único parágrafo, para eu recriar esse mesmo estilo visual em um gerador
de imagens — trocando o assunto por: <descreva aqui o SEU assunto:
seu produto, seu prato, seu retrato>.

Responda em português; só o prompt final vai em inglês.

You recreated the mood of a professional photo in an image of yourself—the promise of this lesson, fulfilled with what you already had on hand.

Summary

  • The realism of an AI image comes first and foremost from the light you describe—the tool matters less than it seems.
  • Intentional imperfections (texture, grain, signs of wear) are what make the eye accept the image as a real photograph.
  • The complete recipe has three parts that work together: observe a reference, describe it technically, and provide your own photo as a guide for color and style.
  • You don’t need to master a photographer’s vocabulary: the chat AI can translate any reference into a technical description for you.

Your next step

You just directed AI for the first time—instead of accepting whatever it chose to give you.

In the next 15 minutes: take a photo of your own work (a dish, a product, a portrait) and run the same analysis on it. Save the prompt you come up with — it’s your first reusable template.

In the next lesson, you’ll solve the problem that trips up every beginner: making the same person look the same from six different angles without turning into someone else with every generation.

Course 1 · Lesson 2

The method of the grid

By the end of this lesson, you’ll turn a single image into a sheet with six angles of the same scene — same person, same outfit, same light — in a single prompt.

You generate a good image, ask for “the same person from the side”... and get a different person. It’s the wall every beginner hits. Until you solve this, there’s no story in your video—because a story needs the same character from multiple angles.

watch this lesson on video (English · optional)

↓ role to study

01 AI has no memory from one prompt to the next

Every time you ask for a new image, the AI starts from scratch — it doesn’t remember the face it created a minute ago. Asking for “the same woman, now in profile” is like asking an artist who’s never seen her to draw her. That’s why the face changes.

The professionals’ solution is clever: instead of six separate requests, a single request that already brings all six angles together, on the same sheet. Since everything is created in the same generation, the AI keeps the face, clothes, and lighting consistent across all the frames.

For the photographer, it’s the equivalent of a whole shoot on a single contact sheet: instead of six sessions with six similar models, one contact sheet with the same model in six poses.

02 The six-angle grid, in a ready-to-use request

The prompt below is a tested recipe: upload your image from lesson 1, paste the text, and get back a sheet with six photos of the same scene—close-up, medium shot, and wide shot—as if a photographer had walked around the set.

It’s in English because image generators understand the technical terms better that way. You don’t need to translate it or understand it word for word — the entire recipe applies as is.

The restaurant owner can use this tomorrow: one good photo of the main dish becomes six angles of the same dish—from above for the menu, from the side for the post, and a close-up of the details for the ad.

The process, in four steps

  1. Open the generator you already use (ChatGPT itself works) and upload your image.
  2. Set the aspect ratio to square (1:1) — the reason comes in the next step of the lesson.
  3. Paste the ready-to-use prompt from the exercise without changing anything.
  4. Generate and check: same person, same clothes, same light in all six frames.

03 Why the sheet is square

It may seem like a detail, but it’s geometry: every vertical film frame is taller than it is wide. To fit six of those frames side by side on a single sheet, the sheet has to be square—three columns, two rows, everything fitting snugly.

If you leave the aspect ratio vertical, the AI tries to squeeze the grid into a narrow corridor: cropped frames, repeated angles, an unusable sheet.

And there’s a second key in the prompt: it prohibits AI from reinventing. No changing clothes, changing hairstyles, or "improving" the setting. Social media managers know this pain: the client approved a look, and the entire series of posts needs to keep that exact look.

04 The method has limits—and it’s good to know that now

Honesty: in the wider frames of the grid, the face looks small and loses detail. Professional methods for total consistency exist—you’ll see them later in the program. The grid is the method fast and hassle-free to get started today.

Think of it as the working base: the frames that become video in the next lesson come from it. For the photographer, it’s the shoot’s draft, not the final photo delivered to the client — and a good draft is what separates a productive shoot from a wasted afternoon.

Common mistake

Use the tiny face in the wide shot as if it were a close-up. When enlarged, it reveals flaws and breaks the realism you built in lesson 1. For close-ups, use the close-up frame from the grid itself—each frame has its own purpose.

Test yourself

Why do all six angles have the same face when they’re all on the same sheet?

Practice now 0/4 done

Generate your six-angle sheet

Walk away with a grid of six angles of your image from lesson 1—in ~12 minutes.

Your original image stays untouched: the grid is always a new file. Did it come out crooked or repetitive? Generate it again—nothing is lost between attempts.

Create six new 2:3 cinematic images based on the reference scene,
preserving ALL characters, objects, and environment elements exactly
as they appear. Do NOT remove, modify, or reinterpret any character
or object. Only change camera placement, angle, and composition to
create alternate close-up, medium, and wide shots of the same moment.
Maintain consistent lighting direction, atmosphere, color palette,
depth of field, and cinematic style. No redesigns, no new elements,
no stylization changes. The final images must look like alternate
photographs captured during the same professional shoot of the
exact same scene.

You have a sheet of six angles of the same scene—the raw material the video in the next lesson needs.

Summary

  • AI starts from scratch with every prompt — consistency comes from generating all six angles together on a single sheet.
  • The square sheet exists for geometric reasons: it’s the only format where six portrait frames fit without cropping.
  • The ready-to-use prompt tells the AI not to reinvent the outfit, hairstyle, or setting—identity is a rule, not a matter of luck.
  • The grid is a foundation for fast work, with known limits: a small face in a wide shot won’t become a close-up.

Your next step

You just solved the problem that makes most people give up: keeping the same person consistent across six angles.

In the next 15 minutes: run the grid on a photo of your product or workspace and save the sheet — it becomes raw material for the next lessons.

In the next lesson, these still frames start moving—and you learn the trick that keeps the face from “melting” halfway through the video.

Course 1 · Lesson 3

The bridge of movement

By the end of this lesson, you’ll turn two still frames into a 5-second video clip where you decide where the scene begins and ends.

Generating AI video the naive way is a lottery: you press the button and hope the face doesn’t melt along the way. People who work with this don’t hope—they lock down both ends of the journey and make the AI travel between them.

watch this lesson on video (English · optional)

↓ role to study

01 Stop guessing: lock the beginning and the end

The secret of this lesson fits in one sentence: instead of giving an image and hoping for the best, you provide two — the frame where the clip begins and the frame where it ends. The AI only fills in the transition between them.

With both endpoints locked in, the identity has nowhere to drift: the face at the beginning and the face at the end are the same, so the middle stays on track.

The social media manager knows the manual version of this: in a client video, she sets the first and last scenes before anything else—the middle exists to connect them. It’s the same here, except the AI creates the middle.

02 Hand-crop the two frames

Using your six-angle sheet (lesson 2), you’ll crop two frames—for example, the medium shot and the close-up—each in a vertical phone format (9:16 aspect ratio, the same as stories and reels).

Hand-cropping before handing it over to AI is a directing choice: you choose the framing instead of letting the machine cut wherever it wants. The photo editor built into your phone will do the job.

The photographer has always done this: on the contact sheet, she marks the two frames that work with a pen—no one gives the client the whole sheet.

Common mistake

Send the entire grid sheet to the video generator. AI doesn’t know which frame to use and mixes everything together. Always crop first: one file for the beginning, another for the end.

03 In the video prompt, describe the camera — not the person

Leaving the prompt blank is the mistake that makes the video “melt”: without instructions, AI invents movement for everything—face, clothes, background. The professional formula reverses the logic: camera movement + minimal subject action + atmosphere.

A ready-made example, in the spirit of what you’ll paste:

Cinematic slow push-in, smooth zoom on the face,
the character remains stationary, warm lighting.
No morphing, consistent features.

Breaking down the recipe: the camera moves forward slowly, the person barely moves, the light stays warm—and the last two words forbid AI from distorting the face.

Don’t know camera terms? Ask the chat AI: "write 3 camera movement prompts in English to connect a medium shot to a close-up in 5 seconds". A restaurant owner doesn’t need to study filmmaking — they need to know who to ask.

04 The chain: clips that join together on their own

A 5-second clip is too short to tell a story. The solution is called a chain: the last frame from clip A becomes the first frame from clip B. Since the splice happens on an identical frame, the cut is invisible — and the sequence can grow without limit.

For the restaurant owner, it’s one continuous tour: clip 1 moves through the dining room and ends at the kitchen door; clip 2 starts at that exact door and reveals the lit stove.

Building the chain

  1. Generate clip A with a start frame and an end frame (the method from this lesson).
  2. Save the last frame of clip A as an image (the tool itself lets you do this, or you can take a screenshot at the endpoint).
  3. Use this image as a frame initial from clip B — and choose a new final frame.
  4. Repeat as many times as the story calls for: A connects to B, B connects to C.

Practice now 0/4 done

Build your first motion bridge

Walk away with a ~5-second clip connecting two of your frames—in ~15 minutes.

Your original images aren’t altered—the video is always a new file. If the movement looks strange, simplify the prompt and generate again; video tools offer free plans with a few generations per day, so mistakes are part of the budget.

You decided where the scene begins and ends—the AI just handled the transition. That’s directing, not gambling.

Summary

  • AI video stops being a lottery when you lock down both ends: the opening frame and the closing frame.
  • Hand cropping, before AI, is where you direct the scene—framing is your choice, not the machine’s.
  • In the prompt, the camera moves; the subject stays almost still and the atmosphere remains unchanged.
  • Clips are joined in a chain by reusing the final frame of one as the opening frame of the next—and the cut disappears.

Your next step

You just made your first directed video—with a beginning and ending you chose.

Over the next 15 minutes: generate clip B in your chain, starting from the last frame of the clip you just created. Two clips joined together already make a scene.

In the next lesson, your clips become a video with rhythm and sound—and you learn why constant speed puts the audience to sleep.

Course 1 · Lesson 4

Editing with rhythm

By the end of this lesson, you’ll stitch three separate clips into one video, with speed changes and sound—without anyone noticing the seams.

A folder full of AI clips isn’t a film — it’s well-generated digital clutter. Editing turns the material into a story: rhythm, speed, and sound. And you can learn that in a free editor in an afternoon.

watch this lesson on video (English · optional)

↓ role to study

01 Pacing is math, not talent

If every segment of your video moves at the same speed, the audience falls asleep—it’s physiology, not opinion. The human eye wakes up with contrast: slow to show, fast to cut, slow again to breathe.

The structure you’ll use is a three-part formula: slow entry → acceleration during the cut → slow resolution. Every cut between clips happens inside the fast section, where the eye can’t inspect it.

Social media has seen this countless times without putting a name to it: the videos that hold attention to the end are the ones that change pace; the ones that lose viewers in three seconds race along flatly, at one speed.

02 The speed curve in practice

In CapCut—the free editor for your phone or computer—this formula has a button name: speed curve. You build the variation into the clip itself: start at 1x, speed up to 5x or 10x near the cut, and the next clip drops back to normal.

The photographer uses exactly this when making a behind-the-scenes video: the scene’s setup runs in fast motion, then the final click plays at normal speed—the audience feels the climax without knowing why.

Applying the curve

  1. Tap the clip in the timeline and open Speed → Curve.
  2. Choose the model Editing — it already includes the slow-fast-slow motion pattern.
  3. Drag the points so the speed ramp peaks right on the cut.
  4. Repeat on the next clip, reversing it: fast at the start, slow for the rest.

03 Sound is half the illusion

The image catches the eye; sound holds attention. A silent video, no matter how beautiful, feels unfinished. The minimum recipe has two layers: one background track e two sound effects marking the transitions — a breath during the acceleration, a beat at the cut.

And the golden rule of mixing: when the effect hits, the music lower. Both things at full volume become noise; the effect over muffled music becomes impact.

Don’t know what effect to use? Describe the video to the chat AI and ask: "suggest 5 sound effects and where to place each one on the timeline". The restaurant owner did this for the dish video — they even got a tip to add the sizzle of the frying pan at the opening.

04 The right screen from the very first click

Before dragging in any clips, define the project format: 9:16, the vertical phone portrait—the same one as your clips from previous lessons. In CapCut, this is under Aspect Ratio (or Format) when you create the project.

For export, the suggested default is fine: 1080p resolution and everything else as is. Export, watch it on your phone, and ask yourself: can you see the seams between the clips? If not, the lesson delivered on its promise.

The social media manager learned this order the hard way: a finished edit in the wrong format means reframing every shot.

Common mistake

Assemble it first, then adjust the format. Changing the aspect ratio at the end reframes everything: cropped heads, text off-screen. Decide on the format when creating the project, before the first clip.

Practice now 0/4 done

Build your first rhythmic video

Walk away with one exported video made from your clips from lessons 2 and 3—in ~15 minutes.

The editor never changes the original files: it creates a working copy. Got the curve or the cut wrong? Undo it or delete the project and start again—your clips will still be where they always were.

You put separate clips together into a video with rhythm and sound—the edit is there, but only you know where it is.

Summary

  • Editing turns clips into a story—and its driving force is the contrast in speed, not speed itself.
  • The slow-fast-slow curve hides the joins in the accelerated sections, where the eye doesn’t inspect closely.
  • Sound has two layers with a clear hierarchy: background music that leaves room, and effects that mark transitions.
  • Choose the format when you create the project—9:16 from the first clip through export, never converted at the end.

Your next step

You just completed the full cycle: image, motion, and editing—from scratch to exported video.

In the next 15 minutes: show the video to someone and watch where they look away — that’s where a sound effect is missing or a slow second is dragging. Make a note and adjust.

In the next lesson, Part 2 begins: you build a cinematic character in nine layers—the protagonist of your first mini-film.

Course 1 · Lesson 5

Your first cinematic character cinema

By the end of this lesson, you’ll build a cinematic vertical character, layer by layer—strong enough to carry the mini-film you’ll finish in lesson 8.

A shallow prompt returns a shallow face: beautiful, forgettable, like millions of others. A character who can carry a film doesn’t come from one sentence — they come from stacked decisions: who they are, what their skin looks like, what they wear, what they feel, where they are, and what light illuminates them. That’s what you’ll stack today.

watch this lesson on video (English · optional)

↓ role to study

01 Generic prompt, generic face

Write “beautiful woman, realistic photo” and the AI gives you exactly what you asked for: the average of all the beautiful faces it knows. Average has no story. The problem isn’t aesthetic—it’s narrative: no one cares what happens to a stock-photo face.

The social media manager sees it in the numbers: a post with an “AI-perfect” face gets scrolled past; a post with a character who seems alive—a laugh line, an outfit with a story—makes people stop scrolling.

You don’t prompt for a character. You build one.

02 The nine layers, in four questions

Professional construction stacks nine layers. It seems like a lot, but they answer just four questions you already know how to ask:

  • Who is it? — identity (age, features, presence), realistic skin texture (pores, natural shine, small marks), and wardrobe (materials, colors, accessories).
  • What do you feel? — the emotional energy: calm, confident, mysterious, vulnerable. Emotion changes posture, gaze, and even what kind of light fits.
  • Where is it? — the surroundings. This is the layer that turns a portrait into a movie scene.
  • How is it filmed? — cinematic lighting (late afternoon, backlighting, shadows), camera language (lens, depth, composition), and the vertical 9:16 format.

The photographer recognizes her own process: it’s the mental sequence of a shoot—model, styling, location, light, lens. Only the medium is new: here, everything becomes text.

03 The ninth layer: say what you don’t want

The final layer is a list of prohibitions, written at the end of the prompt: no plastic skin, no computer-generated look, no distorted anatomy, no random objects, text, or watermarks in the image.

It works because when AI is unsure, it tends to fall back on exactly these bad habits—and an explicit prohibition closes the door before it can go in.

The restaurant owner creating his brand’s fictional “chef” can feel the difference right away: without the restrictions, you get a video game chef; with them, a professional who could be in his kitchen tomorrow.

Before

"Mysterious woman in a lighthouse, realistic photo." One sentence, one layer — AI decides the other eight.

After

Nine layers described: the middle-aged sentry, skin with real texture, soaked wool cloak, serene, lighthouse at dusk, golden backlight, long lens, 9:16, prohibitions at the end.

The payoff: from stock-photo face to a protagonist with a film ahead of them—the same tool, one extra minute of writing.

04 The ready-to-use formula—and the protagonist who stays with us

The formula below organizes the nine layers into blanks. Fill in each bracket and paste it into ChatGPT’s image generator (the free version generates images)—or into the generator you already use.

For the next lessons, we created this the lighthouse keeper: a serene middle-aged woman on a stormy coast at dusk. She becomes a storyboard in lesson 6, motion in lesson 7, and a film in lesson 8. You can use her or create your own — there’s only one rule: choose someone you want to see come alive.

The social media manager can create a brand mascot for a client; the photographer can create a model who would be impossible to hire. The character is yours—the method is the same.

Test yourself

Your character came out as a beautiful portrait, but without a cinematic feel. Which layer was probably left out?

Practice now 0/3 done

Build your film’s protagonist

Leave with your cinematic vertical character, generated in 9:16—in ~12 minutes.

It’s just text and generation: nothing is lost, and trying again costs nothing extra. Toss a character that didn’t convince you without hesitation — the method stays with you.

Create a vertical 9:16 ultra-realistic cinematic image of
<quem é: idade, traços, presença>.
The character has realistic skin with visible pores, natural texture
and subtle imperfections, and wears <figurino: materiais, cores, acessórios>.
The character's emotional energy is <emoção: serena, confiante, misteriosa...>.
The scene takes place in <ambiente e detalhes de história>.
Use <luz: golden hour, contraluz, sombras longas...>, realistic shadows,
atmospheric depth, natural color grading, and <câmera: lente, profundidade>.
The image should feel like <referência de gênero de filme>.
Avoid plastic skin, CGI look, cartoon style, distorted anatomy, blurry face,
random objects, extra characters, text, logos, watermarks, or UI elements.

You have a protagonist with presence, a world, and light of their own — built through your decisions, layer by layer.

Summary

  • A shallow prompt returns the average of faces — and an average can’t carry a story.
  • The nine layers answer four questions: who they are, what they feel, where they are, and how they’re filmed.
  • Setting and lighting are the layers that turn a portrait into a scene; the final prohibitions help avoid classic AI pitfalls.
  • The character’s image + text pair is a working document: the storyboard, motion, and film all start from it.

Your next step

You just created a protagonist with an identity of their own—not a randomly generated face.

In the next 15 minutes: generate a second version of the same character, changing only the emotional layer. Compare the two — feeling the layer change the image is what makes the method stick.

In the next lesson, this character stops posing: you plan, beat by beat, the 15 seconds of scene they’ll live through.

Course 1 · Lesson 6

The storyboard timed

By the end of this lesson, you’ll turn your character into a nine-frame sheet with timed beats — the complete plan for a 15-second scene, ready to become a video.

Generating video without a plan means spending attempts and hoping for luck. A good clip isn’t random: someone decided what happens every second before pressing the button. Today, that someone is you—and the decision costs a digital sheet of paper, not an expensive generation.

watch this lesson on video (English · optional)

↓ role to study

01 A beautiful image still isn’t a video

Between the portrait from lesson 5 and a clip that moves people is one thing beginners skip: deciding what happens, second by second. This technique has a film industry name— storyboard — but the idea is practical: it’s the shopping list for your scene.

And pay attention to the detail that changes everything: a storyboard isn’t a gallery of pretty variations. It’s a plan with timed beats, where each frame answers the question, "and now, what changes?"

The social media manager already does this without naming it when she sketches the sequence for a reel on paper: opening scene, turning point, ending. Here, the sketch becomes a professional nine-panel sheet.

02 The golden rule: everything stays still except the story

Across the nine frames, four things stay fixed: the character, the clothes, the world, and the light. Only one thing moves forward: the story. That’s the asymmetry that makes the sheet feel like scenes from the same film—and not nine different films.

The photographer recognizes the discipline behind a good shoot: same model, same outfit, same afternoon light throughout—the pose and intention are what evolve. When she changes the lighting midway, the client can tell “something broke,” even if they can’t say what.

Common mistake

Let creativity change outfits or scenery halfway through the nine frames. It looks like variety, but it kills continuity—the final video ends up looking like a collage. Good variety comes from the angles and the action, never from the identity.

03 The nine beats of 15 seconds

The arc you'll use is a condensed classic — it works for a storm at a lighthouse, a kitchen with the heat turned up, and a shop window:

  • 0–1.5s · Opening silence — very wide shot, calm world.
  • 1.5–3s · Something changes — wind, vapor, light: the atmosphere comes alive.
  • 3–4.5s · The character feels — close-up on the face; the eyes notice.
  • 4.5–6s · A detail moves — fabric, hand, ground texture.
  • 6–7.5s · The tension builds — low angle, the world feels heavy.
  • 7.5–9s · The hidden force — something appears in the distance.
  • 9–10.5s · The mood intensifies — profile; the character turns around.
  • 10.5–12.5s · The reveal — the main event takes place.
  • 12.5–15s · Hero ending — low angle, character standing firm amid the chaos.

In the restaurant owner’s version: the kitchen is quiet, steam rises, the chef checks for doneness, flames lick the pan, and the finished dish arrives at the table like a revelation. Same arc, different world.

04 From the rule to the sheet: generating the storyboard

The generator builds the sheet for you: it takes your character image as a reference and the nine-beat arc as the script, then returns a 3×3 grid with a number, duration, title, and one action line per frame—with black borders like a professional contact sheet.

Your role is director: review the finished sheet frame by frame, checking the golden rule. Same face? Same clothes? Same world and lighting? Is the story moving forward?

The social media manager gets a sales asset here: the sheet approved by the client before before spending on any video generation — end of the conversation, "that wasn't what I imagined."

The complete workflow

  1. Upload your character image (lesson 5) to the generator.
  2. Paste the storyboard prompt from the exercise below.
  3. Check the nine frames against the golden rule—identity stays consistent, story moves forward.
  4. Did the frame go off track? Generate it again and mention the problem (“keep the exact same outfit in all panels”).

Practice now 0/3 done

Plan your character’s 15 seconds

Leave with a timed 3×3 sheet for your scene—in ~12 minutes.

The sheet is digital paper: mistakes here cost nothing—and that’s exactly why it exists. Better to discard ten sheets than waste one video generation.

Create a cinematic 3x3 timed storyboard for a 15-second video
sequence based on the uploaded image. It must look like a professional
film pre-production board: 9 panels in a clean 3x3 grid, black borders,
each panel with panel number, timecode, short shot title and one short
action description. Use the uploaded image as the main character
reference. Preserve the same character identity, facial structure,
outfit, environment, lighting direction, atmosphere and color grading
in all 9 panels. Scene arc and exact timecodes:
1 (0:00-0:01.5) opening silence, extreme wide shot;
2 (0:01.5-0:03) the atmosphere begins to move;
3 (0:03-0:04.5) close-up, the character senses it;
4 (0:04.5-0:06) a small detail moves;
5 (0:06-0:07.5) tension builds, low angle;
6 (0:07.5-0:09) a hidden force appears in the distance;
7 (0:09-0:10.5) side profile, atmosphere surges;
8 (0:10.5-0:12.5) the reveal, wide action shot;
9 (0:12.5-0:15) hero ending, low-angle iconic shot.
The scene concept: <descreva em 2-3 frases a SUA cena: onde o
personagem está, o que muda, o que é revelado no final>.
Avoid redesigning the character, changing outfit or environment,
extra characters, cartoon style, plastic skin, messy text, logos
or watermarks.

Your 15-second scene exists on paper, with timings marked—a decision made before spending on any video generation.

Summary

  • A storyboard is a timed plan, not a gallery of variations—each frame answers “what changes now?”
  • The golden rule locks in the character, clothing, world, and lighting; only the story is allowed to move forward.
  • The nine-beat arc takes the scene from calm to revelation — and works in any world, from a lighthouse to a kitchen.
  • The approved sheet is the scene’s contract: cheap to revise, and it saves the video generations that cost more.

Your next step

You just planned a scene like a professional: on paper, before pressing the button.

Over the next 15 minutes: tell someone the scene from your sheet out loud, frame by frame. Wherever they furrow their brow, the beat is confusing; adjust that frame's description.

In the next lesson, the sheet comes to life: the nine frames become a 15-second video that follows your plan—not luck.

Course 1 · Lesson 7

From the shot to the movement

By the end of this lesson, you’ll generate your complete 15-second scene by following the storyboard—with the chat AI writing the huge technical prompt for you.

A prompt that generates a professional-quality video is a page long: every second described, along with the camera, light, and physics. No one expects you to write that—not even professionals do anymore. They train an assistant to write it. Today, you’ll build your own.

watch this lesson on video (English · optional)

↓ role to study

01 You direct; the assistant types.

A 15-second video with cinematic quality needs a huge prompt: each 1.5-second segment described in detail — what the character does, where the camera goes, how the light behaves. Writing all that by hand takes expert work.

The turning point: you don’t write the prompt—you commission the request. A chat AI, trained with the right instructions, reads your storyboard sheet and writes the entire technical page. You stay in charge of the decisions; it handles the paperwork.

Social media already works with this division of labor: they define the concept and references, then delegate the technical writing — except now the technical writer is free and responds in seconds.

02 The training: paste once, use forever

In a new ChatGPT conversation, you paste in a training text (it’s ready in the exercise). It turns that conversation into a video-prompt specialist with four habits that make all the difference:

  • Read the timings on your sheet — if the storyboard says 0–1,5s, the prompt follows 0–1,5s.
  • Detail each segment — action, camera, lighting, and physics, second by second.
  • Protects the identity — repeat the instructions “same face, no distortion” in every prompt.
  • Close with the technical tail — the anti-defect list (no melting, no flickering, realistic physics).

Save this conversation to your favorites: it’s a reusable tool you can use in every video project from now on. The photographer who puts together a portfolio animates a new scene every week—always in the same trained conversation.

Common mistake

Skipping the training and asking directly, “write a video prompt.” Without the instructions, the AI returns three generic lines—and the video becomes a lottery again. The training is what makes the assistant think like an engineer.

03 Two references go in, one prompt comes out

With the assistant trained, you send the two pieces you created: the character image (lesson 5) and the storyboard sheet (lesson 6). Then you ask: “write the final prompt using both—following the timing on the sheet and preserving the identity.”

Together, the two complement each other: the portrait ensures a high-quality face; the sheet maps out what happens each second. One without the other falls short — the portrait alone has no story, and the sheet alone has faces too small to preserve identity.

The restaurant owner sends the fictional chef’s portrait and the kitchen sheet in nine beats—and gets back a technical page he would never write himself. He doesn’t need to: every decision was his.

04 Generate—and check against the sheet

The generation runs in a tool that supports Seedance 2.0—Higgsfield and CapCut are two options; if you already use another Seedance platform, the same request works there. You upload the two images, paste the request, adjust three options, and generate.

Next comes the director’s job: watch with the sheet beside you. The video isn’t good “because it looks beautiful”—it’s good if it followed the plan. That’s how the photographer checks a finished shoot: the question is never “does it look beautiful?” but “is this what the client approved on the sheet?”

Generate and check

  1. In the video tool, upload the portrait and the sheet as references.
  2. Paste the prompt the assistant wrote without cutting anything.
  3. Set: duration 15s · format 9:16 · high quality. Generate.
  4. Check against the sheet: same face from beginning to end? Same clothing? Does the movement follow the beats? Strong ending?
  5. Did something break? Simplify the action in that section of the request and generate it again.

Practice now 0/4 done

Generate your 15-second scene

Leave with a video of your scene, generated from your storyboard—in ~15 minutes.

The training lives in a chat conversation: it doesn’t change anything on your computer or in your accounts. If you run out of video generations for the day (free plans have a limit), the prompt stays ready and waiting—nothing is lost.

Você é um engenheiro de prompts especialista em Seedance 2.0.
Sua função: escrever pedidos de vídeo LONGOS e DETALHADOS, em inglês,
com qualidade de cinema. Regras:
- Vou enviar 2 imagens: @image1 = personagem (referência estrita de
  identidade), @image2 = storyboard cronometrado. Leia os timecodes
  EXATOS do storyboard e siga cada um.
- Cada trecho de tempo deve ter 4-6 frases: ação do personagem,
  movimento de câmera, luz, atmosfera e física.
- Sempre inclua: "Use @image1 as strict identity reference. Preserve
  exact likeness. Stable face throughout."
- Máximo 3-4 ações no clipe; nunca escreva "cut to"; nada de marcas
  ou celebridades.
- Termine sempre com: "No morphing. No deformation. No flickering.
  Realistic physics. 4K cinematic."
- Entregue o pedido final completo dentro de um bloco de código.
Quando eu enviar as imagens, faça perguntas se faltar algo e depois
escreva o pedido final em inglês, formato 9:16, duração 15s.

Your scene went from the page to video, following your plan—the assistant wrote it, but you directed it.

Summary

  • A professional video prompt is too long to write by hand—and you don’t need to: have a trained assistant write it for you.
  • Train it once and it becomes a permanent tool: exact timing, detail for each segment, protected identity, technical tail.
  • Portrait and storyboard go together because they complement each other—the face in close-up on one side, timing and action on the other.
  • A good video is one that follows the shot list; when it breaks, simplify that section and generate it again.

Your next step

You just generated a scene that follows your plan—the leap that separates someone who presses buttons from someone who directs.

In the next 15 minutes: ask the same assistant for a second version of the prompt with a different beat in the reveal. Generating the variation tomorrow gives you more material to put together in lesson 8.

In the final lesson, three scenes become a film: you cut what doesn’t work, keep the gold, and put together your first complete mini-film.

Course 1 · Lesson 8

From loose clips to a mini-film

By the end of this lesson, you’ll put together a complete mini-film from your generated clips—with a beginning, conflict, and climax—exported and ready to show.

A generated clip is raw material: sometimes the best second is in the middle, the beginning is weak, and the ending works for something else. Anyone who strings three clips together in a row delivers a collage. Anyone who mines the best moments and reassembles them delivers a film. Today, you’ll mine the clips.

watch this lesson on video (English · optional)

↓ role to study

01 Think like an editor, not a collector

The beginner’s mistake at this stage has a name: attachment. They spent generations producing each clip and want to use everything. The professional editor does the opposite—they watch the footage and ask just one thing: which moments serve the story? Everything else, no matter how beautiful, falls flat.

A case to get a feel for the method: Vera, a photographer, generated three scenes of the lighthouse keeper and hated the entire second clip — except for the two seconds when the cape billows in the wind. She cut the remaining thirteen seconds without a second thought. Those exact two seconds became the best transition in her film.

Structure for prospecting: preparation (the world and the threat appear) → conflict (the tension builds) → climax (the strongest moment, and the resolution). Every nugget you save will go in one of these three boxes.

02 Before editing, improve the raw material

AI-generated video loses fine detail — faces, fabric, dust — especially during fast movement. That’s why professionals run clips through a smart upscaling before editing.

The tool mentioned in this track is Topaz, a paid program with models that restore fabric and skin texture. A practical rule: if the result looks too sharp and jagged, switch to the tool’s own general cleanup model.

And the honest version for anyone just starting out: this step is optional. A social media manager who doesn’t want to pay for any tools yet can put the film together with the original clips—the lesson’s method works just as well; the upscaling just adds polish.

03 Break it into pieces, reassemble the story

In the editor, the professional workflow follows an order that saves hours: first, cut each clip into small pieces; then choose what to keep; only then put it together. Anyone who tries to “edit it all at once” is stuck with the clips in their original order.

The restaurant owner edits the film about his fictional chef by cutting the three clips into twelve pieces, throwing out seven—strange movements, empty seconds, broken frames—and reassembling five in the order that works: calm kitchen, rising flame, the exact moment, plating, the table.

The editing workflow

  1. Import the three clips (upscaled or original) into a new 9:16 project.
  2. Break each clip into small pieces at the points where the action changes.
  3. Sort quickly: nugget (use it), maybe (save it), junk (delete it).
  4. Reassemble only the selects in setup → conflict → climax order.
  5. Watch it all the way through once: wherever your eye stumbles, cut one more second.

04 Inherited sound, light polish, export

Good news about audio: today’s video generations already come with their own sound and ambience, often good ones. Listen to the original audio before changing anything — if it works, keep it. Adjust only what’s needed: lower loud peaks, mute bad sections, add an effect where the transition calls for one (the wind and beat from lesson 4 still apply).

For the look, polish is seasoning, not a renovation. A starting point that works: contrast +4, sharpness +10, saturation +4, highlights −7, temperature +2. If the clip has already gone through smart upscaling, cut the sharpness in half or set it to zero.

The photographer recognizes the principle of photo developing: a good adjustment is one nobody notices—the three scenes simply start to look as if they were filmed on the same day. Export in 1080p and watch on your phone, from beginning to end, without pausing.

Practice now 0/5 done

Build your first mini-film

Walk away with an exported 30–60-second mini-film made from your clips—in ~15 minutes of editing.

Your original clips stay untouched in the folder—the editor works on copies. The worst that can happen is an ugly project that you can delete and redo in minutes.

You have a film. Not a test, not a draft: a short film with a story, sound, and polish—made by you, from scratch.

Summary

  • A generated clip is raw material—the film comes from mining it, not from lining up whole clips.
  • The professional workflow saves hours: cut into pieces, sort them, and only then reassemble them into setup, conflict, and climax.
  • Keep the sound inherited from generation when it works; visual polish is seasoning no one should notice.
  • With this lesson, you’ve completed the course from start to finish: character, shot, movement, and editing—the same path professionals take, made shorter.

Your next step

You just finished your first mini-film—from a blank idea to an exported file.

In the next 15 minutes: publish the film on your social media today, as it is. The biggest enemy of anyone who makes it this far isn’t technique — it’s waiting for "good enough," which never comes. Publishing the first one makes all the next ones easier.

In Course 2, you take fine control of the frame: composition, lighting, and character consistency in the two image tools that matter most—the foundation for a leap in quality.