AI Filmmaking Program · Course 1
In eight lessons, you’ll go from square one: create your first realistic image, bring it to life, and put together your first mini-movie — with no technical background needed.
Lessons
Recreate the mood and lighting of a professional photo you admire in an AI-generated image of your own.
Turn a single image into six camera angles of the same scene—with the same person, the same outfit, and the same lighting.
Turn your still image into controlled video: you decide where the scene starts and ends.
Join your clips in a free editor, with speed and sound — and no one will see the seam.
Build a character layer by layer, strong enough to carry an entire film.
Plan a 15-second scene, beat by beat, before spending even one video generation.
Generate your complete 15-second scene following the storyboard — with the chat AI writing the technical prompt for you.
Choose the best moments, reassemble the story, and export your first film—ready to show.
Course 1 · Lesson 1
By the end of this lesson, you’ll be able to recreate the mood and lighting of a professional photo you admire in an AI-generated image—using only what you have today.
Have you ever asked an AI for an image and gotten something that looked like a plastic doll? Most people give up there, thinking the tool is weak. The problem isn’t the tool: no one taught the AI—or you—to look at light. That’s what this lesson fixes.
watch this lesson on video (English · optional)
↓ role to study
Nine out of ten AI-generated images look like cheap plastic. The cause is almost never the AI you chose: it’s the light no one described. In cinema, there’s a rule of thumb — lighting accounts for about 70% of what makes a scene look expensive. The tool accounts for the other 30%.
The photographer knows this by heart: the same face, photographed at noon in a parking lot and near a window in the late afternoon, looks like two different people. The camera is the same. The light is what changed.
It’s the same with AI. When you write a prompt—the so-called prompt — without mentioning lighting, the AI chooses for you. And it picks the dullest lighting possible.
If you describe the light, you direct. If you don’t, you hope.
Here’s the paradox that separates amateurs from professionals: skin with no pores, fabric with no folds, a floor with no marks — all of it screams “computer.” The human eye trusts imperfection because the real world is imperfect.
That’s why people who work seriously with images do the opposite of what instinct suggests: they ask for texture, grain, visible pores, and small signs of wear. They deliberately make the image “worse”—and it starts to look like an expensive photograph.
A social media manager sees this in practice: a post that’s too perfect looks like a bank ad, and the audience scrolls right past. A photo with a real-life texture makes them stop.
Every cinematic result comes from the same three-part recipe: observation (you look at a photo you admire and notice the lighting, lens, and mood), technical description (you turn what you saw into precise words in the prompt) and reference (you provide your own photo for the AI to copy the color palette and style).
A reference image is the most powerful shortcut of the three: instead of trying to explain a color tone with words, you show it.
A restaurant owner can use this tomorrow: take the cover photo of an award-winning restaurant he admires, provide it as a reference, and ask for his dish in that same window light with soft shadows.
Before
"A beautiful photo of a plate of pasta." AI decides everything — and returns a generic stock photo.
After
"Handmade pasta on a deep plate, late-afternoon window light from the side, long shadows, subtle steam, rumpled linen tablecloth." Same tool, magazine-cover result.
Result: zero extra dollars spent—the entire difference came from the description.
"But I don't know how to talk about lenses and shadows" — you don't need to. You'll use a chat AI (ChatGPT works, even in the free version) as a translator: it looks at the reference photo and writes the technical description for you.
The process is simple: you send the photo you admire to ChatGPT along with a ready-made request, and it returns two things — an analysis of the photo in a photographer's language and a technical prompt ready to paste into any image generator.
For the photographer, it’s like having an assistant who can describe any light she shows them. For the social media manager, it means turning any client reference into a technical request in a minute.
The ready-to-use prompt is in the exercise just below. It’s a recipe: you don’t need to understand every word—just swap in the photo.
Test yourself
Your image came out looking plastic. According to this lesson’s rule, what’s the first thing to suspect?
Practice now 0/4 done
Walk away with an AI-generated image of yours in the mood of a reference you chose—in ~10 minutes.
Nothing here changes your files or costs anything: you’re only talking with the AI. If the result looks strange, just generate it again—each try is free and independent.
Analise esta imagem como um fotógrafo profissional e diretor de fotografia. Descreva com precisão: pose, ângulo, expressão, estilo e direção da luz, sombras, lente, profundidade de campo, cores, texturas, roupa, composição e clima geral. Seja fiel ao que está visível na imagem. Depois da análise, escreva um prompt técnico completo, EM INGLÊS, em um único parágrafo, para eu recriar esse mesmo estilo visual em um gerador de imagens — trocando o assunto por: <descreva aqui o SEU assunto: seu produto, seu prato, seu retrato>. Responda em português; só o prompt final vai em inglês.
You recreated the mood of a professional photo in an image of yourself—the promise of this lesson, fulfilled with what you already had on hand.
Summary
Course 1 · Lesson 2
By the end of this lesson, you’ll turn a single image into a sheet with six angles of the same scene — same person, same outfit, same light — in a single prompt.
You generate a good image, ask for “the same person from the side”... and get a different person. It’s the wall every beginner hits. Until you solve this, there’s no story in your video—because a story needs the same character from multiple angles.
watch this lesson on video (English · optional)
↓ role to study
Every time you ask for a new image, the AI starts from scratch — it doesn’t remember the face it created a minute ago. Asking for “the same woman, now in profile” is like asking an artist who’s never seen her to draw her. That’s why the face changes.
The professionals’ solution is clever: instead of six separate requests, a single request that already brings all six angles together, on the same sheet. Since everything is created in the same generation, the AI keeps the face, clothes, and lighting consistent across all the frames.
For the photographer, it’s the equivalent of a whole shoot on a single contact sheet: instead of six sessions with six similar models, one contact sheet with the same model in six poses.
The prompt below is a tested recipe: upload your image from lesson 1, paste the text, and get back a sheet with six photos of the same scene—close-up, medium shot, and wide shot—as if a photographer had walked around the set.
It’s in English because image generators understand the technical terms better that way. You don’t need to translate it or understand it word for word — the entire recipe applies as is.
The restaurant owner can use this tomorrow: one good photo of the main dish becomes six angles of the same dish—from above for the menu, from the side for the post, and a close-up of the details for the ad.
The process, in four steps
It may seem like a detail, but it’s geometry: every vertical film frame is taller than it is wide. To fit six of those frames side by side on a single sheet, the sheet has to be square—three columns, two rows, everything fitting snugly.
If you leave the aspect ratio vertical, the AI tries to squeeze the grid into a narrow corridor: cropped frames, repeated angles, an unusable sheet.
And there’s a second key in the prompt: it prohibits AI from reinventing. No changing clothes, changing hairstyles, or "improving" the setting. Social media managers know this pain: the client approved a look, and the entire series of posts needs to keep that exact look.
Honesty: in the wider frames of the grid, the face looks small and loses detail. Professional methods for total consistency exist—you’ll see them later in the program. The grid is the method fast and hassle-free to get started today.
Think of it as the working base: the frames that become video in the next lesson come from it. For the photographer, it’s the shoot’s draft, not the final photo delivered to the client — and a good draft is what separates a productive shoot from a wasted afternoon.
Common mistake
Use the tiny face in the wide shot as if it were a close-up. When enlarged, it reveals flaws and breaks the realism you built in lesson 1. For close-ups, use the close-up frame from the grid itself—each frame has its own purpose.
Test yourself
Why do all six angles have the same face when they’re all on the same sheet?
Practice now 0/4 done
Walk away with a grid of six angles of your image from lesson 1—in ~12 minutes.
Your original image stays untouched: the grid is always a new file. Did it come out crooked or repetitive? Generate it again—nothing is lost between attempts.
Create six new 2:3 cinematic images based on the reference scene, preserving ALL characters, objects, and environment elements exactly as they appear. Do NOT remove, modify, or reinterpret any character or object. Only change camera placement, angle, and composition to create alternate close-up, medium, and wide shots of the same moment. Maintain consistent lighting direction, atmosphere, color palette, depth of field, and cinematic style. No redesigns, no new elements, no stylization changes. The final images must look like alternate photographs captured during the same professional shoot of the exact same scene.
You have a sheet of six angles of the same scene—the raw material the video in the next lesson needs.
Summary
Course 1 · Lesson 3
By the end of this lesson, you’ll turn two still frames into a 5-second video clip where you decide where the scene begins and ends.
Generating AI video the naive way is a lottery: you press the button and hope the face doesn’t melt along the way. People who work with this don’t hope—they lock down both ends of the journey and make the AI travel between them.
watch this lesson on video (English · optional)
↓ role to study
The secret of this lesson fits in one sentence: instead of giving an image and hoping for the best, you provide two — the frame where the clip begins and the frame where it ends. The AI only fills in the transition between them.
With both endpoints locked in, the identity has nowhere to drift: the face at the beginning and the face at the end are the same, so the middle stays on track.
The social media manager knows the manual version of this: in a client video, she sets the first and last scenes before anything else—the middle exists to connect them. It’s the same here, except the AI creates the middle.
Using your six-angle sheet (lesson 2), you’ll crop two frames—for example, the medium shot and the close-up—each in a vertical phone format (9:16 aspect ratio, the same as stories and reels).
Hand-cropping before handing it over to AI is a directing choice: you choose the framing instead of letting the machine cut wherever it wants. The photo editor built into your phone will do the job.
The photographer has always done this: on the contact sheet, she marks the two frames that work with a pen—no one gives the client the whole sheet.
Common mistake
Send the entire grid sheet to the video generator. AI doesn’t know which frame to use and mixes everything together. Always crop first: one file for the beginning, another for the end.
Leaving the prompt blank is the mistake that makes the video “melt”: without instructions, AI invents movement for everything—face, clothes, background. The professional formula reverses the logic: camera movement + minimal subject action + atmosphere.
A ready-made example, in the spirit of what you’ll paste:
Cinematic slow push-in, smooth zoom on the face, the character remains stationary, warm lighting. No morphing, consistent features.
Breaking down the recipe: the camera moves forward slowly, the person barely moves, the light stays warm—and the last two words forbid AI from distorting the face.
Don’t know camera terms? Ask the chat AI: "write 3 camera movement prompts in English to connect a medium shot to a close-up in 5 seconds". A restaurant owner doesn’t need to study filmmaking — they need to know who to ask.
A 5-second clip is too short to tell a story. The solution is called a chain: the last frame from clip A becomes the first frame from clip B. Since the splice happens on an identical frame, the cut is invisible — and the sequence can grow without limit.
For the restaurant owner, it’s one continuous tour: clip 1 moves through the dining room and ends at the kitchen door; clip 2 starts at that exact door and reveals the lit stove.
Building the chain
Practice now 0/4 done
Walk away with a ~5-second clip connecting two of your frames—in ~15 minutes.
Your original images aren’t altered—the video is always a new file. If the movement looks strange, simplify the prompt and generate again; video tools offer free plans with a few generations per day, so mistakes are part of the budget.
You decided where the scene begins and ends—the AI just handled the transition. That’s directing, not gambling.
Summary
Course 1 · Lesson 4
By the end of this lesson, you’ll stitch three separate clips into one video, with speed changes and sound—without anyone noticing the seams.
A folder full of AI clips isn’t a film — it’s well-generated digital clutter. Editing turns the material into a story: rhythm, speed, and sound. And you can learn that in a free editor in an afternoon.
watch this lesson on video (English · optional)
↓ role to study
If every segment of your video moves at the same speed, the audience falls asleep—it’s physiology, not opinion. The human eye wakes up with contrast: slow to show, fast to cut, slow again to breathe.
The structure you’ll use is a three-part formula: slow entry → acceleration during the cut → slow resolution. Every cut between clips happens inside the fast section, where the eye can’t inspect it.
Social media has seen this countless times without putting a name to it: the videos that hold attention to the end are the ones that change pace; the ones that lose viewers in three seconds race along flatly, at one speed.
In CapCut—the free editor for your phone or computer—this formula has a button name: speed curve. You build the variation into the clip itself: start at 1x, speed up to 5x or 10x near the cut, and the next clip drops back to normal.
The photographer uses exactly this when making a behind-the-scenes video: the scene’s setup runs in fast motion, then the final click plays at normal speed—the audience feels the climax without knowing why.
Applying the curve
The image catches the eye; sound holds attention. A silent video, no matter how beautiful, feels unfinished. The minimum recipe has two layers: one background track e two sound effects marking the transitions — a breath during the acceleration, a beat at the cut.
And the golden rule of mixing: when the effect hits, the music lower. Both things at full volume become noise; the effect over muffled music becomes impact.
Don’t know what effect to use? Describe the video to the chat AI and ask: "suggest 5 sound effects and where to place each one on the timeline". The restaurant owner did this for the dish video — they even got a tip to add the sizzle of the frying pan at the opening.
Before dragging in any clips, define the project format: 9:16, the vertical phone portrait—the same one as your clips from previous lessons. In CapCut, this is under Aspect Ratio (or Format) when you create the project.
For export, the suggested default is fine: 1080p resolution and everything else as is. Export, watch it on your phone, and ask yourself: can you see the seams between the clips? If not, the lesson delivered on its promise.
The social media manager learned this order the hard way: a finished edit in the wrong format means reframing every shot.
Common mistake
Assemble it first, then adjust the format. Changing the aspect ratio at the end reframes everything: cropped heads, text off-screen. Decide on the format when creating the project, before the first clip.
Practice now 0/4 done
Walk away with one exported video made from your clips from lessons 2 and 3—in ~15 minutes.
The editor never changes the original files: it creates a working copy. Got the curve or the cut wrong? Undo it or delete the project and start again—your clips will still be where they always were.
You put separate clips together into a video with rhythm and sound—the edit is there, but only you know where it is.
Summary
Course 1 · Lesson 5
By the end of this lesson, you’ll build a cinematic vertical character, layer by layer—strong enough to carry the mini-film you’ll finish in lesson 8.
A shallow prompt returns a shallow face: beautiful, forgettable, like millions of others. A character who can carry a film doesn’t come from one sentence — they come from stacked decisions: who they are, what their skin looks like, what they wear, what they feel, where they are, and what light illuminates them. That’s what you’ll stack today.
watch this lesson on video (English · optional)
↓ role to study
Write “beautiful woman, realistic photo” and the AI gives you exactly what you asked for: the average of all the beautiful faces it knows. Average has no story. The problem isn’t aesthetic—it’s narrative: no one cares what happens to a stock-photo face.
The social media manager sees it in the numbers: a post with an “AI-perfect” face gets scrolled past; a post with a character who seems alive—a laugh line, an outfit with a story—makes people stop scrolling.
You don’t prompt for a character. You build one.
Professional construction stacks nine layers. It seems like a lot, but they answer just four questions you already know how to ask:
The photographer recognizes her own process: it’s the mental sequence of a shoot—model, styling, location, light, lens. Only the medium is new: here, everything becomes text.
The final layer is a list of prohibitions, written at the end of the prompt: no plastic skin, no computer-generated look, no distorted anatomy, no random objects, text, or watermarks in the image.
It works because when AI is unsure, it tends to fall back on exactly these bad habits—and an explicit prohibition closes the door before it can go in.
The restaurant owner creating his brand’s fictional “chef” can feel the difference right away: without the restrictions, you get a video game chef; with them, a professional who could be in his kitchen tomorrow.
Before
"Mysterious woman in a lighthouse, realistic photo." One sentence, one layer — AI decides the other eight.
After
Nine layers described: the middle-aged sentry, skin with real texture, soaked wool cloak, serene, lighthouse at dusk, golden backlight, long lens, 9:16, prohibitions at the end.
The payoff: from stock-photo face to a protagonist with a film ahead of them—the same tool, one extra minute of writing.
The formula below organizes the nine layers into blanks. Fill in each bracket and paste it into ChatGPT’s image generator (the free version generates images)—or into the generator you already use.
For the next lessons, we created this the lighthouse keeper: a serene middle-aged woman on a stormy coast at dusk. She becomes a storyboard in lesson 6, motion in lesson 7, and a film in lesson 8. You can use her or create your own — there’s only one rule: choose someone you want to see come alive.
The social media manager can create a brand mascot for a client; the photographer can create a model who would be impossible to hire. The character is yours—the method is the same.
Test yourself
Your character came out as a beautiful portrait, but without a cinematic feel. Which layer was probably left out?
Practice now 0/3 done
Leave with your cinematic vertical character, generated in 9:16—in ~12 minutes.
It’s just text and generation: nothing is lost, and trying again costs nothing extra. Toss a character that didn’t convince you without hesitation — the method stays with you.
Create a vertical 9:16 ultra-realistic cinematic image of <quem é: idade, traços, presença>. The character has realistic skin with visible pores, natural texture and subtle imperfections, and wears <figurino: materiais, cores, acessórios>. The character's emotional energy is <emoção: serena, confiante, misteriosa...>. The scene takes place in <ambiente e detalhes de história>. Use <luz: golden hour, contraluz, sombras longas...>, realistic shadows, atmospheric depth, natural color grading, and <câmera: lente, profundidade>. The image should feel like <referência de gênero de filme>. Avoid plastic skin, CGI look, cartoon style, distorted anatomy, blurry face, random objects, extra characters, text, logos, watermarks, or UI elements.
You have a protagonist with presence, a world, and light of their own — built through your decisions, layer by layer.
Summary
Course 1 · Lesson 6
By the end of this lesson, you’ll turn your character into a nine-frame sheet with timed beats — the complete plan for a 15-second scene, ready to become a video.
Generating video without a plan means spending attempts and hoping for luck. A good clip isn’t random: someone decided what happens every second before pressing the button. Today, that someone is you—and the decision costs a digital sheet of paper, not an expensive generation.
watch this lesson on video (English · optional)
↓ role to study
Between the portrait from lesson 5 and a clip that moves people is one thing beginners skip: deciding what happens, second by second. This technique has a film industry name— storyboard — but the idea is practical: it’s the shopping list for your scene.
And pay attention to the detail that changes everything: a storyboard isn’t a gallery of pretty variations. It’s a plan with timed beats, where each frame answers the question, "and now, what changes?"
The social media manager already does this without naming it when she sketches the sequence for a reel on paper: opening scene, turning point, ending. Here, the sketch becomes a professional nine-panel sheet.
Across the nine frames, four things stay fixed: the character, the clothes, the world, and the light. Only one thing moves forward: the story. That’s the asymmetry that makes the sheet feel like scenes from the same film—and not nine different films.
The photographer recognizes the discipline behind a good shoot: same model, same outfit, same afternoon light throughout—the pose and intention are what evolve. When she changes the lighting midway, the client can tell “something broke,” even if they can’t say what.
Common mistake
Let creativity change outfits or scenery halfway through the nine frames. It looks like variety, but it kills continuity—the final video ends up looking like a collage. Good variety comes from the angles and the action, never from the identity.
The arc you'll use is a condensed classic — it works for a storm at a lighthouse, a kitchen with the heat turned up, and a shop window:
In the restaurant owner’s version: the kitchen is quiet, steam rises, the chef checks for doneness, flames lick the pan, and the finished dish arrives at the table like a revelation. Same arc, different world.
The generator builds the sheet for you: it takes your character image as a reference and the nine-beat arc as the script, then returns a 3×3 grid with a number, duration, title, and one action line per frame—with black borders like a professional contact sheet.
Your role is director: review the finished sheet frame by frame, checking the golden rule. Same face? Same clothes? Same world and lighting? Is the story moving forward?
The social media manager gets a sales asset here: the sheet approved by the client before before spending on any video generation — end of the conversation, "that wasn't what I imagined."
The complete workflow
Practice now 0/3 done
Leave with a timed 3×3 sheet for your scene—in ~12 minutes.
The sheet is digital paper: mistakes here cost nothing—and that’s exactly why it exists. Better to discard ten sheets than waste one video generation.
Create a cinematic 3x3 timed storyboard for a 15-second video sequence based on the uploaded image. It must look like a professional film pre-production board: 9 panels in a clean 3x3 grid, black borders, each panel with panel number, timecode, short shot title and one short action description. Use the uploaded image as the main character reference. Preserve the same character identity, facial structure, outfit, environment, lighting direction, atmosphere and color grading in all 9 panels. Scene arc and exact timecodes: 1 (0:00-0:01.5) opening silence, extreme wide shot; 2 (0:01.5-0:03) the atmosphere begins to move; 3 (0:03-0:04.5) close-up, the character senses it; 4 (0:04.5-0:06) a small detail moves; 5 (0:06-0:07.5) tension builds, low angle; 6 (0:07.5-0:09) a hidden force appears in the distance; 7 (0:09-0:10.5) side profile, atmosphere surges; 8 (0:10.5-0:12.5) the reveal, wide action shot; 9 (0:12.5-0:15) hero ending, low-angle iconic shot. The scene concept: <descreva em 2-3 frases a SUA cena: onde o personagem está, o que muda, o que é revelado no final>. Avoid redesigning the character, changing outfit or environment, extra characters, cartoon style, plastic skin, messy text, logos or watermarks.
Your 15-second scene exists on paper, with timings marked—a decision made before spending on any video generation.
Summary
Course 1 · Lesson 7
By the end of this lesson, you’ll generate your complete 15-second scene by following the storyboard—with the chat AI writing the huge technical prompt for you.
A prompt that generates a professional-quality video is a page long: every second described, along with the camera, light, and physics. No one expects you to write that—not even professionals do anymore. They train an assistant to write it. Today, you’ll build your own.
watch this lesson on video (English · optional)
↓ role to study
A 15-second video with cinematic quality needs a huge prompt: each 1.5-second segment described in detail — what the character does, where the camera goes, how the light behaves. Writing all that by hand takes expert work.
The turning point: you don’t write the prompt—you commission the request. A chat AI, trained with the right instructions, reads your storyboard sheet and writes the entire technical page. You stay in charge of the decisions; it handles the paperwork.
Social media already works with this division of labor: they define the concept and references, then delegate the technical writing — except now the technical writer is free and responds in seconds.
In a new ChatGPT conversation, you paste in a training text (it’s ready in the exercise). It turns that conversation into a video-prompt specialist with four habits that make all the difference:
Save this conversation to your favorites: it’s a reusable tool you can use in every video project from now on. The photographer who puts together a portfolio animates a new scene every week—always in the same trained conversation.
Common mistake
Skipping the training and asking directly, “write a video prompt.” Without the instructions, the AI returns three generic lines—and the video becomes a lottery again. The training is what makes the assistant think like an engineer.
With the assistant trained, you send the two pieces you created: the character image (lesson 5) and the storyboard sheet (lesson 6). Then you ask: “write the final prompt using both—following the timing on the sheet and preserving the identity.”
Together, the two complement each other: the portrait ensures a high-quality face; the sheet maps out what happens each second. One without the other falls short — the portrait alone has no story, and the sheet alone has faces too small to preserve identity.
The restaurant owner sends the fictional chef’s portrait and the kitchen sheet in nine beats—and gets back a technical page he would never write himself. He doesn’t need to: every decision was his.
The generation runs in a tool that supports Seedance 2.0—Higgsfield and CapCut are two options; if you already use another Seedance platform, the same request works there. You upload the two images, paste the request, adjust three options, and generate.
Next comes the director’s job: watch with the sheet beside you. The video isn’t good “because it looks beautiful”—it’s good if it followed the plan. That’s how the photographer checks a finished shoot: the question is never “does it look beautiful?” but “is this what the client approved on the sheet?”
Generate and check
Practice now 0/4 done
Leave with a video of your scene, generated from your storyboard—in ~15 minutes.
The training lives in a chat conversation: it doesn’t change anything on your computer or in your accounts. If you run out of video generations for the day (free plans have a limit), the prompt stays ready and waiting—nothing is lost.
Você é um engenheiro de prompts especialista em Seedance 2.0. Sua função: escrever pedidos de vídeo LONGOS e DETALHADOS, em inglês, com qualidade de cinema. Regras: - Vou enviar 2 imagens: @image1 = personagem (referência estrita de identidade), @image2 = storyboard cronometrado. Leia os timecodes EXATOS do storyboard e siga cada um. - Cada trecho de tempo deve ter 4-6 frases: ação do personagem, movimento de câmera, luz, atmosfera e física. - Sempre inclua: "Use @image1 as strict identity reference. Preserve exact likeness. Stable face throughout." - Máximo 3-4 ações no clipe; nunca escreva "cut to"; nada de marcas ou celebridades. - Termine sempre com: "No morphing. No deformation. No flickering. Realistic physics. 4K cinematic." - Entregue o pedido final completo dentro de um bloco de código. Quando eu enviar as imagens, faça perguntas se faltar algo e depois escreva o pedido final em inglês, formato 9:16, duração 15s.
Your scene went from the page to video, following your plan—the assistant wrote it, but you directed it.
Summary
Course 1 · Lesson 8
By the end of this lesson, you’ll put together a complete mini-film from your generated clips—with a beginning, conflict, and climax—exported and ready to show.
A generated clip is raw material: sometimes the best second is in the middle, the beginning is weak, and the ending works for something else. Anyone who strings three clips together in a row delivers a collage. Anyone who mines the best moments and reassembles them delivers a film. Today, you’ll mine the clips.
watch this lesson on video (English · optional)
↓ role to study
The beginner’s mistake at this stage has a name: attachment. They spent generations producing each clip and want to use everything. The professional editor does the opposite—they watch the footage and ask just one thing: which moments serve the story? Everything else, no matter how beautiful, falls flat.
A case to get a feel for the method: Vera, a photographer, generated three scenes of the lighthouse keeper and hated the entire second clip — except for the two seconds when the cape billows in the wind. She cut the remaining thirteen seconds without a second thought. Those exact two seconds became the best transition in her film.
Structure for prospecting: preparation (the world and the threat appear) → conflict (the tension builds) → climax (the strongest moment, and the resolution). Every nugget you save will go in one of these three boxes.
AI-generated video loses fine detail — faces, fabric, dust — especially during fast movement. That’s why professionals run clips through a smart upscaling before editing.
The tool mentioned in this track is Topaz, a paid program with models that restore fabric and skin texture. A practical rule: if the result looks too sharp and jagged, switch to the tool’s own general cleanup model.
And the honest version for anyone just starting out: this step is optional. A social media manager who doesn’t want to pay for any tools yet can put the film together with the original clips—the lesson’s method works just as well; the upscaling just adds polish.
In the editor, the professional workflow follows an order that saves hours: first, cut each clip into small pieces; then choose what to keep; only then put it together. Anyone who tries to “edit it all at once” is stuck with the clips in their original order.
The restaurant owner edits the film about his fictional chef by cutting the three clips into twelve pieces, throwing out seven—strange movements, empty seconds, broken frames—and reassembling five in the order that works: calm kitchen, rising flame, the exact moment, plating, the table.
The editing workflow
Good news about audio: today’s video generations already come with their own sound and ambience, often good ones. Listen to the original audio before changing anything — if it works, keep it. Adjust only what’s needed: lower loud peaks, mute bad sections, add an effect where the transition calls for one (the wind and beat from lesson 4 still apply).
For the look, polish is seasoning, not a renovation. A starting point that works: contrast +4, sharpness +10, saturation +4, highlights −7, temperature +2. If the clip has already gone through smart upscaling, cut the sharpness in half or set it to zero.
The photographer recognizes the principle of photo developing: a good adjustment is one nobody notices—the three scenes simply start to look as if they were filmed on the same day. Export in 1080p and watch on your phone, from beginning to end, without pausing.
Practice now 0/5 done
Walk away with an exported 30–60-second mini-film made from your clips—in ~15 minutes of editing.
Your original clips stay untouched in the folder—the editor works on copies. The worst that can happen is an ugly project that you can delete and redo in minutes.
You have a film. Not a test, not a draft: a short film with a story, sound, and polish—made by you, from scratch.
Summary