PTENES
Skip to content
MODULE 4 TECHNICAL LEVEL

Multimodal Prompting (Image and Video)

Master the art of creating prompts for multimodal models. Learn to generate, interpret, and control visual content with technical precision.

7
Topics
90
Minutes
12
Exercises
1

Multimodal Prompt Interpretation

How models process different types of input

Multimodal models represent a significant evolution in AI, capable of processing and integrating information from different modalities—text, images, audio, and video. Understanding how these models interpret combined inputs is essential to creating effective prompts.

Vision + Text

  • • Image analysis with context
  • • OCR and information extraction
  • • Visual description and interpretation
  • • Image comparison

Audio + Context

  • • Transcription with sentiment analysis
  • • Speaker identification
  • • Information extraction from audio
  • • Tone and emotion analysis

Interpretation Strategy

When working with multimodal inputs, always provide clear context about what you expect the model to do with each type of input.

# Example of an Effective Multimodal Prompt

[Image: dashboard screenshot]

Analyze this dashboard screenshot and:
1. Identify the main KPIs shown
2. Describe Visible Trends in the Charts
3. Suggest 3 business insights based on the data

2

Prompts for Image Generation

Structure and techniques for visual creation

Generating images via prompts is an art that combines technical description with creativity. The prompt structure directly influences the quality and relevance of the generated images.

Anatomy of an Image Prompt

1
Main Subject

What or who is in the image

2
Visual Style

Photorealistic, illustration, 3D, watercolor, etc.

3
Composition

Angle, framing, perspective

4
Lighting

Natural, studio, dramatic, soft

5
Technical Details

Resolution, quality, specific parameters

Example of a Structured Prompt

// Subject

A scientist working in a futuristic laboratory,

// Style

cyberpunk style with influences from Syd Mead,

// Composition

medium shot, slightly low angle,

// Lighting

purple and blue neon lighting, high contrast,

// Technical

8k, ultra detailed, ray tracing

3

Negative Prompts in Images

Controlling what should NOT appear

Negative prompts are just as important as positive ones in image generation. They let you exclude unwanted elements, fix common artifacts, and refine the final quality.

Positive Prompt

Professional portrait, neutral background, studio lighting, high resolution, sharp focus

Negative Prompt

Distortions, extra hands, extra fingers, crossed eyes, blurry, low quality, watermark

Negative Prompt Categories

Technical Quality

blurry, low resolution, pixelated, jpeg artifacts, noise, grainy, overexposed, underexposed

Anatomy and Proportions

extra limbs, missing fingers, deformed hands, unnatural poses, distorted faces, asymmetrical

Undesirable Elements

watermark, signature, text overlay, frame, border, logo, username, cropped

Specific Style

cartoon (if you want realism), realistic (if you want illustration), anime, photobash, stock photo look

Important Tip

Not all image generation models support negative prompts in the same way. Check the specific documentation for each tool (DALL-E, Midjourney, Stable Diffusion).

4

Visual Consistency

Maintaining Identity Across Multiple Generations

One of the biggest challenges in AI visual projects is maintaining consistency across multiple images. Whether you’re creating a recurring character, a series of illustrations, or brand materials, consistency is crucial.

Consistency Techniques

🎨 Style Reference

Use reference images to establish the base visual style and maintain consistency in new generations.

📝 Prompt Templates

Create reusable prompt templates with fixed and variable elements for each new image.

🔢 Fixed Seeds

When available, use the same seed to maintain similar characteristics across variations.

📋 Character Sheets

Create detailed reference documents describing each element that must remain consistent.

Example: Character Template

# Character: Maya

## Fixed Characteristics:

- Hair: short, black, pixie style

- Eyes: almond-shaped, amber

- Style: casual cyberpunk

- Accessory: gold hoop earrings

## Variables per Scene:

- Clothing: [INSERT]

- Pose: [INSERT]

- Scenario: [INSERT]

5

Prompting for Videos

Adding the time dimension

AI video generation adds significant complexity to prompting because, in addition to static visual elements, you need to describe movement, transitions, and timing.

Elements of a Video Prompt

Camera Movement

Pan, zoom, dolly, tracking shot, steadicam

Subject's Action

Describe the movement: walking, running, gesturing

Pace and Duration

Slow, fast, time-lapse, slow motion

Atmosphere and Mood

Cinematic, documentary, dreamlike, energetic

Video Prompt Example

Slow dolly shot

approaching a steaming cup of coffee

on a rustic wooden table,

golden morning light

coming in through the window,

particles of dust floating in the air,

cinematic style,

shallow depth of field, 24fps, 35mm film

6

Storyboards via Prompt

AI-powered narrative visual planning

Storyboards are essential tools for previewing audiovisual projects. With AI, you can quickly create scene visualizations to validate concepts before production.

AI Storyboard Workflow

1

Define the Narrative Structure

Divide your story into key scenes and important dramatic moments

2

Create a Base Prompt

Establish a visual style, aspect ratio, and recurring elements

3

Generate Keyframes

Create images for each scene using prompts specific to the framing

4

Take Notes and Organize

Add direction notes, dialogue, and technical instructions

Example: Storyboard Scene Prompt

// SCENE 3 - REVEAL

Framing: Close-up

Close-up of the protagonist’s (Maya’s) face,
expression of surprise changing to determination,
dramatic side lighting creating shadows,
blurred background with neon lights,
16:9 aspect ratio, cinematic storyboard style,
clean lines, black and white with touches of color

Note: Transition to scene 4 with zoom out

7

Practical Applications

Real-world use cases across different industries

Multimodal prompting has applications across various industries, from marketing to education. Knowing these use cases helps you identify opportunities in your field.

🎨

Marketing & Advertising

  • • Campaign asset creation
  • • Product mockups
  • • Ad variations for A/B testing
  • • Social media content
🎬

Entertainment

  • • Game concept art
  • • Storyboards for films
  • • Iterative character design
  • • Scenario visualization
📚

Education

  • • Educational illustrations
  • • Custom infographics
  • • Visual teaching materials
  • • Historical representations
🏢

Architecture & Design

  • • Project visualization
  • • Automated mood boards
  • • Interior design variations
  • • Fast rendering

Final Hands-On Exercise

Choose a real project from your field and create:

  1. A base prompt with a defined visual style
  2. A negative prompt appropriate to the context
  3. 3 prompt variations for different scenarios
  4. A reusable template for future generations

💡 Tip: Start with something simple and gradually add complexity. Document what works and what doesn’t to create your own visual prompt “playbook.”

Module 3: Chains and Processes Module 5: System Prompts