M5-A — Text, Image, and Video AI

How to read this note

Until now, AI mostly meant text. Module 5 opens the door: AI can also draw, design, and make videos — if you know which tool to use and how to describe what you want.


🎨 What is "multimodal" AI?

Modal = a type of input or output.

Mode Input example Output example
Text "Write an essay on pollution" Paragraphs
Image "Poster of a cold drink can" PNG / JPG picture
Video "10-second ad of a shoe" MP4 clip
Audio "Read this paragraph aloud" Voice file

Multimodal = AI that handles more than one mode.

Text only     →  ChatGPT (classic chat)
Text → Image  →  Gemini, DALL·E, Midjourney, Ideogram
Text → Video  →  Veo, Runway, Pika, Sora (where available)
Image → Text  →  "What is in this photo?"
Image → Image →  "Make this product photo look professional"
The real insight

Different modes = different AI models trained on different data. You are not "talking to one brain that does everything." You are picking the right specialist tool for the job.


✍️ Text AI — the base layer

You already know this from Modules 2–4.

Your prompt  →  Language model  →  Written answer

Good for: emails, summaries, scripts, captions, code explanations, study notes.

Tools: ChatGPT, Claude, Gemini, Copilot.


🖼️ Image AI — how it works (simple backend picture)

When you ask for an image, this is roughly what happens:

Your text prompt
      ↓
App sends prompt to image model (on company's servers)
      ↓
Model has learned patterns from millions of images + captions
      ↓
It generates pixels that match your description
      ↓
Image file comes back to your screen

The art-school analogy

Imagine an artist who has seen millions of posters, photos, and ads — but never "understood" them like a human. You describe what you want. They paint something that looks like things they've seen before, mixed together.

That is text-to-image AI.

You control You cannot control (easily)
Style words ("cinematic", "minimal") Exact pixel-perfect logo placement
Subject ("silver can, water drops") Real brand trademarks sometimes
Mood ("luxury", "playful") Consistent same face every time (without extra tools)

Tools: Gemini (Nano Banana / image gen), ChatGPT + DALL·E, Adobe Firefly, Canva AI.

Pro move from Module 3

Use text AI to write the image prompt, then paste into image AI. Two-step chain = better posters.


🎬 Video AI — how it works (simple)

Video AI is harder than images because it must make many frames that move smoothly.

Your prompt (+ sometimes a starting image)
      ↓
Video model generates frame 1, frame 2, frame 3 …
      ↓
Frames stitched into a short clip (often 4–10 seconds)
      ↓
You download MP4
Text AI Image AI Video AI
Fast Medium speed Slowest
Cheapest / often free tier Some free limits Usually limited free
Easy to fix (edit words) Regenerate image Regenerate whole clip

Tools: Google Veo, Runway, Pika, CapCut AI, some features inside Canva/Gamma.

Video AI is still early

Expect weird hands, flickering, wrong text on screen. Use for drafts and ideas, not final TV ads — unless you review frame by frame.


🧰 Which tool for which job?

Job Start here
Write ad copy ChatGPT / Gemini (text)
Product poster Gemini image or Prompt Templates workflow
Instagram reel idea Text AI writes script → Video AI makes clip
Presentation slides Gamma.app (text → slides)
Remove background from photo Canva / dedicated tools

🎯 What the brochure means by "Multimodal AI Systems"

It does not mean "use every tool once."

It means: design workflows where text, image, and video steps connect:

Product photo  →  Text AI writes poster prompt  →  Image AI makes poster
       ↓
Text AI writes caption  →  You post on Instagram

That connected chain is a multimodal system — even if you run it by hand today.


Next: M5-B — AI Content Pipelines — how creators and businesses run these chains every day.

Practical: Prompt Templates — copy-paste poster and photoshoot prompts