M5-A — Text, Image, and Video AI
Until now, AI mostly meant text. Module 5 opens the door: AI can also draw, design, and make videos — if you know which tool to use and how to describe what you want.
🎨 What is "multimodal" AI?
Modal = a type of input or output.
| Mode | Input example | Output example |
|---|---|---|
| Text | "Write an essay on pollution" | Paragraphs |
| Image | "Poster of a cold drink can" | PNG / JPG picture |
| Video | "10-second ad of a shoe" | MP4 clip |
| Audio | "Read this paragraph aloud" | Voice file |
Multimodal = AI that handles more than one mode.
Text only → ChatGPT (classic chat)
Text → Image → Gemini, DALL·E, Midjourney, Ideogram
Text → Video → Veo, Runway, Pika, Sora (where available)
Image → Text → "What is in this photo?"
Image → Image → "Make this product photo look professional"
Different modes = different AI models trained on different data. You are not "talking to one brain that does everything." You are picking the right specialist tool for the job.
✍️ Text AI — the base layer
You already know this from Modules 2–4.
Your prompt → Language model → Written answer
Good for: emails, summaries, scripts, captions, code explanations, study notes.
Tools: ChatGPT, Claude, Gemini, Copilot.
🖼️ Image AI — how it works (simple backend picture)
When you ask for an image, this is roughly what happens:
Your text prompt
↓
App sends prompt to image model (on company's servers)
↓
Model has learned patterns from millions of images + captions
↓
It generates pixels that match your description
↓
Image file comes back to your screen
The art-school analogy
Imagine an artist who has seen millions of posters, photos, and ads — but never "understood" them like a human. You describe what you want. They paint something that looks like things they've seen before, mixed together.
That is text-to-image AI.
| You control | You cannot control (easily) |
|---|---|
| Style words ("cinematic", "minimal") | Exact pixel-perfect logo placement |
| Subject ("silver can, water drops") | Real brand trademarks sometimes |
| Mood ("luxury", "playful") | Consistent same face every time (without extra tools) |
Tools: Gemini (Nano Banana / image gen), ChatGPT + DALL·E, Adobe Firefly, Canva AI.
Use text AI to write the image prompt, then paste into image AI. Two-step chain = better posters.
🎬 Video AI — how it works (simple)
Video AI is harder than images because it must make many frames that move smoothly.
Your prompt (+ sometimes a starting image)
↓
Video model generates frame 1, frame 2, frame 3 …
↓
Frames stitched into a short clip (often 4–10 seconds)
↓
You download MP4
| Text AI | Image AI | Video AI |
|---|---|---|
| Fast | Medium speed | Slowest |
| Cheapest / often free tier | Some free limits | Usually limited free |
| Easy to fix (edit words) | Regenerate image | Regenerate whole clip |
Tools: Google Veo, Runway, Pika, CapCut AI, some features inside Canva/Gamma.
Expect weird hands, flickering, wrong text on screen. Use for drafts and ideas, not final TV ads — unless you review frame by frame.
🧰 Which tool for which job?
| Job | Start here |
|---|---|
| Write ad copy | ChatGPT / Gemini (text) |
| Product poster | Gemini image or Prompt Templates workflow |
| Instagram reel idea | Text AI writes script → Video AI makes clip |
| Presentation slides | Gamma.app (text → slides) |
| Remove background from photo | Canva / dedicated tools |
🎯 What the brochure means by "Multimodal AI Systems"
It does not mean "use every tool once."
It means: design workflows where text, image, and video steps connect:
Product photo → Text AI writes poster prompt → Image AI makes poster
↓
Text AI writes caption → You post on Instagram
That connected chain is a multimodal system — even if you run it by hand today.
Next: M5-B — AI Content Pipelines — how creators and businesses run these chains every day.
Practical: Prompt Templates — copy-paste poster and photoshoot prompts