Lesson 7 of 8
Images, Audio and Video
Modern AI can work across multiple kinds of input and output. The important skill is choosing the modality that best fits the task.
Active reading
Understand the idea
Text is best for precise instructions and structured reasoning; images carry visual context; audio carries timing and voice; video adds motion over time. Multimodal systems combine these signals, but each added modality also adds uncertainty.
Example
For a product support task, use text for the question, an image for the damaged part, and ask for a concise checklist instead of requesting an open-ended answer.
Practice
Apply this lesson to one real task you already have, then write down how you would verify the result.
Common mistake
Adding every available modality even when it does not improve the decision.
Key takeaway
Choose the smallest set of modalities that gives the model the evidence it actually needs.