Prompt Engineering FoundationsModern assistant features: modes, memory, research and media · Lesson 10 of 16

Your first image and video prompts

Article · 11 min · 8 min lecture

Video lecture

Your first image and video prompts

11 chapters · about 8 min · captions · full transcript

Watch now

Chapters

Describing pictures, not intentions

Many assistants can now create images from a description, and some can generate short videos. Image models respond to visual description. "Make a picture that increases sales" gives them nothing to draw. "A bright flat-lay photo of a glass of mango lassi on a white marble table, fresh mango slices beside it, soft morning light, lots of empty space on the right for text" does.

The five-part image prompt

  1. Subject: what is in the picture, with specifics (colours, materials, number of items).
  2. Setting: where it is and what is around it.
  3. Style: photo, illustration, watercolour, 3D render, flat icon, vintage poster.
  4. Light and mood: soft daylight, golden hour, studio lighting, cosy, energetic.
  5. Layout and format: close-up or wide, where the subject sits, space for text, and the shape (square for a feed post, tall for a story, wide for a banner).
A cosy illustration of a small bookshop at night in a rainy London
street, warm yellow light glowing from the window, a cat asleep on a
stack of books in the window display. Style: soft watercolour with ink
outlines. Wide format, with empty sky at the top for a headline.

Editing by conversation

The newest image tools let you refine an image by talking to it: "Make the cat ginger", "Move the text space to the left", "Same scene but in daylight". You can also upload a reference image, for example your own product photo, and ask for a new background. Change one thing at a time so you can see what each instruction does.

Text inside images

Image models have become much better at writing text inside images, but mistakes still happen, especially with long phrases, small fonts or non-Latin scripts. For anything you will publish (prices, dates, Arabic or Urdu headlines), add the text afterwards in a design tool such as Canva, or check every letter carefully.

Short videos

Video generators create short clips from a description or from a starting image. The same principles apply, plus two extras:

  • One clear action per shot: "steam rises from a cup as a hand places a pastry beside it."
  • Camera words: close-up, wide shot, slow push-in, slow pan, static camera.
Close-up of a hand pouring chai from a steel kettle into a glass cup on
a wooden table, steam rising, early morning light through a window.
Slow push-in, calm and warm, realistic.

Longer stories are built by generating several short shots and editing them together. Video generation is slower and often costs more than images, so plan the shot before you generate.

Rules you must follow

  • People and likeness: never create realistic images or videos of real people without their consent, and never to mislead.
  • Honest marketing: if you show a product, it must look like what customers will receive. Advertising rules in the UK, US, UAE, Saudi Arabia and Pakistan all prohibit misleading ads. Use real product photos for the product itself.
  • Labels and disclosure: many platforms require or offer labels for realistic AI-generated content. Follow each platform's rules.
  • Copyright: avoid asking for copyrighted characters or a named living artist's style. Check the tool's terms for commercial use.
  • Representation: default images can reflect stereotypes. Describe people deliberately and review results.

Worked example: a café launch post

A café in Islamabad launches a new saffron cake.

  1. They photograph the real cake.
  2. They ask an image tool: "Place this cake on a rustic wooden table with soft window light and dried rose petals around it; keep the cake exactly as it is." They check the cake has not been changed.
  3. They generate three versions, pick one, and add the headline and price in Canva.
  4. They post with the platform's AI label where required.

Hands-on: your first three images

1. Write a five-part prompt (subject, setting, style, light, layout) for
   an image you actually need.
2. Generate it, then make exactly two conversational edits, one change
   each time.
3. Write a second prompt for the same idea in a completely different
   style and compare which suits your audience better.

How to judge an image

Score each result from 1 to 3 on four questions: Does it match the prompt? Does it fit the brand? Are there visual errors (hands, text, odd objects)? Is it honest about the product? Only publish images that score well on all four.

Key takeaways

  • Image models need visual description: subject, setting, style, light and mood, layout and format.
  • Refine images conversationally, one change at a time, and use your own product photos as references.
  • Add important text in a design tool and check any text the model renders.
  • Never depict real people without consent, keep product images honest and follow platform labelling rules.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which prompt is most likely to produce a usable image?
  2. You need a promotional image with the exact price in Urdu. What is the safest approach?
  3. Which use of generated video is most appropriate for a product ad?

Put it into practice

Write a five-part image prompt for something you actually need, generate it, make two single-change edits, and score the result on prompt match, brand fit, visual errors and honesty.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.