---
title: "Your first image and video prompts | Optimize All Academy"
description: "Describing pictures, not intentions Many assistants can now create images from a description, and some can generate short videos. Image models respond to…"
url: https://optimizeall.com/learn/prompt-engineering-foundations/first-image-and-video-prompts
updated: 2026-10-05
---

Prompt Engineering Foundations · Modern assistant features: modes, memory, research and media · lesson 10 of 16 · 11 min

# Your first image and video prompts

## Describing pictures, not intentions

Many assistants can now create images from a description, and some can generate short videos. Image models respond to **visual description**. "Make a picture that increases sales" gives them nothing to draw. "A bright flat-lay photo of a glass of mango lassi on a white marble table, fresh mango slices beside it, soft morning light, lots of empty space on the right for text" does.

## The five-part image prompt

1. **Subject:** what is in the picture, with specifics (colours, materials, number of items).
2. **Setting:** where it is and what is around it.
3. **Style:** photo, illustration, watercolour, 3D render, flat icon, vintage poster.
4. **Light and mood:** soft daylight, golden hour, studio lighting, cosy, energetic.
5. **Layout and format:** close-up or wide, where the subject sits, space for text, and the shape (square for a feed post, tall for a story, wide for a banner).

```text
A cosy illustration of a small bookshop at night in a rainy London
street, warm yellow light glowing from the window, a cat asleep on a
stack of books in the window display. Style: soft watercolour with ink
outlines. Wide format, with empty sky at the top for a headline.
```

## Editing by conversation

The newest image tools let you refine an image by talking to it: "Make the cat ginger", "Move the text space to the left", "Same scene but in daylight". You can also upload a reference image, for example your own product photo, and ask for a new background. Change one thing at a time so you can see what each instruction does.

## Text inside images

Image models have become much better at writing text inside images, but mistakes still happen, especially with long phrases, small fonts or non-Latin scripts. For anything you will publish (prices, dates, Arabic or Urdu headlines), add the text afterwards in a design tool such as Canva, or check every letter carefully.

## Short videos

Video generators create short clips from a description or from a starting image. The same principles apply, plus two extras:

- **One clear action per shot:** "steam rises from a cup as a hand places a pastry beside it."
- **Camera words:** close-up, wide shot, slow push-in, slow pan, static camera.

```text
Close-up of a hand pouring chai from a steel kettle into a glass cup on
a wooden table, steam rising, early morning light through a window.
Slow push-in, calm and warm, realistic.
```

Longer stories are built by generating several short shots and editing them together. Video generation is slower and often costs more than images, so plan the shot before you generate.

## Rules you must follow

- **People and likeness:** never create realistic images or videos of real people without their consent, and never to mislead.
- **Honest marketing:** if you show a product, it must look like what customers will receive. Advertising rules in the UK, US, UAE, Saudi Arabia and Pakistan all prohibit misleading ads. Use real product photos for the product itself.
- **Labels and disclosure:** many platforms require or offer labels for realistic AI-generated content. Follow each platform's rules.
- **Copyright:** avoid asking for copyrighted characters or a named living artist's style. Check the tool's terms for commercial use.
- **Representation:** default images can reflect stereotypes. Describe people deliberately and review results.

## Worked example: a café launch post

A café in Islamabad launches a new saffron cake.

1. They photograph the real cake.
2. They ask an image tool: "Place this cake on a rustic wooden table with soft window light and dried rose petals around it; keep the cake exactly as it is." They check the cake has not been changed.
3. They generate three versions, pick one, and add the headline and price in Canva.
4. They post with the platform's AI label where required.

## Hands-on: your first three images

```text
1. Write a five-part prompt (subject, setting, style, light, layout) for
   an image you actually need.
2. Generate it, then make exactly two conversational edits, one change
   each time.
3. Write a second prompt for the same idea in a completely different
   style and compare which suits your audience better.
```

## How to judge an image

Score each result from 1 to 3 on four questions: Does it match the prompt? Does it fit the brand? Are there visual errors (hands, text, odd objects)? Is it honest about the product? Only publish images that score well on all four.

## Video lecture: Your first image and video prompts

Video: [Your first image and video prompts](https://www.youtube-nocookie.com/embed/_fwd1mmOXzA) — Describing pictures, not intentions Many assistants can now create images from a description, and some can generate short videos. Image models respond to…

11 chapters · about 8 minutes · captions and full transcript below.

1. Your first image and video prompts
2. Describe what the viewer sees
3. The five-part image prompt
4. Edit by conversation
5. Short video prompts
6. The rules
7. Worked example: a café launch
8. Judge every image
9. Example 1: a birthday card
10. Example 2: a smoothie launch (illustrative)
11. Recap

## Lecture transcript

### Your first image and video prompts

Make an image that boosts our sales. If you have ever typed something like that into an image generator, you know the result: something shiny and useless. In this lecture you will learn why image models need visual descriptions, a five-part structure for image prompts, how to edit images by conversation, how to prompt a short video clip, and the rules that keep you honest and legal. By the end, you will write prompts that produce images you can actually use.

### Describe what the viewer sees

Image models respond to what they can draw. Boost sales is an intention, and it gives the model nothing to render. A bright flat-lay photo of a glass of mango lassi on a white marble table, fresh mango slices beside it, soft morning light, and empty space on the right for text: that is a picture the model can make. So the golden rule is: describe what the viewer should see.

### The five-part image prompt

Use five parts. Subject: what is in the picture, with specific colours, materials and numbers. Setting: where it is and what surrounds it. Style: photo, illustration, watercolour, three D render, flat icon, vintage poster. Light and mood: soft daylight, golden hour, studio lighting, cosy or energetic. And layout and format: close-up or wide, where the subject sits, space for text, and the shape, square for a feed, tall for a story, or wide for a banner. For example: a cosy watercolour illustration of a small London bookshop at night in the rain, warm light from the window, a cat asleep on a stack of books, wide format with empty sky at the top for a headline.

### Edit by conversation

The newest image tools let you edit by conversation. Make the cat ginger. Move the text space to the left. Same scene, but in daylight. You can also upload your own photo, for example your real product, and ask for a new background while keeping the product exactly as it is. Change one thing at a time, so you can see what each instruction does. And be careful with text inside images. It has improved a lot, but long phrases, tiny fonts and non-Latin scripts like Arabic or Urdu can still come out wrong. For prices, dates and headlines you will publish, add the text yourself in a design tool.

### Short video prompts

Video generators create short clips from a description or a starting image. Two extra rules apply. One clear action per shot, such as steam rising from a cup as a hand places a pastry beside it. And camera words: close-up, wide shot, slow push-in, slow pan, static camera. Try this: close-up of a hand pouring chai from a steel kettle into a glass cup on a wooden table, steam rising, early morning light, slow push-in, calm and realistic. Longer stories are built by generating several short shots and editing them together. Video is slower and often costs more, so plan before you generate.

### The rules

Now the rules. Never create realistic images or videos of real people without their consent, and never to mislead. If you show a product, it must look like what customers will receive; advertising rules in the UK, US, UAE, Saudi Arabia and Pakistan all prohibit misleading ads. So use real photos of the product itself. Follow each platform's rules for labelling AI-generated content. Avoid copyrighted characters and named living artists' styles, and check the tool's terms for commercial use. And watch for stereotypes: describe people deliberately and review the results.

### Worked example: a café launch

Here is how a café in Islamabad launched a new saffron cake. They photographed the real cake. They asked an image tool to place it on a rustic wooden table with soft window light and dried rose petals, keeping the cake exactly as it is, and checked it had not changed. They generated three versions, picked one, and added the headline and price in a design tool. Then they posted with the platform's AI label where required. Beautiful, fast, and honest.

### Judge every image

Here is how to judge any image before you use it. Score it from one to three on four questions. Does it match the prompt? Does it fit the brand, in colours, mood and style? Are there visual errors, like odd hands, garbled text or strange objects? And is it honest about the product, showing what customers will actually receive? Only publish images that score well on all four. Keeping this little scorecard also teaches you which prompt details make the biggest difference for your brand.

### Example 1: a birthday card

A simple example first. You want a birthday card illustration for your mother, who loves her garden. Weak prompt: a nice birthday picture. Stronger prompt, using the five parts: a watercolour illustration of a small garden with pink roses and a wooden bench, a teacup on the bench, soft morning light, with a blank area at the top for a handwritten message, portrait format. Then one edit at a time: make the roses yellow; add a small bird on the bench. Each edit changes one thing, so you stay in control. Ten minutes, and the card feels personal.

### Example 2: a smoothie launch (illustrative)

Now a business scenario, with illustrative numbers. A juice bar in Jeddah is launching a mango and saffron smoothie and wants ten social images for the launch week. Instead of a full photo shoot, they photograph the real drink once, in good light. For each post, they generate a different setting around that real photo: a beach at sunset, a marble counter, a picnic blanket, always saying keep the drink exactly as it is. They add the price and Arabic headline in a design tool. They reject two images where the cup shape changed. If a shoot would have cost around two thousand riyals and a day of work, this approach took a few hours and a small subscription. The drink shown is the drink customers get. That is the line you never cross.

### Recap

To recap. Describe what the viewer should see, using subject, setting, style, light and layout. Edit one change at a time, and use your own photos as references. Add critical text yourself. For video, one action per shot, with camera words. And always keep people consented, products honest and content labelled where required. Try this now: write a five-part prompt for an image you actually need, make two single-change edits, and score the result on prompt match, brand fit, visual errors and honesty. Next: research with web search and sources.

## Key takeaways

- Image models need visual description: subject, setting, style, light and mood, layout and format.
- Refine images conversationally, one change at a time, and use your own product photos as references.
- Add important text in a design tool and check any text the model renders.
- Never depict real people without consent, keep product images honest and follow platform labelling rules.

## Try it

Write a five-part image prompt for something you actually need, generate it, make two single-change edits, and score the result on prompt match, brand fit, visual errors and honesty.

- [Previous: Prompting with files, photos, screenshots and voice](https://optimizeall.com/learn/prompt-engineering-foundations/prompting-with-files-photos-and-voice)
- [Next: Research with web search, deep research and sources](https://optimizeall.com/learn/prompt-engineering-foundations/research-with-web-search-and-sources)
- [All lessons of Prompt Engineering Foundations](https://optimizeall.com/learn/prompt-engineering-foundations)
