---
title: "Image generation and vision — Mastering ChatGPT (OpenAI)"
description: "Two different capabilities - Vision means ChatGPT analyses images you give it: screenshots, photos, charts, handwriting, product labels, design drafts. -…"
url: https://optimizeall.com/learn/mastering-chatgpt/image-generation-and-vision
updated: 2026-10-05
---

Mastering ChatGPT (OpenAI) · Files, data, images and voice · lesson 7 of 19 · 15 min

# Image generation and vision

## Two different capabilities

- **Vision** means ChatGPT *analyses* images you give it: screenshots, photos, charts, handwriting, product labels, design drafts.
- **Image generation** means ChatGPT *creates* new images from descriptions (and edits existing ones). ChatGPT's image generation is powered by OpenAI's GPT Image models. In September 2026 OpenAI released **ChatGPT Images 2.5**, which added a **Sketch** feature for turning your own drawings into finished images and reduced generation time. The older standalone DALL·E GPT was retired at the end of August 2026; image creation now lives directly in ChatGPT and its images library.

Availability, limits and quality settings vary by plan.

## Vision: the describe-first pattern

```text
1. Describe everything you see in this image: text, numbers, labels, colours, layout.
   Don't interpret yet.
2. Using only that description, answer: [your question].
3. List anything you were unsure about reading.
```

This separates perception from reasoning, so a misread number is caught before it becomes a conclusion. Crop to what matters, use legible resolution, and label multiple images ("Image 1 = before, Image 2 = after").

## Image generation: the prompt formula

A strong image prompt covers:

| Element | Example |
|---|---|
| Subject | a ceramic coffee cup with Arabic calligraphy |
| Setting | on a marble café counter in Old Dubai at golden hour |
| Style | editorial product photography, natural light |
| Composition | cup on the left third, clean space at the top for headline text |
| Lighting and mood | warm, soft shadows, inviting |
| Format | 4:5 vertical for Instagram feed |
| Text (if any) | exact words in quotes, and check spelling in the result |

Iterate **one change at a time** ("same image, but cooler morning light") and keep a prompt log so you can reproduce what worked. Use Sketch when you know the composition you want: draw rough shapes and let the model render them.

## Editing images

You can upload an image and ask for edits (change the background, remove an object, adjust the style), or select part of a generated image and describe the change. Always compare the edit to the original for unintended changes, especially to faces, logos and product details.

## Responsible use

- **Do not misrepresent real products.** Use AI images for concepts, mood boards and backgrounds; use real photos of the actual product in ads, or composite real product shots into AI backgrounds.
- **People and likeness:** do not create images that impersonate real people or imply endorsements they did not give. Respect OpenAI's usage policies.
- **Disclosure and provenance:** OpenAI adds C2PA provenance metadata to ChatGPT-generated images; platforms such as Meta, TikTok and YouTube have their own AI-content labelling rules. Follow them, and your advertising codes.
- **Intellectual property:** avoid prompts that copy protected characters, logos or a living artist's distinctive work; check your organisation's policy for commercial use.
- **Check details:** hands, text, logos, reflections and product features often contain errors.

## Worked example: a restaurant launch

A Karachi restaurant launching a new biryani platter uses ChatGPT to generate three mood-board images for the menu design and social teaser (clearly concept art), and Sketch to test a flat-lay layout. For the launch ads, the team photographs the real dish in the layout ChatGPT helped design, and uses ChatGPT's vision to review the ad drafts for text legibility on mobile. No customer ever sees an AI image presented as the real food.

## Hands-on

1. Use the prompt formula to create one concept image for your brand or project.
2. Refine it three times, changing one element each time; log what changed.
3. Upload a screenshot of an ad or landing page (with personal data cropped) and run the describe-first pattern.

## A reusable creative-review prompt (vision)

```text
You are reviewing social ad creatives for [brand]. For each attached image:
1. Describe layout, headline text, CTA, colours and any small print.
2. Score 1-5: offer clarity at thumbnail size, brand fit, visual hierarchy,
   mobile legibility.
3. Flag any claim that needs evidence and any missing ad disclosure label.
4. Suggest one concrete fix.
End with a ranking table and "Uncertain readings".
```

Treat scores as hypotheses to test with real audiences, not results.

## Pitfalls

- Changing five things per iteration and losing what worked.
- Publishing AI images with garbled text or extra fingers.
- Uploading screenshots with customer emails or account details visible.

## How to measure success

Concept rounds that used to take days take hours, no AI image misrepresents a real product or person, and vision reviews catch issues before publication.

## Video lecture: Image generation and vision

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Images in ChatGPT
2. Why images need a method
3. Two capabilities
4. Vision: describe first
5. The image prompt formula
6. Iterate one change at a time
7. Editing images
8. Responsible use
9. Simple example
10. Worked example: a Karachi restaurant
11. A creative-review prompt
12. Try this now
13. Watch me do it, part 1
14. Watch me do it, part 2
15. Recap and next step

## Lecture transcript

### Images in ChatGPT

ChatGPT can look at images and it can make them. Used well, that means faster concept work, sharper creative reviews and fewer design bottlenecks. Used carelessly, it means misleading ads and embarrassing typos. In this lecture you will learn both capabilities, a prompt formula that works, what changed in September twenty twenty six, and the rules that keep you honest.

### Why images need a method

Why do images need a method? Because a beautiful image is not the same as an accurate one. AI can produce a gorgeous product shot where the label text is garbled or the product shape is subtly wrong. For concept work that is fine. For an ad, it can mislead customers and break advertising rules. And on the vision side, a quick glance at a screenshot can misread a number. The methods in this lecture keep your images both beautiful and honest.

### Two capabilities

Vision means ChatGPT analyses images you give it, screenshots, photos, charts, handwriting and design drafts. Image generation means it creates new images from descriptions and edits existing ones, powered by OpenAI's GPT Image models. In September twenty twenty six OpenAI released ChatGPT Images two point five, which added Sketch, so you can turn your own rough drawing into a finished image, and made generation faster. The old standalone DALL E GPT was retired at the end of August, and image creation now lives directly in ChatGPT.

### Vision: describe first

For vision, use the describe first pattern. Ask ChatGPT to describe everything it sees, text, numbers, labels, colours and layout, without interpreting. Then answer your question using only that description. Then list anything it was unsure about reading. This separates seeing from thinking, so a misread number is caught before it becomes a conclusion. Crop to what matters, and label multiple images clearly.

### The image prompt formula

For generation, use a formula. Subject, a ceramic coffee cup with Arabic calligraphy. Setting, a marble café counter in Old Dubai at golden hour. Style, editorial product photography. Composition, the cup on the left third with clean space at the top for a headline. Lighting and mood, warm and inviting. Format, four by five vertical for the Instagram feed. And if you need text in the image, put the exact words in quotes and check the spelling in the result.

### Iterate one change at a time

Then iterate one change at a time. Same image, but cooler morning light. Same image, cup moved right. Keep a prompt log so you can reproduce what worked. And if you already know the composition you want, sketch rough shapes and let Sketch render them. It is often faster than describing a layout in words.

### Editing images

You can also edit. Upload an image and ask to change the background or remove an object, or select part of a generated image and describe the change. Always compare the edit with the original, because faces, logos and product details can change in ways you did not ask for.

### Responsible use

Now the rules. Do not misrepresent real products. Use AI for concepts, mood boards and backgrounds, and real photos of the actual product in ads, or composite real shots into AI backgrounds. Do not create images that impersonate real people or imply endorsements. OpenAI adds C two P A provenance metadata to generated images, and social platforms have their own AI labelling rules, so follow them and your advertising codes. Avoid copying protected characters, logos or a living artist's distinctive style. And check the details, hands, text, logos and reflections.

### Simple example

A simple example. You need a header image for a blog post about working from cafés. Using the formula, you ask for a laptop and a cup of coffee on a wooden café table, soft morning light, editorial photography style, subject on the right third with space on the left for a headline, wide sixteen by nine format. The first version is close but a little cold. You change one thing, warmer golden hour light, and log the change. Two iterations, and you have a header you can use.

### Worked example: a Karachi restaurant

A restaurant in Karachi is launching a new biryani platter. The team uses ChatGPT to generate three mood board images for the menu design and social teaser, clearly as concept art, and uses Sketch to test a flat lay layout. For the launch ads, they photograph the real dish in the layout ChatGPT helped design, then use vision to check the ad drafts for text legibility on mobile. No customer ever sees an AI image presented as the real food. The team also used vision on the draft ads. ChatGPT flagged that the price on the mobile story ad sat too close to the edge and would be cut off by the platform's interface, and that the Urdu headline was too small at thumbnail size. Both were fixed before launch. Concept images helped them design faster, and vision helped them publish cleaner, but every customer facing image showed the real dish.

### A creative-review prompt

Here is a vision workflow worth saving. Ask ChatGPT to review your ad creatives one by one. Describe the layout, headline, call to action, colours and small print. Score offer clarity at thumbnail size, brand fit, hierarchy and mobile legibility. Flag any claim that needs evidence and any missing ad disclosure label. Suggest one concrete fix per creative, and end with a ranking table. Then treat the scores as hypotheses, and let real audience tests decide.

### Try this now

Try this now. Create one concept image for your brand or a current project using the seven part formula. Iterate three times, changing only one element each time, and log each change. Notice which change made the biggest difference. Then take a screenshot of one of your ads or landing pages, crop out any personal data, and run the describe first review. Did the description catch anything you had not noticed, like small text that becomes unreadable on mobile? Save both prompts for next time.

### Watch me do it, part 1

Let me create the café concept image. I type the prompt using the formula. Subject, a ceramic coffee cup with Arabic calligraphy. Setting, a marble counter in an Old Dubai café at golden hour. Style, editorial product photography. Composition, cup on the left third, clean space at the top for a headline. Lighting, warm with soft shadows. Format, four by five vertical. And the text, the words Morning Ritual in quotes. The first render is lovely, but when I zoom in, the headline reads Morning Ritaul. Text in images always needs checking.

### Watch me do it, part 2

I iterate one change at a time. Same image, fix the headline spelling to Morning Ritual. Version two is correct. Same image, cooler morning light. Version three is fresher, and the team prefers it. Then I try Sketch. I draw a rough flat lay, a cup in the centre, a date bowl top right and a notebook bottom left, and upload it. The result follows my layout closely. I log every version in a small table, the version, the one change and the result, so anyone can reproduce the approved look. And because this is concept art for a real café, the final ad will use a real photo shot in this layout.

### Recap and next step

Recap. Vision analyses, generation creates. Describe before you conclude. Use the seven part prompt formula, iterate one change at a time, and use Sketch for layouts. And never let AI images misrepresent real products or people. Your next step: create one concept image for your brand, refine it three times while logging each change, and run a describe first review on one of your ads.

## Key takeaways

- Vision analyses images you upload; image generation (GPT Image, Images 2.5 with Sketch since September 2026) creates and edits images.
- Describe subject, setting, style, composition, lighting, mood, format and exact text; iterate one change at a time with a prompt log.
- Don't use AI images to misrepresent real products or people; follow provenance, platform labelling and advertising rules.
- Check hands, text, logos and product details, and crop personal data from screenshots before uploading.

## Try it

Use the image prompt formula to create one concept image for your brand or project. Refine it three times and note which change had the biggest effect.

- [Previous: File uploads and data analysis](https://optimizeall.com/learn/mastering-chatgpt/file-uploads-and-data-analysis)
- [Next: Voice: talk, think and practise out loud](https://optimizeall.com/learn/mastering-chatgpt/voice-conversations)
- [All lessons of Mastering ChatGPT (OpenAI)](https://optimizeall.com/learn/mastering-chatgpt)
