Multimodal & Reasoning Models in Practice · Image and video generation concepts · lesson 8 of 17 · 11 min
Image generation: how it works and how to direct it
The core idea
Most leading image generators are based on diffusion or related techniques. Conceptually, a model learns to reverse a noising process: starting from random noise, it gradually refines an image step by step, guided by your text prompt (and optionally reference images). A text encoder turns your prompt into a representation that steers each refinement step. Some newer systems generate images natively within large multimodal models, which can improve text rendering and instruction following. The field changes fast; the practical skills below transfer across tools.
Why prompts for images differ
Image models respond to visual description, not instructions about intent. "Make an ad that increases conversions" means little to an image model. "Flat-lay product photo of a matte black water bottle on a light oak table, soft morning window light, minimal, space on the left for text" gives it something to render.
A structure for image prompts
- Subject: what is in the image, with specifics.
- Composition: framing, angle, placement, negative space for text.
- Style and medium: photo, illustration, 3D render, watercolour; era or aesthetic.
- Lighting and colour: soft daylight, studio, golden hour; brand palette.
- Details and constraints: aspect ratio, what to exclude (many tools support negative prompts or exclusions).
Subject: a young woman in a modest, light-blue abaya reviewing a laptop
spreadsheet in a bright co-working space in Riyadh
Composition: medium shot, subject on the right third, empty wall space
on the left for a headline
Style: natural, editorial photography, realistic
Lighting: soft daylight from large windows
Constraints: 16:9, no visible logos, no text in the image
Editing and control
Modern tools increasingly support:
- Image-to-image and reference images for consistent style or product appearance.
- Inpainting (edit a selected region) and outpainting (extend the canvas).
- Instruction-based editing ("change the background to a beach at sunset").
- Structural control (guiding pose, edges or depth from a reference) in some tools.
- Character and product consistency across a series, still one of the harder problems.
Known weaknesses
- Text in images has improved considerably in some systems but can still produce misspellings; add text in a design tool when accuracy matters.
- Hands, fine mechanical detail and counting can go wrong.
- Brand accuracy: logos and exact products are unreliable unless you use your own reference images and check outputs.
- Bias and stereotyping: default depictions of professions, nationalities or genders can reflect stereotypes. Specify diversity deliberately and review.
Legal and ethical guardrails
- Copyright and training data: the legal position of generated images and training data is contested and varies by jurisdiction. Check your tool's terms on commercial use and any indemnities, and avoid prompting for specific living artists' styles or copyrighted characters.
- Likeness: do not generate realistic images of real people without consent, and never in misleading or harmful contexts.
- Disclosure and provenance: some platforms and jurisdictions require labelling AI-generated or manipulated media, and many tools attach provenance metadata or watermarks. Content credential standards (such as C2PA) exist to help trace origin. Follow platform rules and advertising standards.
- Misleading ads: a generated "product photo" that exaggerates a product can breach consumer protection rules. Keep generated imagery honest about what customers will receive.
Worked example: campaign visuals
A small skincare brand needs lifestyle images for a campaign in the UK and Gulf markets. Workflow:
- Mood board and brand palette defined in text.
- Generate backgrounds and scenes; composite the real product photo in a design tool, so the product shown is accurate.
- Review for representation, cultural appropriateness and brand fit, with a local team member for each market.
- Add headlines in the design tool.
- Label as required by platform rules; keep records of prompts and tools used.
The current generation of image models
Two families dominate practical work:
- Diffusion and related models in dedicated image tools, strong on aesthetics and style control.
- Native multimodal image models built into large language model systems, such as OpenAI's GPT Image models and Google's Gemini image models. These follow long, detailed instructions well, render text far more reliably than earlier systems, and support conversational editing ("keep everything, but make the jacket navy") and multiple reference images.
The prompting consequence: with native multimodal models, write like a creative brief (full sentences, relationships between objects, the purpose of the image) rather than a list of comma-separated keywords.
Hands-on: generate an image through an API
import base64, os
from openai import OpenAI
client = OpenAI()
prompt = ("Editorial photo for a Riyadh co-working space's website hero: a woman in a light-blue "
"abaya reviewing a spreadsheet on a laptop, bright modern interior, soft daylight from "
"large windows. Subject on the right third; clean empty wall on the left for a headline. "
"Realistic, natural colours, no logos, no text in the image.")
result = client.images.generate(
model=os.environ.get("IMAGE_MODEL", "gpt-image-1.5"), # check current image models
prompt=prompt,
size="1536x1024",
)
with open("hero.png", "wb") as f:
f.write(base64.b64decode(result.data[0].b64_json))
Gemini's image models are called through the Gemini API with the image returned as inline data, and dedicated tools offer their own APIs. Supported sizes, quality settings and prices vary; check the docs before building.
A reusable brief template
PURPOSE: [where it will be used and what it must achieve]
SUBJECT: [who/what, specifics, actions, expressions]
SETTING: [place, time, surrounding details]
STYLE: [photo / illustration / 3D; references to eras or techniques, not living artists]
LIGHT & COLOUR: [lighting, palette, brand colours as hex if supported]
COMPOSITION: [framing, subject placement, negative space for text, aspect ratio]
MUST NOT: [logos, text, extra people, specific objects]
Store the filled brief with the output file name, model and settings. That record is your reproducibility and your audit trail.
Going further
For systematic work, keep a prompt library with seeds or reference images (where tools expose them) to reproduce looks. Evaluate generators on your own briefs with a simple rubric: prompt adherence, brand fit, artefacts, and time to acceptable result. Speed to a usable image often matters more than peak quality.
Video lecture: Image generation: how it works and how to direct it
Lecture coming soon · 11 chapters · about 7 minutes. Read the full transcript below.
- Image generation
- Two families
- Visual briefs
- Worked example + API
- The brief template
- Known weaknesses
- Guardrails
- Worked example: campaign visuals
- Example 1: a blog header image
- Example 2: off-plan property ads (illustrative)
- Recap
Lecture transcript
Image generation
Make an ad that increases conversions. Type that into an image generator and you get something glossy and useless. Describe what the viewer should see, and you get something you can use. In this lecture you will learn how image generators work, the two families of models you will meet in 2026, why image prompts need visual description, a reusable brief template, how to generate images through an API, the known weaknesses, and the legal and ethical guardrails.
Two families
Many image generators are based on diffusion. A model learns to reverse a noising process. Starting from random noise, it refines an image step by step, guided by your text prompt and optionally reference images. A text encoder turns your prompt into a representation that steers each step. The second family is native multimodal image models built into large language model systems, such as OpenAI's GPT Image models and Google's Gemini image models. These follow long, detailed instructions well, render text far more reliably than earlier systems, and support conversational editing and multiple reference images.
Visual briefs
The prompting consequence is important. Image models respond to visual description, not to intent. And with native multimodal models, write like a creative brief, in full sentences that describe relationships between objects and the purpose of the image, rather than a pile of comma-separated keywords. Use five core parts: the subject with specifics; composition, meaning framing, angle, placement and space for text; style and medium; lighting and colour; and constraints like aspect ratio and what to exclude.
Worked example + API
Here is an example from the lesson for a co-working space in Riyadh. An editorial photo for the website hero: a woman in a light-blue abaya reviewing a spreadsheet on a laptop in a bright modern interior, soft daylight from large windows. Subject on the right third, a clean empty wall on the left for a headline. Realistic, natural colours, no logos and no text in the image. The code sends that prompt to OpenAI's image generation endpoint with a wide size and saves the result. Gemini's image models work through the Gemini API, and dedicated tools have their own APIs. Sizes, quality settings and prices vary, so check the docs.
The brief template
For repeatable work, use a brief template with seven fields. Purpose: where it will be used and what it must achieve. Subject. Setting. Style, referring to eras or techniques rather than living artists. Light and colour, with brand colours as hex codes where supported. Composition, including aspect ratio and space for text. And must not: logos, text, extra people or specific objects. Store the filled brief with the output file name, model and settings. That record is your reproducibility and your audit trail.
Known weaknesses
Know the weaknesses. Text in images has improved a lot but can still misspell, especially long phrases and non-Latin scripts, so add critical text in a design tool. Hands, fine mechanical detail and counting can go wrong. Exact brand accuracy, logos and real products, is unreliable unless you use your own reference images and check. And default depictions of professions, nationalities or genders can reflect stereotypes, so specify people deliberately and review.
Guardrails
Now the guardrails. Copyright and training data remain contested and vary by jurisdiction, so check your tool's terms on commercial use and indemnities, and avoid prompting for specific living artists' styles or copyrighted characters. Do not generate realistic images of real people without consent. Follow platform and legal disclosure rules; many tools attach provenance metadata or watermarks, and content credential standards such as C2PA help trace origin. And do not mislead: a generated product photo that exaggerates the product can breach consumer protection and advertising rules.
Worked example: campaign visuals
Here is a campaign workflow for a small skincare brand targeting the UK and the Gulf. Define the mood board and brand palette in text. Generate backgrounds and scenes, then composite the real product photo in a design tool, so the product shown is exactly what customers get. Review for representation, cultural appropriateness and brand fit, with a local team member for each market. Add headlines in the design tool. Label as platforms require, and keep records of prompts and tools. To evaluate, score each image for prompt adherence, brand fit and artefacts, and track time to an acceptable result.
Example 1: a blog header image
A simple worked example. You want a header image for a blog post about working from home. Keyword prompt: home office, laptop, cosy, plants, 4k. The result is generic stock. Brief prompt: a small, sunlit home office in a Karachi apartment in the late afternoon, a laptop on a wooden desk with a cup of chai beside it, a potted money plant, warm light through a patterned curtain, wide format with empty space on the right for the blog title, realistic photo style, no text. The second image feels specific and local, and the empty space means the title fits without cropping. The difference is purpose, place and layout, not magic words.
Example 2: off-plan property ads (illustrative)
Now a business scenario, with illustrative numbers. A real estate marketing agency in Dubai produces about one hundred and twenty social ads a month for off-plan projects. Photography is not possible, because the buildings do not exist yet. They use a strict workflow. Architectural renders supplied by the developer are the only depiction of the building. Image generation is used only for lifestyle scenes, a family on a balcony at sunset, a café by the water, with no claims about views or amenities that the developer has not confirmed. Every brief is logged with the output. Each image is checked against the developer's approved amenity list. And ads follow the platform rules and local advertising standards. Production time per ad falls by about half, and not one ad depicts a feature the project will not have. Illustrative figures.
Recap
To recap. Diffusion and native multimodal models both turn visual descriptions into images; native models follow detailed briefs and edit conversationally. Write full-sentence briefs covering purpose, subject, setting, style, light, composition and exclusions, and keep records. Expect weaknesses in text, detail and brand accuracy, and respect copyright, consent, disclosure and consumer protection. Try this now: write a brief for one real marketing need, generate three variants, and score them on adherence, brand fit and artefacts. Next: editing, references and brand consistency.
Key takeaways
- Diffusion-style models refine noise into images guided by your prompt; native multimodal generation is also emerging.
- Image prompts need visual description: subject, composition, style, lighting and constraints.
- Weak points include text rendering, fine detail, exact brand accuracy and stereotyped defaults; composite real products for accuracy.
- Respect copyright uncertainty, likeness consent, disclosure rules and consumer protection.
Try it
Write a structured image prompt (subject, composition, style, lighting, constraints) for one real marketing need. Generate three variants and score them against a simple rubric.