Multimodal & Reasoning Models in PracticeImage and video generation concepts · Lesson 10 of 17
Image editing, references and brand consistency
Video lecture
Image editing, references and brand consistency
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Image editing and brand consistency
One beautiful AI image is easy now. Twenty images for a campaign, all with the same product, the same palette and the same lighting, across three aspect ratios and two languages, with the product shown honestly: that is the real job. In this lecture you will learn the four editing modes of modern image models, how to use reference images without letting the model alter your product, how to build a consistency kit, and how provenance, platform labels and regulation fit in.
0:36 Four editing modes
Native multimodal image models make consistency far more practical, because they can edit an existing image by instruction and accept multiple reference images. There are four editing modes. Instruction edits: change the background to a sunset beach, keeping the product, lighting direction and shadows. Masked edits, or inpainting, which change only a selected region. Reference-guided generation: use image one for the product, image two for the palette, image three for the composition. And outpainting or reframing, extending a square image into nine by sixteen or sixteen by nine versions.
1:15 Hands-on: edit with references
In the lesson's hands-on code, an edit request sends two images to OpenAI's image edit endpoint: the real product photo of a bottle, and a palette mood board. The prompt places the exact bottle from the first image on a pale sandstone ledge at golden hour, with soft shadows and dried desert flowers, uses the palette of the second image, and explicitly says do not alter the bottle's shape, label, text or colour. Square format, headline space above, no added text. Then comes the essential step: compare the product in the output against the original at full size. If anything drifted, composite the real photo over the generated scene.
2:02 The consistency kit
For any series, build a consistency kit. Product references, front and side. A style reference. A palette as hex codes, for example sand, terracotta and sage. The light, golden hour with a low sun from the left. The camera look, like a fifty millimetre look at eye level with shallow depth of field. Composition rules, product on the right third, headline space top-left. A never list: extra products, text in the image, people's faces, logos. And formats: a square master, then four by five for feeds and nine by sixteen for stories, each checked after reframing. Put the kit in every prompt or template, and log every output with its prompt, references and model.
2:52 Provenance and labels
Provenance is becoming part of the workflow. Content Credentials, based on the C2PA standard, are metadata describing how an image was made; many generators, including OpenAI's image models, attach them, and some platforms read them. Stripping metadata can remove them, so preserve it where you can. Google embeds SynthID invisible watermarks in content from its generative models, and other providers use their own methods. Social platforms such as Meta, TikTok and YouTube have labels or disclosure settings for realistic AI-generated content. And the EU AI Act includes transparency obligations for AI-generated and manipulated content.
3:33 Honesty first
None of that replaces honesty. Advertising regulators in the UK, the US, the UAE and Saudi Arabia expect ads not to mislead. A generated result photo for a skincare product, showing effects the product does not deliver, is a problem no matter how clearly it is labelled. Keep generated imagery to settings, moods and concepts, show the real product accurately, and never imply real customer results or endorsements with generated people.
4:04 Worked example: a Lahore lookbook
A worked example from Lahore. A small fashion label shoots each garment once, flat-lay and on a mannequin, in plain light. For the lookbook, it generates settings, the Lahore Fort at dusk, a modern café and a garden mehndi, using the real garment photos as references. Then it checks the embroidery and colour against the originals. Two images where the embroidery pattern changed are rejected. Models shown are generated, not real people, and captions note AI-assisted imagery as the platforms require. The label saves most of a location shoot, while keeping every garment honest.
4:45 Example 1: one candle, three scenes
A simple worked example. You have one good photo of your homemade candle and want three versions: a cosy winter scene, a spring garden scene and a minimal white background for your shop page. Use the photo as the reference each time and say: keep the candle, jar and label exactly as they are; change only the setting. Then compare each output to the original at full size. The spring version changed the label's font slightly. You reject it, run it again, and the second attempt keeps the label intact. Three consistent images from one photo, in twenty minutes.
5:28 Example 2: a 10-shade lipstick range (illustrative)
Now a business scenario, with illustrative numbers. A cosmetics brand in Lahore launches a ten-shade lipstick range and needs about ninety assets: each shade in three settings, in square, portrait and story formats. They build a consistency kit: product references for every shade, the palette as hex codes, golden-hour light from the left, product on the right third, and a never list that bans faces and text. They generate square masters first, check shade accuracy against the physical swatch photos, then reframe to other formats and check again. About twelve percent of first attempts fail the shade check, usually a colour drift, and are regenerated or composited. Everything is logged with prompts, references and model. A shoot for ninety assets would have taken several days; this takes about two days of generation and checking. Illustrative figures, but the shade check is non-negotiable for cosmetics.
6:31 Pitfalls and scorecard
Four pitfalls to watch: letting the model improve the product's label, shape or colour; generating realistic images of real people or implying real results; losing track of which prompt and references made which asset; and relying on in-image text for prices or legal lines. Measure each asset with a simple scorecard: product fidelity, pass or fail against the reference; brand consistency on a one to three scale; artefacts, pass or fail; and time to approval. Track the rejection rate per model and per kit. A rising rate usually means the kit needs tightening, or the model changed.
7:13 Recap
To recap. Use instruction edits, masked edits, references and reframing to build consistent series. Always feed the real product as a reference, forbid changes to it, and verify at full size, compositing when needed. Build a consistency kit and log everything. Preserve provenance, follow platform labels and regulation, and never mislead. Try this now: create a consistency kit for one campaign, generate three assets with product and style references, check product fidelity at full size, and record the prompt, references and model for each. Next: video generation.
From one-off images to consistent campaigns
A single striking image is easy now. The hard part in real marketing work is consistency: the same product, character, palette and style across twenty assets, three aspect ratios and two languages, with honest product depiction and a clear record of what was generated.
Native multimodal image models (such as OpenAI's GPT Image models and Google's Gemini image models) have made this much more practical, because they can edit an existing image by instruction and accept multiple reference images.
Four editing modes
- Instruction edits: "Change the background to a sunset beach; keep the product, lighting direction and shadows."
- Masked edits (inpainting): edit only a selected region, useful for removing objects or replacing a label area.
- Reference-guided generation: "Use Image 1 for the product, Image 2 for the colour palette and Image 3 for the composition."
- Outpainting and reframing: extending the canvas to create 9:16 and 16:9 versions of a 1:1 image.
Hands-on: an edit with reference images
import base64, os
from openai import OpenAI
client = OpenAI()
result = client.images.edit(
model=os.environ.get("IMAGE_MODEL", "gpt-image-1.5"), # check current models and parameters
image=[open("refs/product-bottle.png", "rb"), # the real product photo
open("refs/palette-moodboard.png", "rb")], # style reference
prompt=("Place the exact bottle from the first image on a pale sandstone ledge at golden hour, "
"with soft shadows and a few dried desert flowers. Use the colour palette of the second "
"image. Do not alter the bottle's shape, label, text or colour. Square format, empty space "
"above for a headline, no added text."),
)
with open("campaign-01-square.png", "wb") as f:
f.write(base64.b64decode(result.data[0].b64_json))Then check the product against the original at full size: label text, cap shape, colour. If anything drifted, composite the real product photo over the generated scene in a design tool.
A consistency kit
For any series, create and reuse:
CONSISTENCY KIT: "Desert Bloom" autumn campaign
Product refs: refs/product-bottle.png (front), refs/product-side.png
Style ref: refs/palette-moodboard.png
Palette: #E9D8C3 sand, #B5653A terracotta, #3F5A4A sage
Light: golden hour, low sun from the left, soft shadows
Camera: 50mm look, eye level, shallow depth of field
Composition: product on right third; headline space top-left
Never: extra products, text in image, people's faces, logos
Formats: 1:1 master -> 4:5 feed, 9:16 story (outpaint, then check)Put the kit in every prompt (or in a project or template), and log each output with its prompt, references and model.
Provenance and disclosure
- Content Credentials (C2PA): many generators, including OpenAI's image models, attach C2PA metadata describing how an image was made. Some editing and social platforms read it. Preserve it where possible; stripping metadata can also remove provenance.
- Invisible watermarks: Google embeds SynthID watermarks in content from its generative models, and other providers use their own methods.
- Platform labels: Meta, TikTok and YouTube have labels or disclosure settings for realistic AI-generated or altered content; follow each platform's current rules.
- Regulation: the EU AI Act includes transparency obligations for AI-generated and manipulated content, and advertising regulators in the UK, US, UAE and Saudi Arabia expect ads not to mislead. A generated "result" photo for a skincare product is a problem regardless of labels.
Worked example: a Lahore fashion label
A small fashion label shoots each new garment once, flat-lay and on a mannequin, in plain light. For the lookbook it generates settings (Lahore Fort at dusk, a modern café, a garden mehndi) using the real garment photos as references, then checks embroidery and colour against the originals. Two images where the embroidery pattern changed are rejected. Models shown are generated, not real people, and captions note AI-assisted imagery as the platforms require. The label saves most of a location shoot while keeping the garment honest.
Pitfalls
- Letting the model "improve" the product (label text, shape, colour).
- Generating realistic images of real people or implying real customer results.
- Losing track of which prompt and references produced which asset.
- Relying on in-image text for prices or legal lines.
How to measure success
Score each asset for product fidelity (pass/fail against the reference), brand consistency (1-3), artefacts (pass/fail) and time to approval. Track the rejection rate per model and per kit; a rising rate usually means the kit needs tightening or the model changed.
Key takeaways
- Native multimodal image models support instruction edits, masked edits, multiple references and reframing.
- Use the real product as a reference, forbid changes to it, and verify fidelity at full size.
- A consistency kit (refs, palette, light, camera, composition, never-list, formats) keeps series coherent.
- Preserve provenance (C2PA, watermarks), follow platform labels and never mislead about products or people.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Create a consistency kit for one campaign, generate three assets with product and style references, check product fidelity at full size, and record prompt, references and model for each.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.