Multimodal & Reasoning Models in PracticeImage and video generation concepts · Lesson 10 of 17

Image editing, references and brand consistency

Article · 12 min · 8 min lecture

Video lecture

Image editing, references and brand consistency

11 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 11

Image editing and brand consistency

  • Four editing modes
  • References without product drift
  • The consistency kit
  • Provenance and disclosure

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

From one-off images to consistent campaigns

A single striking image is easy now. The hard part in real marketing work is consistency: the same product, character, palette and style across twenty assets, three aspect ratios and two languages, with honest product depiction and a clear record of what was generated.

Native multimodal image models (such as OpenAI's GPT Image models and Google's Gemini image models) have made this much more practical, because they can edit an existing image by instruction and accept multiple reference images.

Four editing modes

  1. Instruction edits: "Change the background to a sunset beach; keep the product, lighting direction and shadows."
  2. Masked edits (inpainting): edit only a selected region, useful for removing objects or replacing a label area.
  3. Reference-guided generation: "Use Image 1 for the product, Image 2 for the colour palette and Image 3 for the composition."
  4. Outpainting and reframing: extending the canvas to create 9:16 and 16:9 versions of a 1:1 image.

Hands-on: an edit with reference images

import base64, os
from openai import OpenAI

client = OpenAI()
result = client.images.edit(
    model=os.environ.get("IMAGE_MODEL", "gpt-image-1.5"),   # check current models and parameters
    image=[open("refs/product-bottle.png", "rb"),          # the real product photo
           open("refs/palette-moodboard.png", "rb")],      # style reference
    prompt=("Place the exact bottle from the first image on a pale sandstone ledge at golden hour, "
            "with soft shadows and a few dried desert flowers. Use the colour palette of the second "
            "image. Do not alter the bottle's shape, label, text or colour. Square format, empty space "
            "above for a headline, no added text."),
)
with open("campaign-01-square.png", "wb") as f:
    f.write(base64.b64decode(result.data[0].b64_json))

Then check the product against the original at full size: label text, cap shape, colour. If anything drifted, composite the real product photo over the generated scene in a design tool.

A consistency kit

For any series, create and reuse:

CONSISTENCY KIT: "Desert Bloom" autumn campaign
Product refs:     refs/product-bottle.png (front), refs/product-side.png
Style ref:        refs/palette-moodboard.png
Palette:          #E9D8C3 sand, #B5653A terracotta, #3F5A4A sage
Light:            golden hour, low sun from the left, soft shadows
Camera:           50mm look, eye level, shallow depth of field
Composition:      product on right third; headline space top-left
Never:            extra products, text in image, people's faces, logos
Formats:          1:1 master -> 4:5 feed, 9:16 story (outpaint, then check)

Put the kit in every prompt (or in a project or template), and log each output with its prompt, references and model.

Provenance and disclosure

  • Content Credentials (C2PA): many generators, including OpenAI's image models, attach C2PA metadata describing how an image was made. Some editing and social platforms read it. Preserve it where possible; stripping metadata can also remove provenance.
  • Invisible watermarks: Google embeds SynthID watermarks in content from its generative models, and other providers use their own methods.
  • Platform labels: Meta, TikTok and YouTube have labels or disclosure settings for realistic AI-generated or altered content; follow each platform's current rules.
  • Regulation: the EU AI Act includes transparency obligations for AI-generated and manipulated content, and advertising regulators in the UK, US, UAE and Saudi Arabia expect ads not to mislead. A generated "result" photo for a skincare product is a problem regardless of labels.

Worked example: a Lahore fashion label

A small fashion label shoots each new garment once, flat-lay and on a mannequin, in plain light. For the lookbook it generates settings (Lahore Fort at dusk, a modern café, a garden mehndi) using the real garment photos as references, then checks embroidery and colour against the originals. Two images where the embroidery pattern changed are rejected. Models shown are generated, not real people, and captions note AI-assisted imagery as the platforms require. The label saves most of a location shoot while keeping the garment honest.

Pitfalls

  • Letting the model "improve" the product (label text, shape, colour).
  • Generating realistic images of real people or implying real customer results.
  • Losing track of which prompt and references produced which asset.
  • Relying on in-image text for prices or legal lines.

How to measure success

Score each asset for product fidelity (pass/fail against the reference), brand consistency (1-3), artefacts (pass/fail) and time to approval. Track the rejection rate per model and per kit; a rising rate usually means the kit needs tightening or the model changed.

Key takeaways

  • Native multimodal image models support instruction edits, masked edits, multiple references and reframing.
  • Use the real product as a reference, forbid changes to it, and verify fidelity at full size.
  • A consistency kit (refs, palette, light, camera, composition, never-list, formats) keeps series coherent.
  • Preserve provenance (C2PA, watermarks), follow platform labels and never mislead about products or people.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A generated campaign image slightly changes the text on your product's label. What is the best response?
  2. What is the main purpose of a consistency kit?
  3. Which statement about C2PA Content Credentials is accurate?

Put it into practice

Create a consistency kit for one campaign, generate three assets with product and style references, check product fidelity at full size, and record prompt, references and model for each.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.