---
title: "Prompt anatomy: a structure that works"
description: "From wishes to directions A prompt like \"cool image of a coffee shop\" asks the model to guess everything. Professional prompts read like a creative…"
url: https://optimizeall.com/learn/ai-image-generation-and-design/prompt-anatomy
updated: 2026-10-05
---

AI Image Generation and Design · Prompting for images · lesson 4 of 18 · 7 min

# Prompt anatomy: a structure that works

## From wishes to directions

A prompt like "cool image of a coffee shop" asks the model to guess everything. Professional prompts read like a creative director's brief to a photographer or illustrator: specific about subject, style, composition, lighting and mood, and clear about constraints.

## The seven-part prompt structure

| Part | What to specify | Example |
|---|---|---|
| 1. Subject | Who/what, with concrete details | "a young woman in a mustard hijab reading on a tablet" |
| 2. Action / context | What's happening, where | "sitting by a window in a small modern café" |
| 3. Style / medium | Photo, illustration, 3D, painting, graphic | "editorial lifestyle photograph" |
| 4. Composition | Framing, angle, placement, negative space | "medium shot, subject on left third, empty space on right for text" |
| 5. Lighting | Quality, direction, time of day | "soft natural window light, late afternoon" |
| 6. Color / mood | Palette, atmosphere | "warm neutral tones, calm and focused mood" |
| 7. Technical / constraints | Camera, lens, aspect ratio, what to avoid | "shallow depth of field, 4:5 aspect ratio, no text, no logos" |

Put together:

```
Editorial lifestyle photograph of a young woman in a mustard hijab
reading on a tablet, sitting by a window in a small modern café.
Medium shot, subject on the left third, clean empty space on the right
for a headline. Soft natural window light, late afternoon. Warm neutral
tones, calm focused mood. Shallow depth of field. 4:5. No text, no logos.
```

Different tools interpret prompts differently. Some favor natural sentences; some respond to comma-separated keywords; some have dedicated parameters for aspect ratio, style strength or negative prompts. Learn your tool's syntax, but the underlying structure transfers.

## Vocabulary that helps

- **Shot types:** extreme close-up, close-up, medium shot, wide shot, overhead/flat lay, low angle, bird's-eye.
- **Lens and camera language:** wide-angle, 50mm, macro, shallow depth of field, motion blur, film grain.
- **Lighting:** soft box, rim light, backlit, golden hour, overcast, neon, high-key, low-key, chiaroscuro.
- **Styles and media:** flat vector illustration, watercolor, isometric 3D, claymation, risograph print, editorial photo, product packshot.
- **Mood words:** add them, but support them with concrete visual choices ("serene" + "soft light, muted palette, open space").

## Negative prompting and constraints

Some tools support negative prompts (what to exclude); others work better with positive phrasing. Common constraints:

- "No text, no watermark, no logos" — reduces unwanted artifacts.
- "Plain background" or "seamless studio backdrop" — for cut-outs.
- "Natural skin texture" — reduces over-smoothed, plastic-looking skin.
- Counts and positions: "exactly three cups, arranged left to right" (verify — counts can still fail).

## Prompting for design use

Remember the image is usually part of a layout. Prompt for:

- **Space for text:** "negative space in the top third".
- **Aspect ratio:** generate at the target ratio rather than cropping heavily.
- **Brand colors:** describe them in words ("deep teal and warm sand palette") — hex codes may or may not be understood depending on the tool; references work better.
- **Consistency hooks:** reuse the same style phrase across a campaign.

## Worked example: from vague to directed

- **V1:** "fitness app ad image" → generic gym with muscular model; no space for text; harsh lighting.
- **V2:** "Photo of a woman in her 40s doing a home workout on a yoga mat in a bright living room, early morning light, candid moment, smiling slightly, wide shot with space on the right for text, natural colors, 9:16, no logos" → closer; one hand is distorted.
- **V3:** V2 + "hands relaxed at her sides" + generate four variations → choose the best, fix minor details with inpainting.

## Prompt template

```
[Style/medium] of [subject with details], [action/context/setting].
[Shot type/angle], [composition incl. space for text].
[Lighting]. [Palette/mood]. [Technical details].
[Aspect ratio]. [Exclusions/constraints].
```

## Common mistakes

- Stacking vague adjectives ("beautiful, amazing, stunning, 8k, masterpiece") instead of concrete direction.
- Conflicting instructions ("minimalist" plus a long list of objects).
- Forgetting composition, so the image cannot hold a headline.
- Expecting exact counts, text or anatomy without checking.

## Prompting assistant models versus parameter-driven tools

In 2026 you will write prompts in two dialects:

- **Brief style (assistant-built models such as GPT Image and Gemini image models, Firefly Image 5):** full sentences, with the purpose stated ("This is a 4:5 Instagram ad; leave the top third empty for a headline"). These models also accept follow-up instructions ("keep everything, make the light warmer").
- **Parameter style (for example Midjourney):** a descriptive prompt plus parameters such as `--ar` for aspect ratio and `--no` for exclusions, and reference parameters or panels for style and subjects. Parameter names change between versions — check the current documentation.

The seven parts stay the same in both; only the syntax changes.

## Hands-on: four prompt templates with before/after results

Copy these, fill the brackets, and keep them in your prompt library.

```text
T1 PRODUCT SCENE (use with a reference photo of the real product)
Use the attached product photo exactly as it is: do not change its shape,
label or color. Place it [surface] in [setting]. [Shot type], [angle].
[Lighting: quality + direction]. Palette: [colors in words].
Leave [area] empty for text. [Aspect ratio]. No extra text or logos.

T2 LIFESTYLE PHOTO
Editorial photograph of [specific person description, activity],
in [specific location]. [Shot type], subject on [left/right third].
[Lighting]. [Mood] shown through [concrete visual choices].
Natural skin texture, realistic hands. [Aspect ratio]. No text.

T3 SPOT ILLUSTRATION (series)
Flat vector spot illustration of [subject] for a [topic] article.
Style: [style spec snippet]. Palette: [3-5 named colors].
Centered on a plain [color] background, generous margin. No text.

T4 BACKGROUND PLATE FOR A LAYOUT
Empty [setting] background for a [format] design, [aspect ratio].
Soft focus, low detail in [area] where text will sit, strongest detail
in [area]. [Lighting], [palette]. No people, no text, no logos.
```

| Template | Before (vague prompt) | After (template) |
|---|---|---|
| T1 | "perfume bottle luxury" — invented bottle, random gold text | Real bottle preserved, marble surface, clean top third for headline |
| T2 | "woman working happy" — stock-looking, centered, no space for copy | Specific person and setting, subject on left third, copy space on right |
| T3 | "money illustration" — mixed styles, cash clichés | Consistent vector style that matches the rest of the series |
| T4 | "nice office background" — busy detail behind where text goes | Calm left half for text, detail concentrated on the right |

These before/after descriptions are typical results rather than guaranteed ones; run both versions in your tool and keep the pair in your library as evidence.

## Prompt QA before you press generate

- Is the **purpose** stated (format, placement, text area)?
- Does every mood word have a **visual** partner (light, palette, space)?
- Are **counts and positions** stated only where they matter?
- Are **brands, artists and real people** absent unless you have rights and consent?

## Summary

Structure prompts around subject, context, style, composition, lighting, mood and constraints; use precise photographic and artistic vocabulary; prompt for how the image will be used in the layout; and check details every time.

## Video lecture: Prompt anatomy: a structure that works

Lecture coming soon · 10 chapters · about 8 minutes. Read the full transcript below.

1. Prompt anatomy
2. Why structure matters
3. The seven parts
4. Two dialects, same seven parts
5. Worked example 1: a fitness app ad
6. Worked example 2: Ramadan perfume scenes
7. Watch me do it: writing a prompt
8. Constraints that help
9. Common mistakes
10. Recap and try this now

## Lecture transcript

### Prompt anatomy

Picture two briefs landing on a photographer's desk. The first says, cool photo of a coffee shop. The second says, editorial lifestyle photo, a young woman in a mustard hijab reading on a tablet by the café window, late afternoon light, subject on the left third, space on the right for a headline. Which photographer delivers what the client wanted? Obviously the second. AI image models are the same. In this lecture, you'll learn a seven-part prompt structure that turns vague wishes into clear directions, how to write it for both conversational models and parameter-based tools, and four templates you can reuse tomorrow. By the end, your prompts will read like a creative director's brief.

### Why structure matters

Why is structure worth the effort? Because the model fills every gap with its defaults. Leave out lighting, and you get generic bright light. Leave out composition, and you get a centered subject with no room for your text. Leave out palette, and it picks whatever is most common. Every gap becomes a random decision. Think of ordering at a tailor. If you just say, a nice suit, you'll get a suit. It might even be nice. But it won't be your suit. The seven parts are your measurements. And there's a bonus. Structured prompts are easier to reuse, compare and hand to a teammate, which is how one good image becomes a consistent campaign.

### The seven parts

Here are the seven parts. One, subject, with concrete details. Two, action and context, what's happening and where. Three, style or medium, like editorial photo, flat vector illustration or 3D render. Four, composition, meaning shot type, angle, where the subject sits and where the empty space goes. Five, lighting, its quality, direction and time of day. Six, color and mood, with the mood backed by visual choices. And seven, technical details and constraints, like depth of field, aspect ratio, and what to avoid, such as no text and no logos. Here's the key idea. Mood words alone don't do much. Serene needs soft light, a muted palette and open space to actually look serene.

### Two dialects, same seven parts

Now, the two dialects. Assistant-built models, like OpenAI's GPT Image models, Google's Gemini image models and Adobe's Firefly Image 5, understand full sentences, so write like a brief and state the purpose. This is a four by five Instagram ad, keep the top third empty for a headline. Then refine with follow-up instructions. Keep everything, make the light warmer. Parameter-driven tools, like Midjourney, prefer a descriptive prompt plus parameters, such as an aspect ratio setting and a no setting for things to exclude, with separate controls for references. Parameter names change between versions, so check the current docs. The good news? The seven parts transfer perfectly. Only the grammar changes.

### Worked example 1: a fitness app ad

Worked example one. Let's fix a simple prompt. Before: fitness app ad image. What comes back? A generic gym, a very muscular model, harsh light, and no room for text. Now let's add the parts. Subject: a woman in her forties in comfortable workout clothes, checking a phone. Context: a small living room with a yoga mat. Style: warm editorial photo. Composition: medium shot, subject on the left third, empty wall on the right for the headline. Lighting: soft morning window light. Mood: calm and encouraging, with a warm neutral palette. Constraints: four by five, natural skin texture, no text, no logos. The after image looks like your audience, and it drops straight into the layout.

### Worked example 2: Ramadan perfume scenes

Worked example two, a business scenario with illustrative details. Omar runs social media for a perfume retailer in Riyadh. He needs twelve product scenes for Ramadan, all using the real bottles. Before: perfume bottle luxury. The model invents a bottle and scatters gold text everywhere. After, he uses template one, the product scene template. Use the attached product photo exactly as it is, do not change its shape, label or color. Place it on a dark marble surface beside dates and a brass lantern, three-quarter view, warm low light from the right, deep green and gold palette, top third empty for Arabic and English headlines, four by five, no extra text or logos. Then he swaps the surface and props for each scene, keeping everything else fixed.

### Watch me do it: writing a prompt

Watch me do it. I'll write a prompt from scratch using the checklist. I start with purpose, because it shapes everything. This is a background for a sixteen by nine webinar banner, with the title on the left. Next, subject and context: an empty modern co-working space. Then style: soft-focus editorial photograph. Composition: the left half calm and low in detail, the strongest detail on the right. Lighting: late afternoon, warm, from the right. Palette: navy and warm sand to match the brand. Constraints: no people, no text, no logos. Before I press generate, I run my QA. Purpose stated? Yes. Every mood word has a visual partner? Yes. Any brands, artists or real people? None. Now I generate a batch of four.

### Constraints that help

A quick word on negative prompts and constraints. Some tools support a separate field for things to exclude, while others respond better to positive phrasing, like plain seamless background instead of no clutter. Useful constraints include no text, no watermark, no logos, natural skin texture, and exact counts, such as exactly three cups arranged left to right. But counts can still fail, so always check. And when you need brand colors, describe them in words, like deep teal and warm sand. Some tools understand hex codes and many don't reliably, so reference images usually beat hex codes for color accuracy. You'll fix exact brand colors in your design tool anyway.

### Common mistakes

Common mistakes. Stacking adjectives like stunning, amazing, beautiful, which add almost nothing. Forgetting the layout, so there's no space for text. Generating at the wrong aspect ratio and cropping away half the image. Writing contradictions, like minimalist and highly detailed in the same prompt. Naming living artists or other brands, which creates ethical and legal risk. And changing six things at once, so you never learn what worked. Here's a habit that fixes most of these. Read your prompt out loud as if you were briefing a photographer. If a human would ask you a follow-up question, the model has to guess. Answer that question in the prompt.

### Recap and try this now

Let's recap. A professional prompt is a brief with seven parts: subject, action and context, style, composition, lighting, color and mood, and technical constraints. Write in full sentences with a stated purpose for assistant-built models, and add parameters for tools that use them. Back every mood word with a visual choice, design for the layout, and use references when accuracy matters. Your try this now: take the vaguest prompt you've used recently. Rewrite it with the seven parts, or with one of the four templates in the lesson. Generate both versions, drop them into your real layout, and save the before and after pair in your prompt library.

## Key takeaways

- Write prompts as briefs with seven parts: subject, action/context, style, composition, lighting, color/mood and technical constraints.
- Assistant-built models respond to full-sentence briefs with a stated purpose; parameter-driven tools add settings — the structure transfers.
- Back every mood word with concrete visual choices and design the image for its layout (copy space, aspect ratio).
- Use a reference photo and an explicit 'do not change' instruction when a real product must stay accurate.
- Keep before/after template pairs in a prompt library so results are reusable and teachable.

## Try it

Take a vague prompt you have used before and rewrite it with the seven-part structure and template. Generate both versions and compare how usable each is in a real layout.

- [Previous: Strengths, limits and when not to use AI imagery](https://optimizeall.com/learn/ai-image-generation-and-design/strengths-limits-and-when-not-to-use-ai)
- [Next: Styles, references and responsible style direction](https://optimizeall.com/learn/ai-image-generation-and-design/styles-and-references)
- [All lessons of AI Image Generation and Design](https://optimizeall.com/learn/ai-image-generation-and-design)
