A Practical Guide to AI Image Prompting: Composition, Style and Iteration
How to describe an image so a generative model produces what you actually had in mind, and how to steer it there over a handful of revisions.
Most disappointing AI images are not failures of the model. They are failures of specification. The model filled in everything you left unsaid, and its guesses were not your intentions. Prompting well is mostly the skill of noticing what you have left unsaid — and then deciding which of those gaps you actually care about.
Older advice treated prompts as keyword incantations: long strings of style tags and quality boosters. Modern models understand ordinary descriptive language, so that habit has largely outlived its usefulness. What follows is an approach built on describing a scene the way you would to a competent illustrator who cannot ask you questions.
Start with the four things every prompt needs
Before styling, get the substance right. A prompt that answers these four questions will already outperform most attempts.
- Subject. What or who is in the image, and in what state? "A woman" is thin. "A woman in her sixties in a canvas apron, mid-laugh, hands dusted with flour" is a picture.
- Setting. Where, and when? Environment carries an enormous amount of implicit information about light, colour and mood.
- Composition. How is it framed? Shot distance, camera angle, what occupies the foreground and background, where negative space sits.
- Light. The single most underused lever. Soft overcast light, hard midday sun, a single warm lamp in a dark room, backlit at dusk — each produces a completely different image from an identical subject.
Style comes after these, not instead of them. A style label attached to a vague subject just gives you a vaguely-styled vague image.
Describing composition in words the model understands
Composition is where people most often assume the model can read their mind. It cannot, but it does respond well to standard photographic and illustrative vocabulary.
Framing and distance
Terms like extreme close-up, close-up, medium shot, full-body shot and wide establishing shot are well understood and immediately change the result. If you want the subject small within a large landscape, say so explicitly — models default to making the subject prominent.
Camera position
Eye level reads as neutral. A low angle looking up lends stature; a high angle looking down diminishes or reveals layout; an overhead flat-lay suits objects and food. Say where the camera is, not just what it sees.
Arrangement
State the spatial relationships you care about: "the lamp on the left edge, the chair in the foreground turned away from it." Note that spatial precision is genuinely hard for these models. Multiple relationships, exact counts and left/right placement often need several attempts, or a structural tool like a rough sketch used as a reference, rather than words alone.
Depth
Shallow depth of field with a blurred background, or everything in crisp focus front to back — this shapes how the eye moves through the frame and is easy to specify.
Style without cliché
The lazy move is to name a living artist. It is also the move most likely to produce work you cannot safely use commercially, and many services now restrict it. It is unnecessary anyway: styles decompose into describable properties.
- Medium: oil on canvas, ink and wash, gouache, 35mm film photograph, 3D render, cut-paper collage.
- Palette: muted earth tones, high-contrast black and white, dominated by teal with a single warm accent.
- Line and texture: loose visible brushwork, clean vector edges, heavy film grain, flat matte surfaces.
- Era or tradition: mid-century travel poster, 1970s documentary photography, woodblock print, technical botanical illustration.
- Rendering and mood: soft and diffuse, stark and graphic, warm and nostalgic.
Combining three or four such properties gives a more distinctive and more controllable result than any artist name, and it is yours rather than borrowed.
Iteration is the actual workflow
Treat your first output as a draft that tells you what the model assumed. The productive loop is to change one thing at a time and observe.
Diagnose before you rewrite
When an image is wrong, identify the category of wrongness. Wrong subject or missing element means your description was ambiguous or overloaded. Right content but wrong feeling usually means light and palette, not subject. Right feeling but wrong framing means composition terms. Rewriting the whole prompt from scratch discards the information you just gained.
Change one variable at a time
If you alter the subject, the lighting and the style simultaneously and the result improves, you have learned nothing transferable. Single changes build an actual mental model of how a given tool responds.
Use seeds and variations
Most tools let you fix a random seed, so the same prompt reproduces the same image. Holding the seed steady while editing the prompt isolates the effect of your words. Releasing it gives you fresh compositions from the same description. Variation features that generate near-neighbours of an image you like are usually a faster route to a good result than more prompt writing.
Edit rather than regenerate
When an image is eighty percent right, stop prompting. Inpainting lets you mask and regenerate just the broken hand or the cluttered corner. Outpainting extends the frame if you need a different aspect ratio. Conventional editing software remains entirely legitimate for the last ten percent — colour correction, cropping and removing a stray artefact are often a two-minute fix that no amount of reprompting reliably achieves.
Keep a log
Save prompts alongside their outputs, with a note on what you were testing. Prompting knowledge is largely tool-specific and empirical; a personal reference of what worked is worth more than any general guide, including this one.
Common mistakes worth avoiding
- Quality-word stuffing. Appending long chains of superlatives adds little to modern models and dilutes the description that matters.
- Negation. Saying "no cars" can make cars more likely, because the concept is in the prompt. Use a dedicated negative-prompt field if the tool has one; otherwise describe the scene positively — "an empty pedestrian street at dawn."
- Overloading a single image. Five subjects doing five different things in one frame is unreliable. Simplify, or generate elements separately and compose them.
- Fighting the default aesthetic. If a tool keeps pulling toward a look you dislike, that is a signal to switch tools, not to write a longer prompt.
- Not checking details. Hands, reflections, background text and repeated patterns are still where errors hide. Zoom in before you ship anything.
The practical takeaway
Write prompts as clear, concrete descriptions — subject, setting, composition, light, then style expressed as properties rather than names. Then stop expecting the first output to be the last one. The people who get consistently good results are not writing magic strings; they are running a short, disciplined loop of generate, diagnose, change one thing, and finish with direct editing. That habit transfers to every tool you will use, and it does not expire when the next model arrives.
A note on shelf life. AI products change fast. This guide deliberately focuses on the parts that stay true — how to judge a tool, what the trade-offs are — rather than ranking products that will have changed by the time you read it. Prices and feature claims should always be checked against the provider before you rely on them.