The four-module structure behind stable image generation, with concrete English phrasing for each module and the mistakes that make results drift.
2026-09-25 · 约 3 分钟读完
作者:何屿(CPCX 撰稿人 · 图像与影音)
I kept the subject description identical and changed only the lighting words across eight generations — the results looked like they came from eight different models. Light is now the first module I write, and this article explains why.
Generating images by typing a phrase and rerolling until something looks right works about as well as it sounds. Text-to-image models respond to your words in a roughly predictable way, and the people who get consistent results are simply describing the image in a consistent order.
The order that works for most current models has four modules: subject, style, light, and composition. Written in that order, a prompt reads like a compact art brief — and each module can be varied independently, which is what makes iteration cheap instead of random.
Style words do the heaviest lifting after the subject itself, and vague ones (“beautiful”, “artistic”) do almost none. Choose from families of real treatments: for illustration — flat vector, watercolor, ink line art, children’s book illustration; for photography — 35mm film, documentary photo, studio product shot; for rendering — isometric 3D, claymation, pixel art.
A practical exercise: take one fixed subject and generate it in three different named styles. Comparing the three teaches more about what style words actually control than any list of examples.
If results look flat or amateur, the missing word is usually about light. “Soft window light”, “golden hour backlight”, “overcast diffusion”, “single hard key light with deep shadows” — each phrase moves mood, depth, and perceived realism more than any quality token.
Light also decides the story of the image. The same street scene in harsh noon sun and in blue-hour rain are two different photographs. When a generation is close-but-not-right, change the light before changing anything else.
Keep a small private list of light phrases that worked for your use case. It is the highest-leverage five lines a frequent image generator can own.
For matching sets — covers, stickers, product views — freeze everything except the subject. Write the style, light, and composition modules once, keep them identical, and swap only the subject text. Series drift almost always comes from re-wording the “fixed” part each time.
When a result is nearly right, prefer correction over reroll: image-to-image edits that keep the subject and change one module preserve everything you already liked. Reroll only when the overall direction is wrong.
The templates in our drawing category are pre-assembled with this four-module structure — swap the subject line and the rest of the prompt holds its shape.