← All classes
CLASS 5June 9 · AI Image Generation & Visual Prompting

"How do we hold onto our creative processes,in a world that both encourages and fears AI use?"

Sydney Seifert-Gram
Taught bySydney Seifert-GramArt director, photographer, designer
Educator, Creative AI Academy
Tony Jones
WithTony JonesCo-instructor, Pratt AI Design Certificate

The class in brief

Sydney and Tony's image class opened on three brand cases already living with AI in production, then built the image-prompt companion to GOLD: Format, Subject, Details, Style, Parameters. The back half was reference craft: weighting composition against style as a tug-of-war, evolving a reference instead of restarting it, and a five-step method for diagnosing a bad output before you regenerate. After this page you can write an image prompt in five deliberate slots, weight a reference on purpose, and troubleshoot a hallucination instead of just rerolling it.

The night at a glance

Why this matters · 6:37 PM

A person, a product, or a brand can now be extended, simulated, or created with AI.

Three cases set the stakes before any prompting started. H&M builds digital twins down to the birthmark , and the twin comparison ran under the line "finally a way for me to be in New York and Tokyo on the same day" . The human models own the twins outright and can license them to any brand, rivals included.

Hugo Boss runs AI imagery beside hand-stitched and stop-motion craft, treating it as "another visual language" rather than a shortcut . Four Seasons Condoms built Snapchat-native AI monsters to dodge Meta's safe-sex ad limits .

3
brand cases opened the night: H&M digital twins, Hugo Boss AI-plus-craft, Four Seasons Condoms' AI monsters.
6
months of production behind Four Seasons' two-minute film, "90% pro bono," including an eight-week continuous stretch.

The framework · 7:03 PM

The image prompt structure

01

Formatwhat kind of visual

Photo, oil painting, tintype, comic strip, Polaroid, watercolor. The mosaic test: a mosaic is the subject, a photo of a mosaic is the format.

02

Subjectthe hero

An elephant, a woman, a mountain, a mug, flowers, a couch. Whatever the image is actually about.

03

Detailswhat's happening

Wearing a beret, in a snowglobe, in a field, mid-stride, sneezing, on a bike. The action or situation around the subject.

04

Stylethe vibe, colors, textures, feelings

Fisheye, antique, warm light, monochrome, faded, double-exposed. Color lives here, not under subject.

05

Parameterswhat else to establish

Aspect ratio, --weird, character ref, remix. The Midjourney cheat sheet alone runs eleven flags, from --ar to --repeat.

Worked live as a tintype elephant on a bicycle , with one rule attached: longer prompts are not necessarily more effective, start with exactly what you need. The class then split image tools into conversational versus command-based lineups : chat-and-iterate tools like ChatGPT and Nano Banana Pro versus every-word-is-the-instruction tools like Krea, Midjourney, and Firefly.

The craft · 7:45 PM

Whatever weight you max out is what the model obeys.

A composition reference and a style reference get weighted separately, and the weighting is a tug-of-war . At a character reference weight of 50 and a style reference weight of 100 (Midjourney's Cref and Sref settings) on a wilting-tulip prompt, the AI tried to copy the grain and lighting, and composition played a smaller role .

Flip it to Cref 100 / Sref 50 and composition locks in while style gets deprioritized . Split it evenly at 50/50 and the AI finds balance, pulling from the prompt, both references, and its training data all at once .

50/50
is where the tulip demo found balance: enough flexibility for the AI to ask itself what flowers actually look like.
3
weight combinations tested live on the same prompt, same subject, before the class moved to reference evolution.

The exercise · 7:52 PM onward

Evolve the reference instead of restarting it

The activity

15 minutes: generate from the Miro reference images: style and composition.

"The output aspect ratio should mirror the aspect ratio of the reference image." Composition references count as Midjourney "image prompts": start with one of each, then stack .

Deliver: the reference images used, prompt notes, a few favorite outputs.

Feed your best output back in

A reference doesn't have to be perfect to be useful; every output shows what the model understands, ignores, or invents . Sydney's white-horse output wasn't the image she wanted, but the haze, lighting, and framing were perfect, so it became the next reference .

Five steps before you regenerate

Press pause instead of starting over from scratch : analyze how consistently the hallucination appears, check your references, evaluate your prompt for vague or sticky words, consider the training data, and use a multimodal tool to troubleshoot the meta question .

Use the specific word instead of the vague one

She chose "tobiano," a real marking term she googled, over "painted," because she knew the AI might get stuck on the vaguer word. Every generation costs money and resources, so the rule was to reuse an interesting hallucination rather than generate-and-discard it.

The judgment · 8:51 PM

Two ways to hold a series together, chosen by what you're trying to preserve.

Command-based consistency keeps every input fixed and changes one variable at a time. The full horse series ran white to brown to tobiano to dark brown to dappled grey, same composition throughout , ending in a five-image recap of the whole iteration .

Conversational consistency works the opposite direction: ask a multimodal tool to analyze an existing image's visual logic, then extend it into new subjects . ChatGPT's read of the oceanside bird image named it exactly: "realistic enough to identify as a bird, but softened into something dreamlike" .

5
horse portraits in the command-based series, one variable changed at a time: color, then marking, then setting.
8
images in the final conversational series: bird, jellyfish, sea holly, hibiscus, wave, coral, breaking wave, fish.

Methods and prompts

Five methods to take with you

METHOD 01 · TAUGHT 7:03 part of: structure your ask

Write it as Format, then Subject, Details, Style and Parameters

The image companion to GOLD. Five slots, filled in order, so you know exactly which word is doing which job when an output goes wrong.

Where it came fromSydney presented the image prompt structure as the visual companion to GOLD, five slots filled in order, format, subject, details, style, and parameters, so a stuck output can be traced back to exactly which word in which slot caused the problem.Use it whenUse this as the starting structure any time you write an image prompt from scratch, so each descriptive word has a clear job.

Working prompt

Format: [photo, painting, tintype, etc.]. Subject: [the hero]. Details: [what's happening]. Style: [the vibe, colors, textures, feelings]. Parameters: [aspect ratio, weirdness, character ref]. Start with exactly this, nothing extra.

You will know it worked whenthe image matches the five fields you filled in, with nothing extra added that you didn't specify in any of them.

METHOD 02 · TAUGHT 7:45 part of: build tools not chats

Weight your references on purpose

Composition and style are separate dials, and whichever one you max out wins. State the split explicitly, even in a tool without weight sliders.

Where it came fromSydney demonstrated with a tulip image that setting a style reference at full weight made the AI hallucinate to copy grain texture, while setting a composition reference at full weight dropped the described subject entirely, and that splitting both references around fifty percent found the actual balance.Use it whenUse this when you are working from both a composition reference and a style reference and one of them keeps overriding your prompt.

Working prompt

Use this as my composition reference at roughly 50% strength, and this as my style reference at roughly 50% strength: influence the output, don't copy either one directly. My prompt: [describe the subject and scene].

You will know it worked whenthe output echoes both references without copying either one outright, and the composition and style read like separate, blended influences.

METHOD 03 · TAUGHT 7:52 part of: context beats prompts

Evolve the reference instead of restarting

Feed your best imperfect output back in as the new reference and change exactly one thing. Every generation is a stepping stone, not a discard.

Where it came fromSydney described feeding her own imperfect output, a haunted-looking horse with perfect haze and framing, back in as the new reference and changing only one variable at a time, treating every generation as a stepping stone rather than something to discard.Use it whenUse this once you have a generated image that gets part of what you want right, so you build on it instead of starting over from scratch.

Working prompt

Here's my last output [attach]: the [lighting/framing/haze] is exactly right, but the [specific element] is wrong. Use this as the new reference. Change only the [one variable] and keep everything else as close to identical as possible.

You will know it worked whenonly the one element you named actually changed, and everything else in the new output still matches your last one.

METHOD 04 · TAUGHT 7:57 part of: reject the first draft part of: keep the judgment human

Diagnose your own hallucination first

Before handing a bad output to the model for troubleshooting, name your own guess. The five-step check works better when you've already staked a claim it can confirm or correct.

Where it came fromSydney's five-step hallucination check has you analyze the output, check your references, evaluate your prompt for vague or sticky words, consider the training data, and only then bring it to a multimodal tool, and she said the check works better when you have already staked a guess it can confirm or correct.Use it whenUse this when an image output has gone wrong and you want to diagnose it systematically before asking AI to fix it.

Working prompt

Here's my output and my prompt: [attach both]. I answer first, then you check me. My guess at what went wrong: [your diagnosis: reference, prompt wording, or training data]. Now walk the five-step check yourself and tell me where my read was right or wrong.

You will know it worked whenit states plainly whether your guess about the reference, the wording, or the training data was right, not a generic list of causes.

METHOD 05 · TAUGHT 8:51 part of: context beats prompts

Study your own image and then extend it

Conversational consistency the way ChatGPT did it for the ocean series: analyze an existing image's visual logic before asking for new subjects, and re-attach the original every round to fight drift.

Where it came fromSydney's conversational consistency method, illustrated with an ocean image series, has the model analyze an existing image's visual logic before generating new subjects in the same style, and re-attaching the original image every round to fight drift.Use it whenUse this when you want a series of new images that stay visually consistent with one image you already like.

Working prompt

I want more images in the same style as this one [attach]. Analyze its style, color, and rendering first. Then suggest new subjects that fit the same collection. Every time we generate one, I'll re-attach this original as the reference, not the desired output.

You will know it worked whenthe new subjects it suggests share the color and rendering it just described from your original, not a random unrelated set.

Where it broke · all night

The model does not understand absence and it fixates on the words you give it

"A still life of a pear on a background with no horizon line" gave back horizon lines, every time . AI models don't process a negative; describing what you don't want can still pull that exact idea into the frame .

The fix, "a seamless background," worked, until the word "paper" got added for texture and quietly reintroduced the horizon lines and odd textures it was supposed to avoid . The lesson holds past the pear: say exactly what you want, and watch for the one word the model won't let go of.

Try this prompt

Quiz me on the image prompt structure and the difference between composition and style references. Then give me a prompt full of negatives and make me rewrite it to say only what I want.

You will know it worked whenit quizzes you on the image prompt structure and the composition-versus-style split first, then makes you rewrite a negative-heavy prompt to say only what you want.

The shelf

Tools and references

Tools that night

  • Adobe Firefly, the only model Sydney called technically commercially safe, trained on licensed stock
  • Krea (native large/turbo, Flux, Imagen, Nano Banana, Ideogram, Seedream, Recraft), Midjourney, Figma Make
  • Seedance and Seedream named as the video and image models pointed at for tomorrow's class

Named in the room

  • H&M's digital-twin partner Uncut · Hugo Boss's analog craft artists
  • Four Seasons Condoms' agencies, AiCandy Australia and Emotive
  • Midjourney's public user galleries, browsable without an account, for harvesting vocabulary
← Class 4 · Prompt Design for LLMs CLASS 5 OF 24 Class 6 → Creative AI Video + Coded Design