top of page

Anatomy of an AI Image: There’s More Going On Here Than You Might Think

immortalconcepts
6 days ago
5 min read

One of the misconceptions about AI-generated images is that you type a sentence into a box, hit a button, and out pops exactly what you had in your head.

Sometimes you get lucky.


Most of the time, not so much.


Take Enzo and Sophia.


They’re two AI models I created and have been developing so I can use them in different locations, wardrobes and situations. There’s actually quite a bit that goes into creating the characters themselves, but that’s only the beginning.


The real challenge is being able to take them from Paris to Venice, change almost everything around them, and still make sure Enzo looks like Enzo and Sophia looks like Sophia.


And that’s where creating AI images starts feeling a lot more like directing a photoshoot than simply writing a prompt.


First, You Have to Create the Model


Before I can send Sophia walking through Paris, I have to decide what Sophia actually looks like.


You can certainly ask AI to create a random person for you. But if I have something specific in mind, I have to communicate those details and refine the results until the character looks right.


Face, hair, age, body type, proportions and other physical characteristics all become part of creating the model. And this is where the human eye starts playing a role very early in the process. AI might generate something that is technically very good, but that doesn’t necessarily mean it’s what I want. I still have to look at the result and decide what works, what doesn’t and what needs to change.



Sophia’s character reference. Establishing her appearance gives me a foundation for recreating the same character in different images.


Once I have a character I’m happy with, I can start putting her to work.

And unlike human models, she hasn't asked me about her day rate yet.


Now Try Getting Her to Stay Sophia


This is where things get interesting.


AI has a tendency to reinterpret characters each time you generate a new image. Facial features can change. Hair changes. Body proportions shift. Sometimes the differences are subtle. Other times Sophia suddenly looks more like Sophia’s vaguely related cousin.

For a single image, that may not matter.


For a series of images, it matters a lot.



Here I have Enzo and Sophia walking together in Paris.



Now, let's move the entire production to Venice.



Different city. Different environment. Different wardrobe, lighting, poses and composition.


But they still need to be Enzo and Sophia.


That becomes particularly important if you're creating images for a fashion brand, travel company, advertising campaign or any project where the same people need to appear across multiple pieces of content.


You could follow the same couple through Paris, Venice, Rome and Santorini. A fashion company could put the same model in an entire collection. A resort could create a series following the same guests through different experiences.


At that point, you're not simply generating individual pictures.


You're building a visual story.


Then You Have to Direct the Photoshoot


Once I have the characters and can keep them reasonably consistent, I still have to create the actual image.


Where are they?

What are they wearing?

How are they posed?

How are they interacting?

Where is the camera positioned?

What does the lighting look like?

What's happening in the background?

And most importantly, what is this image supposed to feel like?


You can leave a lot of those decisions up to AI. Sometimes it works. A lot of times it doesn't.


I've been photographing people for years, including portraits and fashion, and over time, you develop an eye for what looks right and what doesn't.


During a traditional photoshoot, I might tell someone to turn their shoulder, bring their chin down slightly, move closer to the other person, change what they're doing with their hands or look somewhere other than directly at the camera.


I'm essentially doing the same thing here, except now I have to communicate those directions to AI. And trust me, sometimes the human model was easier. 😂


The Camera Still Matters, Even When There Isn't One


The same thinking applies to composition and lighting.


I still have to decide where I want the viewer to be in relation to the subjects. I might want the image at eye level, slightly below them or shot from a three-quarter angle. I may want the characters walking toward the viewer rather than standing perfectly posed.


Then there's lighting.


Do I want hard midday sunlight? Soft window light? Golden-hour light? A cloudy day? Dramatic studio lighting? The environment needs direction, too.


In the Paris image, I want you to know they're in Paris, but I don't necessarily want the Eiffel Tower screaming, “HEY! LOOK! PARIS!” from the background.


It's part of the story, not necessarily the subject.


Those are the kinds of decisions photographers and creative directors have always made. The difference is that instead of controlling a physical camera, lights, location and models, I'm describing those choices to an AI system and then refining what comes back.


Knowing what you want is one skill.


Knowing how to communicate what you want to AI is another.



Pretty Isn't Enough


AI is remarkably good at creating pretty pictures. But pretty isn't necessarily useful. The image needs a reason to exist.


Why are Enzo and Sophia in Paris? Why are they walking instead of standing and posing? What does their wardrobe tell you about them? How close are they to each other? What does their body language communicate?


Change those decisions, and you change the story.


The Paris image could be part of a luxury travel campaign. Change the wardrobe, and it could become a fashion advertisement. Put Enzo and Sophia aboard a cruise ship, at a resort, having dinner overlooking the Caribbean or walking through Venice, and you're telling completely different stories with the same characters.


That's where this becomes much more interesting to me than simply generating an image.


You're building a world around them.


AI Can Create the Pixels. Someone Still Has to Decide What They Should Say.


That's the part of AI-generated imagery I think sometimes gets lost.


Yes, the technology is remarkable.


It can create images in minutes that could require a photographer, models, wardrobe, hair and makeup, travel, locations, lighting equipment and a production crew to create traditionally. But access to the technology and knowing what to create with it are two very different things.


Composition still matters. Lighting still matters. Wardrobe still matters. Body language still matters. Consistency still matters.


And story definitely still matters.


Most importantly, someone still needs to look at what AI produces and say:

“Nope. That's not it. Let's try again.”


The tools have changed.


The need for a creative eye hasn't.

 
 
 

Comments


bottom of page