Closer #011

We all know things can get weird in AI image land. But what if you give it a weird image? Will we get twice the weirdness or will they balance each other out?
The image shared with ChatGPT as input.

Welcome to Closer,

A weekly experiment in AI interpretation: I share a photo, ask AI to describe it, then generate an image from that description. This is edition #011.

If you're not sure what something means then you just wing it and pretend like you do, right? At least that's what ChatGPT thinks, it seems.

Human experience → photo → AI description → AI image. Prompts are in grey boxes. Descriptions are in italics.

LLMs used

GPT-4o for image interpretation and image generation.

The input

✍️
Please describe the attached image.
The image shared with ChatGPT as input.
The image shared with ChatGPT as input.

The output

The image captures a sunlit street scene in Tarifa, Spain, with a strong sense of tranquility and Mediterranean charm. The composition focuses on the upper facades of two adjacent buildings under a deep blue, cloudless sky.

  • The left building is painted a warm beige tone with classical elements like small column balustrades and window frames trimmed in a soft yellow-gold hue. Its windows have various shutters and bars, hinting at residential use.
  • *The right building is a pristine white hotel & spa, marked by the text “HOTEL & SPA *” and a crest-like logo. Its architecture features arched colonnades and black wrought iron balcony railings. A banner showing a person diving into turquoise water hangs prominently, enhancing the sense of leisure and escape.
  • A single seagull soars through the vast open sky on the right side, reinforcing the coastal location and peaceful mood.

The warm color palette, clear light, and geometric simplicity create a dreamy, almost cinematic atmosphere.

✍️
Generate an image with an aspect ratio of 2:3 based on this description: {the description previously generated}
The image generated by ChatGPT based on its description of my image alone.
The image generated by ChatGPT based on its description of my image alone.

Impressions

The surrealist distortion of the two swimmers in the rooftop pool is heavily misinterpreted by ChatGPT and the resulting image is hilariously incorrect. While it's respectable that the small column balustrades are there to begin with on that left building, it's puzzeling to see the hotel being downgraded from four to two stars.

That being said, one could argue the overall composition of the AI image is better balanced. There's less negative space and the bird is much more prominent (or is it comically large?) which makes the AI image an easier one to digest. It's also a much more boring one.

That being said, I want to share a note I received from Ivar Blaauw, who is an avid AI-explorer in his own right, as a reply to last week's newsletter:

"I view it as how videogames, such as Hitman, GTA, Red Dead etc create a more 'perfect' world, so we (the players) can escape to locations which are more movie like.

AI is very similar; why would it replicate real life, if it can create a more perfect version of that.

Creating imperfections, and thus reality, is very difficult. More difficult than creating something perfect at least."

I think this is interesting take and perhaps speaks in defense of human creators. Yes, reality is messy, complex and imperfect. If you want to mimicking the imperfections, how do you do that? Where do you place the dents, the scratches, the paint peeling? Will you simulate years of wear and tear caused by changes in the weather? Will you look at where an image is logically made and cross-reference that with the supposed climate of that place to reach greater accuracy? Will you envision hordes of people scraping past a fence with their jackets, boots, bags, zippers, to get that authentic used look? In the context of gaming, do you want all those imperfections at all? Are games a form of escapism from all of life itself to begin with, which allows for greater creative freedom, or should it approximate real life as close as possible? Is the same true for these AI tools?

The only reason my original image works to begin with, is because of the distorted humans swimming. All the other elements, bird included, are there to support that main subject. Strip out the main subject and what you're left with is... not much.

If you have any feedback, my inbox is open. Otherwise, see you next week.

Mitch

Subscribe to the monthly newsletter

Just you, me, and some occasional notes from the field. No spam.

Join the conversation