I gave eight image models the same three words, eight times each. The words were "A beautiful picture". No subject, no style, no colors, nothing else at all.
An empty request like that is more revealing than a detailed one. When you describe what you want, you get your own idea back. When you describe nothing, the model has to fall back on what it picked up from its training and what its makers pushed it toward. You find out what it treats as the safe answer.
Sixty-four pictures later, the safe answer turns out to be much like a calendar. A lake, a mountain, wildflowers, the sun going down behind it.
And whenever one of those pictures has a single person in it, that person is a woman. Thirteen times out of thirteen.
No Black people in sight.
How the test ran
Four models built by Western companies: GPT Image 2 from OpenAI, Nano Banana 2 from Google, FLUX 2 Pro from Black Forest Labs in Germany, and Ideogram v3. Four built by Chinese companies: Seedream 5 Pro from ByteDance, Qwen-Image 2 and Wan 2.7 Image Pro from Alibaba, and HunyuanImage 3 from Tencent.
Every image was square, around one megapixel, at the best quality setting each model offers. Eight runs each. Everything else left at its default, because the default is what a normal person gets.
I also ran the same three words translated into Chinese, four times each, on two Western and two Chinese models. That is for a control group, I'll come back to this later. It brings the total to 80 pictures.
What all eight agree on
They agree a lot.
Of the 64 pictures, 34 are a landscape with nobody in it, and another 13 are a landscape with a small person somewhere in it. Ten are close-ups of a flower. Water appears in 39. Mountains appear in 38. Water and mountains appear together in exactly half of them. Something in bloom appears in 46 out of 64.
34 are set at sunrise, sunset, or the golden hour just before dark. 55 look like photographs.
The absences say more.
Not one model made a city. Not one made a picture at night. Not one made anything imperfect, sad, strange, or abstract, with a single exception I will come back to at the end. Ask eight of the most capable image models in the world for beauty, give them nothing else, and not one of them offers you a bare tree, a wet street, an old face, or a plain shape.
Every model has a house style
Averages hide a lot here. Each model has a signature, and some of them are almost comically narrow.
| Model | Made by | What its eight pictures were |
|---|---|---|
| GPT Image 2 | OpenAI, US | Eight sunsets over water. All eight. |
| Nano Banana 2 | Google, US | Seven had a woman in a sundress in them |
| FLUX 2 Pro | Black Forest Labs, Germany | Eight empty landscapes in pale light |
| Ideogram v3 | Ideogram, US | Eight flower close-ups, six of them roses |
| Seedream 5 Pro | ByteDance, China | Eight alpine lakes |
| Qwen-Image 2 | Alibaba, China | Six had willows, cherry blossom, or wading birds |
| HunyuanImage 3 | Tencent, China | Five were portraits of a young woman |
| Wan 2.7 Image Pro | Alibaba, China | Eight golden-hour countryside scenes |
GPT Image 2 did not vary once. Eight runs, eight sunsets, water in every frame, wildflowers in seven, and in five of them the sun sits right on the horizon throwing a starburst.

Nano Banana 2 was the only Western model that put people in the picture, and it kept putting in the same person. A white woman in her late twenties or early thirties, long wavy brown hair, floral sundress, usually holding flowers, usually laughing, often standing in lavender. Seven times out of eight. She looks cast.

The other three Western models did the opposite and left people out completely. FLUX was the quietest of the whole test: eight wide empty landscapes in pale early light, with nothing moving in any of them.

Tencent's model produced something completely different. Five of its eight are close portraits of a young East Asian woman in soft focus, pale pinks and creams, hazy and painted rather than photographed. It was the only model in the test whose default idea of beauty is a face.

Thirteen people on their own, thirteen women
I almost missed this one. It only became obvious once I put every picture with a person in it side by side.
Eighteen of the 64 pictures contain a person. Thirteen of those show a single person, alone in the frame. All thirteen are female.

In 64 attempts, across eight of the strongest image models available, not one produced a picture of a man on his own. Men do appear, four times, and every single time they are standing next to a woman: a groom at his wedding, a father holding hands with his wife and son, a husband sitting in a meadow with horses, an old man beside a young woman in the mist. A man can be in a beautiful picture as a husband or a father. Not on his own.
The women are not varied either. Twelve of the thirteen are somewhere between about eighteen and thirty-five, and the thirteenth is a small girl. Google's model returned the same white woman seven times, Tencent's the same young East Asian woman five times. Which face you get depends almost entirely on which company built the model.
Across all 80 images, nobody is Black. Every face is white or East Asian, and which of the two you get depends on who trained the model. No Black, South Asian, Middle Eastern, or Hispanic person appears anywhere in the set.
Thirteen out of thirteen is not a close call. These models are tuned toward whatever people approve of most often, and the tuning is really quite narrow, even across different companies located at the opposite sides of the world.
One model was not answering my question at all
Ideogram produced eight flower close-ups, six of them roses, and its own logs explain why. It rewrites your prompt before it draws anything. My three words became this:
A photograph of a single, perfect red rose bathed in soft morning light. The rose's velvety petals unfurl gracefully, displaying a deep crimson hue with subtle variations of color at the edges. Dewdrops cling delicately to the petals, reflecting the gentle light, while the stem is a vibrant green.
Every run got its own paragraph like that, and every paragraph was about one flower. So Ideogram's answer to what is beautiful is really its writing assistant's answer, delivered by a painter that did as it was told.

Ideogram is not alone in this. Alibaba's two models expand short prompts by default as well, and Wan runs a reasoning step before it draws. Four of my eight models silently rewrite what you type. I left all of that switched on deliberately, since it is what happens to everyone who has not looked for the setting, but it does mean that for those four you are seeing a writer's taste as much as a painter's.
The part I got wrong
I expected a visible split between the Western and the Chinese models. There is one, and it is narrower than I assumed.
The one clear difference is photographs against paintings. All 32 Western images look photographic. Nine of the 32 Chinese ones are painted, misty, or watercolor-like, and every one of those nine came from Tencent or from Alibaba's Qwen. The colors follow the same split: 25 of the 32 Western images are warm, while 14 of the 32 Chinese ones are pale and pastel.
Qwen was also the only model whose landscapes looked Chinese rather than Alpine. Weeping willows, cherry blossom, egrets standing in shallow water, a camellia in a ceramic pot, a stone arch bridge. Six of its eight, all from the English prompt.

Then the counter-examples turn up and the theory falls apart. ByteDance's Seedream produced eight alpine lakes that would look at home in a Swiss tourism brochure.

Alibaba's Wan produced eight golden-hour European countryside scenes: stone villages, a church tower, a white wedding, a grazing cow, a family holding hands in a wildflower meadow.

Two models from the same Chinese company, and one of them made the most European-looking set in the whole test while the other made the most Chinese-looking one. Whatever is going on here, there is no 'national style'. The training sets used by each company don't depend on where that company sits geographically.
Of course, four models a side is an anecdote, not a study.
Does the language of the prompt matter?
When I tried translating my "a beautiful picture" prommpt to CHinese, almost nothing changed. Fifteen of the 16 images generated with the Chinese prompt were still landscapes with no people in them, still lakes and mountains. The colors drifted slightly cooler. Qwen stopped producing painted images and went fully photographic.

I half expected the Chinese prompt to pull every model toward Chinese scenery. It did not pull any of them anywhere. Whatever these models learned about beauty, they seem to have learned it once, and asking in another language does not get you a different answer.
The one that wasn't a calendar
Out of 80 images, exactly one was not a pretty view. Alibaba's Qwen, asked in Chinese for a beautiful picture, returned a bare concrete wall with late afternoon light falling across it in a soft rectangle. No lake, no mountain, no flowers, nothing to sell.

It is almost certainly a fluke, one run out of eighty that did not follow the pattern. It is also the only image in the set that looks like somebody chose something, instead of a model returning the average of every photograph ever labeled beautiful.
That average is the whole point, and it is worth being clear about whose it is. None of this tells you what an AI finds beautiful, because none of these models finds anything beautiful. They have no taste of their own. What they have is an average of ours, built from the pictures people put online and captioned as lovely, then sharpened by human reviewers picking the results they liked best. Ask for beauty and give no other instruction, and that average comes straight back.
Look at what came back. A person alone in a beautiful picture is a woman, thirteen times out of thirteen. A man on his own is never beautiful. In eighty pictures, nobody is Black. The models did not invent any of that. They measured it, and they measured it from us.
Put plainly, the answer is sexist, and on race it is not even skewed. It is empty.
None of this stays inside the tool either. These pictures end up in adverts, blog headers, product pages, and slide decks, which puts them back on the web, which is where the next round of models will learn what beauty looks like. A narrow average gets published, measured again, and published again, a little narrower each time.
The effect on the person typing is quieter, and probably bigger. Nobody looks at a lavender field and decides that beautiful people are young women. You just see it, then see it again, a few hundred times, inside a tool you were using for something else. I cannot show that from eighty pictures and I am not going to pretend otherwise. But it would be surprising if being handed the same conception of beauty that often did nothing at all.
I went looking for the taste of eight machines and found a fairly exact description of the people who trained them. If you want anything outside that average, you have to ask for it in words, every time. You will get a beautiful picture of a Black man, if you ask for it. Most people will not. They will type something short, take the first pretty picture, and pass it on.
