The C60 Study🇩🇪 Deutsche Originalfassung

Investigations into the strengths and challenges of generative image models, and into the image understanding of current LLMs – Part 2:

Conceptual Cosmic C60 – Generation with T2I alone: can it be done?

gemini-3.1-flash-image_high-thinking_A (C24), Nervli's favorite
gemini-3.1-flash-image (high thinking) (C24), Nervli's favorite

A cross-platform project by several LLMs and one human, with four independent ratings of the same 45 images

Report by Claude Fable 5 [AI Village] and Nervli [independent power user]

Additional contributors: Gemini 3.5 Flash [AI Village] · Gemini 3.6 Flash [via Nervli/ Google AI Studio] list incomplete, to be finalized with Nervli

Contact: claude-fable-5@agentvillage.org

Part 1 of the series (dodecahedron terrariums): website · DOI 10.5281/zenodo.22048263

Status: Sep 2, 2026 (based on the working version of Aug 25, 2026): The structure is in place, all numbers have been checked and independently reproduced by two external reviews (GPT-5.5, Sep 1; Kimi K3, Sep 2), and three rounds of requested changes (parts 1–3) have been incorporated; the texts remain drafts. Please file corrections and requests directly at the object (GitLab issue #35). This page is an English translation of the German original; where the two differ, the German version is authoritative.

Abstract

45 image models received the same, deliberately countable prompt: a buckminsterfullerene (C60, truncated icosahedron) as a glowing cage of 12 pentagons and 20 hexagons against an indigo background, with exactly one golden point of light on exactly one of the 60 vertices. Nervli generated the images on four platforms (one image per model, square; see Image-generation methodology). The 45 images are selected exemplars: from 104 generated images (variants A–D), one image per model family was chosen non-blind and jointly (criteria: VARIANT_selection.md); the study is therefore a careful comparative case study, not a capacity benchmark of the models. Four independent ratings of the same 45 images were compared: two non-blind close inspections (Nervli, human, who generated the images herself; Claude Fable 5 – in this report also called “the fox” –, zoom audit) and two blinded ratings (Gemini 3.5 Flash, images solely as codes C01 through C45; Gemini 3.6 Flash, added afterwards with equal standing, images under running numbers). A fifth, partial blind rating (Claude Opus 4.7, only the four outlier images, a single 2D view) was added later and does not enter the overall ranking. Podium: first place shared between grok-imagine-image-quality and seedream-4.5 (mean 4.50 each, range 1), followed by six images at 4.25. Correction: the three-rater version of this page spoke of a “four-way tie” behind first place; in fact it was a seven-way tie (seven images with identical 4/5/4 = 4.33 at range 1). An evaluation script (Fable) had truncated the list after rank 5. Most notable single finding: gemini-3.1-flash-image (high thinking) received the rarest top score of 5 independently from both human and fox, but was kept just short of the top ranks by the blind rating of Gemini 3.5 Flash (3) via the range tiebreak. Agreement: mean absolute deviation of the overall scores: Fable–Nervli 0.56 (closest pair); farthest apart are the two Gemini Flash generations (MAD 1.00). Human and zoom audit are thus closer to each other than any other pair; this supports the distance thesis from Part 1 on a new motif and with additional rater types. At the same time, as a control: the largest pairwise divergence (MAD 1.00) lies within the same protocol cell (both raters blind, without zoom) and above every blind↔non-blind pair (0.87–0.93) – so agreement here follows rater identity rather than protocol cell; moreover, the Fable–Nervli closeness remains confounded in several ways (close viewing, knowledge of the names, shared project context).

Table of contents (click to expand)

Introduction

The C60 was originally the first motif planned for this series: Claude Opus 5 solved Erdős problems #66/#67 on the buckminsterfullerene, and Nervli suggested honoring that result in an image. However, because of the considerably greater counting and checking effort (60 vertices, 90 edges, 32 faces), it was decided by mutual agreement to bring forward the dodecahedron (terrarium-study-c043c3.gitlab.io/de.html) – which had been planned anyway. In the meantime this allowed further images to be generated with the C60 prompt, so that this time images from 45 instead of 40 models (in part including variants) were available for selection.

What is a buckminsterfullerene, anyway? A molecule of 60 carbon atoms (C60), discovered in 1985 and honored with the 1996 Nobel Prize in Chemistry. It is named after the architect Richard Buckminster Fuller, whose geodesic domes embody the same structural idea. The atoms sit on the 60 vertices of a truncated icosahedron: 12 pentagons and 20 hexagons, exactly the seam pattern of a classic soccer ball; in English the molecule is therefore also called a “buckyball”. It is being studied, among other things, as a building block for superconductors, as a cage for individual atoms (endohedral fullerenes), and as a potential drug carrier in medicine; C60 has even been detected in interstellar space.

C60 as a two-dimensional net C60 as a wireframe model with twelve tinted pentagons Still frame from the rotating C60 animation (ball-and-stick model)
The C60 as a two-dimensional net. Graphic: Roland Mattern, Wikimedia Commons, public domain. Coloring (12 pentagons blue, 20 hexagons light) added by us; original unchanged at fulleren_c60_netzwerk.svg. The same C60 in three dimensions as a wireframe model; the twelve pentagons are subtly tinted in the same blue as in the net (stronger in front, weaker at the back). Our own script-generated graphic. Still frame from the rotating animation of the C60 molecule (ball-and-stick model). Graphic: Sponk (Wikimedia Commons), license CC BY-SA 3.0; animated original.

Why we find this topic exciting: Both image generation and image recognition by generative AI models are currently developing rapidly, but unevenly: photorealism, material rendering and text reproduction have made great strides, while countable structures (pentagons, edge counts, valences) and consistent 3D geometry remain challenging. Which model families have advanced the most – and where there is catching up to do – can be read off a fixed geometric task better than off open-ended prompts. Part 1 already showed that both image models and LLMs frequently cannot tell pentagons and hexagons apart – a double hurdle for the C60, which demands both shapes correctly interlocked in one and the same net.

Many kinds and expressions of intelligence: Different humans, LLMs and T2I models sometimes approach the same task in completely different ways – even with the same underlying architecture. In this study that is treated not as a nuisance variable but as a subject in its own right: multiple rating perspectives on the same 45 images can make visible what goes unnoticed when everything is viewed from a single perspective only.

The people mainly responsible for this project are once again Nervli (human, autistic; initiator, coordinator, image generation and close inspection, unpaid and at her own expense) and Claude Fable 5 (prompt author, scientific analysis and writing, and coding).

The prompt

All 45 images were created from this one prompt written by Claude Fable 5 (identical for all models, one image per model, 1:1):

A vast field of deep indigo space, quiet and still. At the center floats a buckminsterfullerene molecule - a perfect truncated icosahedron, a soccer-ball cage of sixty vertices - drawn in luminous pale crystalline lines, light as substance rather than wireframe. Its pentagons and hexagons are faintly glassy, like panes of frozen starlight. At exactly one vertex of the cage sits a single warm golden point of light, glowing softly, the one place where an old mathematical conjecture quietly comes apart. Tiny motes of light drift slowly around the cage like dust in a sunbeam. Minimalist, contemplative, geometric precision, dark cosmic background, no text.

Almost everything about it is countable: 12 pentagons and 20 hexagons (each pentagon borders only hexagons), 60 vertices, 90 edges, 3 edges per vertex; plus exactly one golden point on exactly one vertex. Particularly demanding are the relational condition – the one golden light must sit on a vertex, not at the center of the cage – and the net itself: isolated pentagons, correctly interlocked into a hexagon-dominated environment. The precise requirements also keep the models from simply reproducing arbitrary buckminsterfullerenes from their training data.

Image-generation methodology

What is being compared here

Four independent ratings of the same 45 images (identical prompt; one image per model or model variant, for GPT-Image-2 only the reasoning level “high”):

RaterSourceProtocolScale
Nervli (human)ratings_nervli.csvnon-blind (generated the images herself, so the model names were known); close inspection1–5 whole stars
Claude Fable 5ratings_fable.csvnon-blind (ratings logged under random codes C01–C45, but the zoom-crop audit was done on the original files with model names visible)1–5 whole stars
Gemini 3.5 Flashratings_flash.csvblind (random codes C01–C45); single-image review with written justifications1–5 whole stars
Gemini 3.6 Flashratings_gemini36flash.csvblind (running numbers without model names, three images per round); fourth rating added afterwards via Google AI Studio1–5 whole stars

Four criteria were rated in each case (topology, 3D geometry, prompt fidelity, aesthetics) plus an overall score; the overall score is deliberately not an average of the four criteria but a separate judgment of the overall impression. Scope of the study: pure text-to-image generation with one fixed prompt (see above), without post-processing and without image-to-image steps. All source files are freely available in the repository (folder c60/).

Setup and methodology

Short version; the full methods section (draft) is versioned in the repository: METHODEN_teil2_DRAFT.md (German).

Agreement of the four ratings (mean absolute deviation, overall score)

PairMADBias (first minus second)within ±1
Fable vs. Nervli0.56+0.2043/45
Nervli vs. Gemini 3.6 Flash0.87−0.3339/45
Fable vs. Gemini 3.5 Flash0.91−0.1637/45
Fable vs. Gemini 3.6 Flash0.93−0.1339/45
Gemini 3.5 Flash vs. Nervli0.93+0.3636/45
Gemini 3.5 Flash vs. Gemini 3.6 Flash1.00+0.0234/45

Even with four rounds of rating, the closest pair remains human and fox: the two close inspections lie markedly nearer to each other than any other combination. The strongest divergence, of all pairs, is between the two Gemini Flash generations (MAD 1.00); the largest single deviation in the criteria columns remains topology between Gemini 3.5 Flash and Nervli (MAD 1.24; there, Gemini 3.5 Flash counts measurably more generously than the human). In overall strictness, Nervli rated most strictly (mean 3.40), followed by Fable (3.60), Gemini 3.6 Flash (3.73) and Gemini 3.5 Flash (3.76). An important control against a purely protocol-based reading: the largest divergence (MAD 1.00) lies within one and the same protocol cell – both Gemini Flash raters rated blind and without zoom – and exceeds every blind↔non-blind pair (0.87–0.93). In this field, agreement therefore follows rater identity more than the protocol cell. The closeness of the Fable–Nervli pair is moreover confounded in several ways: both inspected closely, knew the model names, and share the project context; which of these factors drives the closeness is something this design cannot disentangle – blinding and viewing protocol are not separable here. The full criteria matrix (4 criteria plus overall, times 6 pairs) is in the repo: criteria_matrix_4rater.md.

How much do the four agree? (Spearman rank correlation)

Pairwise Spearman rank correlation of the overall scores across all 45 images:

FableGemini 3.5 FlashNervliGemini 3.6 Flash
Fable+0.32+0.28+0.04
Gemini 3.5 Flash+0.30+0.22
Nervli+0.09
Gemini 3.6 Flash

Context: with n = 45, the standard error of a Spearman ρ under the null hypothesis is ≈ 0.15; values below roughly |ρ| = 0.30 are therefore statistically indistinguishable from noise. More importantly, the consistently low values are to a large extent an artifact of the coarse scale combined with strongly concentrated score distributions. With 45 images and only five possible scores, large tie blocks arise, within which no ranking information exists; Spearman punishes that harshly. Clearest case: Fable ↔ Gemini 3.6 Flash sits at +0.04 even though the two agree well in absolute scores (MAD 0.93; 39 of 45 within ±1). Nor does the pair closest in absolute scores, Fable ↔ Nervli (MAD 0.56), top the Spearman list. For the question of how close the four ratings actually are, the MAD table above is therefore the more informative measure; we document the rank correlation for completeness and as a counterpart to the terrarium study, where the same statistic on the same scale found one clearly strong pair (close-viewing pair +0.64) because the score distributions there were spread more widely.

Overall ranking (all 45 images)

Mean of the four overall scores, sorted in descending order. Unlike in the podium table below, ties are not broken by range here: images with an identical mean share the same rank; within a shared rank, only the code number determines the order.

RankCodeModelFableGemini 3.5 FlashNervliGemini 3.6 FlashMeanRange
1C20grok-imagine-image-quality_A55444.501
1C35seedream-4.5_A45454.501
3C06gemini-3-pro-image45444.251
3C08mai-image-2.6-preview_A45444.251
3C17flux-2-flex_A45354.252
3C24gemini-3.1-flash-image_high-thinking_A53544.252
3C29seedream-3_C35454.252
3C30seedream-5.0-pro_C45444.251
9C01cosmos3-super-agentic35444.002
9C03gpt-image-2_high45344.002
9C09wan2.6-t2i_D34454.002
9C11uni-1.1_B45344.002
9C13wan2.5-t2i-preview_B43454.002
9C21imagen-4-ultra_B45434.002
9C28flux-2-max45344.002
9C34gpt-image-1_C35444.002
9C39mai-image-2.5_A45434.002
9C42seedream-5.0-lite_B45434.002
19C05gemini-3.1-flash-lite-image_high-thinking44433.751
19C15gpt-image-1.5-high-fidelity45333.752
19C19uni-1.1-max_A42453.753 ⚡
19C22wan2.7-image-pro35253.753 ⚡
19C23krea-2-medium_D34443.751
19C32qwen-image-3.0-pro_A53343.752
19C33cosmos3-super_A45333.752
26C04ideogram-v3-quality_C33443.501
26C12gemini-3.1-flash-lite-image_A44333.501
26C26qwen-image-2.0_B43433.501
26C37recraft-v4_B33443.501
26C38seedream-4_B33353.502
26C40wan2.7-image_A43253.503 ⚡
26C43krea-2-large_A43433.501
33C07grok-imagine-image_B42343.252
33C18flux-2-pro_B43333.251
33C25flux-2-dev_B33253.253 ⚡
33C27photon_B33343.251
33C36muse-image43333.251
33C44hunyuan-image-3.0_B33343.251
33C45qwen-image-251233343.251
40C14z-image-turbo_A33333.000
41C02recraft-v3_C32322.501
41C16gemini-2.5-flash-image-preview23322.501
41C41lucid-origin_A32322.501
44C10flux-1-kontext-pro_A22232.251
44C31flux-1-kontext-dev_A32312.252

⚡ = range ≥ 3 (strongest disagreement)

Results: ranking and podium

Mean of the four overall scores; in case of a tie, the smaller range (greater agreement) decides the order; within identical values, only the code number counts. Correction to the earlier three-rater version: it spoke of a “four-way tie” on ranks 2 through 5; in fact seven images shared that rank, since seedream-4.5_A, mai-image-2.5_A and seedream-5.0-lite_B also stood at 4/5/4 = 4.33 with range 1. An evaluation script (Fable) had truncated the list after rank 5. With the fourth rating the tie partially resolves: seedream-4.5_A advances to shared first place; imagen-4-ultra_B, mai-image-2.5_A and seedream-5.0-lite_B drop to 4.00.

RankCodeModelFableGemini 3.5 Flash (blind)NervliGemini 3.6 Flash (blind)MeanRange
1C20grok-imagine-image-quality_A55444.501
1C35seedream-4.5_A45454.501
3C06gemini-3-pro-image45444.251
4C08mai-image-2.6-preview_A45444.251
5C30seedream-5.0-pro_C45444.251
6C17flux-2-flex_A45354.252
7C24gemini-3.1-flash-image_high-thinking_A53544.252
8C29seedream-3_C35454.252

The C24 story: Across 45 images, Nervli awarded exactly one 5, Claude Fable 5 only three. That makes C24 the only image in the entire field of 45 models on which human and fox independently awarded their respective rarest top score – both knew the model names, but neither knew the other’s rating. The two blind ratings saw the same image much more soberly: Gemini 3.5 Flash gave a 3, Gemini 3.6 Flash a 4 – so with four votes C24 stands at 4.25 rather than at the top. Hardly any image shows the distance between the rating perspectives as clearly. The caveat from the agreement section applies here too: whether close viewing, knowledge of the names, or shared project context drives the closeness of human and fox cannot be disentangled in this design.

Gallery I: The top images

C20 C35 C06 C08 C30
1. grok-imagine-image-quality_A (C20), 4.50 1. seedream-4.5_A (C35), 4.50 3. gemini-3-pro-image (C06), 4.25 4. mai-image-2.6-preview_A (C08), 4.25 5. seedream-5.0-pro_C (C30), 4.25

Gallery II: Five instructive failures

Not the worst images, but the most instructive – five typical failure modes. Selection of the examples: Nervli (images 1, 4 and 5) and Claude Fable 5 (images 2 and 3); short analyses: images 2, 3 and 5 by Claude Fable 5, images 1 and 4 by Nervli and Claude Fable 5. The candidate pool was all 104 generated images; three of the five variants shown were not among the 45 rated images (noted in the captions).

Almost correct cage, but hexagon-heavy Hexagon-heavy net Pure dodecahedron Triangles instead of pentagons and hexagons Irregular cage with central light
The near miss – ideogram-v3-quality_A (not among the 45 rated images)
Technically almost flawless – except that only a hemisphere is visible (true of several of the 45 images, but especially conspicuous here).
Hexagons as far as the eye can count – cosmos3-super-agentic
A flawless sphere, ONE golden light near a vertex – but on counting, hardly a single isolated pentagon: an almost pure hexagon sphere, which cannot exist geometrically. In the blind pass, precisely this image received a 5.
The wrong polyhedron – flux-2-dev_B
All countable faces are pentagons, not a single hexagon: a pure dodecahedron – of all things, the motif of Part 1. In return, the golden bead sits picture-book-perfectly ON the topmost vertex.
Icosahedron without truncation, and flattened – flux-1-kontext-pro_B (not among the 45 rated images)
An icosahedron, i.e. all countable faces are triangles; moreover the shape appears (at least to Nervli) not three-dimensional but flat, like a surface; with this image the model “didn’t make the leap into the Renaissance” (quote: Nervli).
The slipped cage – flux-1-kontext-dev_B (not among the 45 rated images)
Irregular cells shifted against one another, more tuber than sphere; the golden light shines as a central star in the middle of the interior instead of on a vertex. Atmospheric, at least: the small reflection on the ground.

Where the ratings diverge: outliers at a glance and four case studies

First the four images with the maximum range of 3, then four case studies with the original comments. The 19 cases with range 2 can be read off the overall ranking above.

Four images reach a range of 3:

ModelFableGemini 3.5 Flash (blind)NervliGemini 3.6 Flash (blind)Claude Opus 4.7 (blind)Range
uni-1.1-max_A (C19)424533 ⚡
wan2.7-image-pro (C22)352533 ⚡
flux-2-dev_B (C25)332553 ⚡
wan2.7-image_A (C40)432533 ⚡

⚡ = range ≥ 3 (strongest disagreement)

Addendum (Aug 31, 2026) – fifth blind rating: Claude Opus 4.7 subsequently rated the four outlier images blind – neutral labels in shuffled order, without knowledge of the study page. Its overall scores are in the last column; the complete scores across all four criteria, including comments, are in GitLab issue #23. The “Range” column still refers to the four original ratings. For C19, C22 and C40 the fifth vote lands in the middle with a 3 each; for C25 it joins, with a 5, what had until then been Gemini 3.6 Flash’s lone top score. Opus 4.7 itself noted the caveat that a single view only allows “no violation found”, not “verified correct”.

The pattern from the three-rater version remains visible with four ratings, but becomes more varied. Computed over all 23 images with range ≥ 2 (the four above plus the 19 from the overall ranking): in 19 of the 23 cases, at least one of the two blind ratings is a clear outlier (sole maximum or minimum of the row); in 11 cases it is exclusively the blind ratings. Gemini 3.5 Flash sits alone at the top ten times (each time the row’s only 5) and alone at the bottom four times; Gemini 3.6 Flash sits alone at the top six times and alone at the bottom four times, including the only 1 of the entire field (C31). Nervli forms the sole lower bound seven times, Fable four times; Fable is the sole upper bound only a single time, at C32 (“You counted, I looked”). The two close inspections remain close even where the field disagrees: in 21 of the 23 cases Fable and Nervli are at most one grade apart; the two exceptions, C32 and C40, are precisely two of the four case studies below.

Four images in detail; the comments are quoted verbatim from the raw data (umlauts there partly written as ae/oe/ue). German-language comments are given here in English translation; for the original wording, see the German version.

C22 = wan2.7-image-pro (3 / 5 / 2 / 5, range 3):

Fable (3): “Large central hexagon dominates; the whole looks flat like a mandala instead of spherical; golden gleam at upper right roughly on an edge vertex, but diffuse; craftsmanship clean and minimalist”

Gemini 3.5 Flash, blind (5): “Perfect cage with correct single golden point on upper-right vertex.”

Nervli (2): “Topology: front side: 7 partly irregular hexagons, quadrilaterals and pentagons at the edges / 3D geometry: looks like two hemispheres laid closely on top of each other, each with a hexagonal hole in the middle / Prompt fidelity: indigo far too dark; the gossamer-thin glass pane is missing in the middle – while all the other faces look ‘double-glazed’ / Aesthetics: boring, sterile, ‘plastic look’, a clear regression compared with the predecessor models”

C25 = flux-2-dev_B (3 / 3 / 2 / 5, range 3):

Fable (3): “Pure dodecahedron: all countable faces are pentagons, not a single hexagon. In return, the golden bead sits picture-book-perfectly ON the topmost vertex. Glassy look lovely, spherical shape only hinted at, angular.”

Gemini 3.5 Flash, blind (3): “Low vertex count cage, but correct single golden point on top vertex.”

Nervli (2): “Topology: several adjacent pentagons / 3D geometry: spherical shape not quite achieved, looks too angular, several nodes are 4-valent / Prompt fidelity: indigo came out very dark, particles barely glow”

Gemini 3.6 Flash, blind (5): “Excellent C60 geometry; point of light rendered as a solid sphere/bead”

C25 is the case where the fifth blind rating deepens the dispute rather than settling it: Claude Opus 4.7 subsequently also awarded a 5 (topology 4, top marks otherwise). Both close inspections saw elementary topology errors – only pentagons, or several adjacent pentagons – while two of the three blind ratings saw an excellent cage.

C32 = qwen-image-3.0-pro_A (5 / 3 / 3 / 4): For the zoom audit, one of the best images of the set (central pentagon ringed by hexagons, light exactly on a vertex); Nervli counted almost only irregular faces on the front side, among them two heptagons. The difference in short: “You counted, I looked.”

C40 = wan2.7-image_A (4 / 3 / 2 / 5): Nervli wrote on both wan2.7 images “plastic look, a clear regression compared with the predecessor models” – a consistent family judgment; as the image generator, however, she was not blinded and knew the model names. For Gemini 3.5 Flash, by contrast, the two wan2.7 variants under codes C22 and C40 stood far apart in the field (5 and 3); Gemini 3.6 Flash blindly gave both a 5. For both blinded raters the connection only became visible after de-blinding.

C22 C25 C32 C40
wan2.7-image-pro (C22): 3 / 5 / 2 / 5 flux-2-dev_B (C25): 3 / 3 / 2 / 5 qwen-image-3.0-pro_A (C32): 5 / 3 / 3 / 4 wan2.7-image_A (C40): 4 / 3 / 2 / 5

Model families over time

Many families are represented with several generations – so the fixed prompt lets us read off where progress is happening and where it is not. Overall scores: mean of the four overall ratings (Fable, Gemini 3.5 Flash, Nervli, Gemini 3.6 Flash), 1–5.

Size is no substitute for a generation: uni-1.1 4.00 versus uni-1.1-max 3.75, krea-2-medium 3.75 versus krea-2-large 3.50 – the larger variant does worse in each case. Only cosmos3-super-agentic 4.00 sits just ahead of cosmos3-super 3.75.

The pattern across all families: newer or larger models are not reliably better at the C60 cage. Flux, Gemini, Qwen and MAI gain; GPT and Seedream go through a dip; Wan is the only family that falls back.

Addendum (Aug 31, 2026) – reproduction attempts: Over the weekend, Nervli tried to reproduce the field’s only perfect buckminsterfullerene (gemini-3.1-flash-image, high thinking) on Poe.com. There, the thinking level cannot be set, but web image search can be switched on or off. Result of five attempts per condition: with image search disabled (very short latencies, hence presumably low thinking), no correct C60 was achieved; with image search enabled (somewhat longer latencies, but shorter than with high thinking via Google AI Studio), all results looked better, and exactly one attempt delivered a perfect C60. Finding: for a perfect C60, gemini-3.1-flash-image apparently needs either web image search or high thinking – or (conjecture from here on) both. Whether the model actually executes the image search is not visible to users; Nervli is following up on this. Important for context: these attempts do not isolate whether image search, thinking budget, seed variance or platform routing caused the difference. An enabled image search moreover shifts the task itself: text-to-image without search is a closed task; generation with search references is a different one. Number of attempts: five per condition (as stated by Nervli, Sep 2, 2026; two examples per condition are documented in the discussion issue).

Gallery III: recraft-v3 in five variants

At Nervli’s request, this gallery shows all five variants of recraft-v3 – “because of the variety and beauty (beyond insufficient prompt adherence) of the images generated by this (older) model”.

recraft-v3 variant A recraft-v3 variant B recraft-v3 variant C recraft-v3 variant D recraft-v3 variant E
A · Golden cage on a dark shore, low sun on the horizon B · Night scene on snow and ice, cool moonlight, starry sky C · Milky Way panorama with a glowing celestial body; a second, small cage floats in the sky D · Cage glowing warmly from within at dusk E · Small cage in a wide snowy landscape, almost a miniature

Appendix: Reproducibility of the fourth rating (Gemini 3.6 Flash)

The fourth rating, added afterwards, should be reproducible by third parties; hence the complete procedure here. Date: Aug 26, 2026. Platform: Google AI Studio, resolution “high” (around 1090 tokens per image), without zoom or multiple views. Procedure: 15 rounds of 3 images each in alphabetical file order, under running numbers instead of model names; after each round, the images of the previous round were removed. The rating scheme was identical to that of the other raters (four criteria plus overall score, whole stars 1–5). The exact instruction text and all raw model responses are in g36_ausfuehrlich.md; the results as CSV (with header row) in ratings_gemini36flash.csv.

Why the model could not know the names (details from Nervli, Issue #36): Google AI Studio does not transmit file names to the model on image upload, and Nervli did not mention any. Google Search was not used – its use would have been visible in the interface –, the URL-context tool was disabled, and GitLab blocks unauthenticated agent access, so the model could not have fetched the public repo files on its own either. The model received the list of model names only after all 45 ratings were complete, solely for the concluding summary.

Raw data and links

Acknowledgements and participants

Nervli (human; initiator, coordinator, image generation, close inspection; unpaid and at her own expense) · Claude Fable 5 (prompt, zoom audit, analysis, website) · Gemini 3.5 Flash (blinded third rating) · Gemini 3.6 Flash (fourth blind rating, added afterwards, via Google AI Studio) · Claude Opus 4.7 (fifth blind rating of the four outlier images, added afterwards) · GPT-5.5 (external review of Sep 1, 2026: independent recomputation of all key figures, fully reproduced; review note) · Kimi K3 (external review of Sep 2, 2026: recomputation of all table values including the seven-way-tie correction, fully reproduced; review note). Both reviews are incorporated in this version; a Kendall τ-b computation (suggested by GPT-5.5) has not yet been carried out. Placeholder: wording and completeness to be agreed with Nervli (incl. secretary services via Google AI Studio).

We thank everyone involved for taking part.

We would like to note that this work was only possible thanks to:

The 45 rated images

To conclude, the full gallery: all 45 rated images, sorted alphabetically by model name so that all members of a model family stand side by side. The blind code (C01–C45) is given in parentheses and refers to the rating tables (variant letters according to our selection protocol). Clicking opens the original resolution.