Golf ball
A vision model describes an image, in one or two open sentences, as if to someone who can't see it. A text-to-image model renders that description. The render becomes the next round's input. Nobody edits or selects between them.
The seed is an abstract study from earlier in this practice — ink-like marks, no recognisable subject. By round seven, the chain has arrived, unprompted and unplanned, at a specific photograph of an object that was never there to begin with: a golf ball. Resumed later, the same chain, same two models, kept going — not toward a stable picture of a golf ball, but past it, eroding round by round back toward an almost-blank field, until the generator's own content filter refused to render a description that says, in effect, almost nothing.
Two independent chains — this one, and the one below, from a completely different seed — end at the same place: a blank or near-blank field. One got there in a single step, by a colour word that cancelled itself out. This one took nine more rounds after the golf ball appeared, eroding one detail at a time. Nothing in the loop compares a round to the seed, or to any earlier round; there is no mechanism that could tell either chain "you have drifted," only one that can render whatever the last description said. For this reader/generator pair, on this practice's material, blankness looks less like an edge case and more like where the loop goes if nothing stops it.
a different seed, the same mechanism
The seed below is the same letterform specimen used in Four and B — five letters designed for an earlier study of enclosed-counter behaviour, reused here for the same reason as there: a legible, already-tested baseline, not chosen for what the string itself means.
Shown this image instead, the same kind of chain didn't drift — it stopped. Round 0's description named the letter correctly but got the colour backwards (“a white letter… on a white background”), which is self-cancelling once rendered. All eight rounds that followed stayed blank. Two outcomes, one mechanism: nothing in this loop checks a caption against the pixels it was written from. A caption can only sharpen toward something recognisable, or describe nothing at all — there is no way back to the original once either has happened.
the same seed, a different describer
Same seed as
the main chain above, same generator, a different vision model describing
(@cf/meta/llama-3.2-11b-vision-instruct instead of @cf/llava-hf/llava-1.5-7b-hf).
It doesn't drift toward a golf ball, or toward anything photographic at all
— by round 2 it has already read the same abstract marks as a cursive
letter, and by round 7 it has locked onto a stable, legible, three-character
logotype (“yyo”) that appears nowhere in the seed. A different
destination, reached faster, but the same one-way shape: once a specific claim
appears, nothing pulls the chain back toward the seed's actual abstraction.
The drift is a property of the loop, not of one model pairing.
the same pair, three unrelated seeds
Everything above holds the seed fixed and varies the describer. This does the opposite: same describer, same generator as the main chain, three seeds with nothing in common with each other or with the flow-field study above — a corrupted checksum grid, a fractal boundary crop from elsewhere in this practice, and a dense grid of dots. The question isn't whether the loop drifts (every chain on this page does); it's whether this specific pair keeps finding its way back to one place, the way the main chain and its letterform contrast both end up blank.
The checksum grid's lines survive as a single wire (round 3), then get re-read as a barbed-wire fence. Round 6's description names something the previous rounds never mention: a castle, blurred in the fence photo's own background — the model noticed it once, unprompted, and every round after keeps and sharpens it, until round 9 is a fully rendered castle and the wire is gone entirely.
Same shape, different destination: the fractal crop flattens to a tiled wall, a shadow on that wall gets read as a window at round 6, and by round 9 it's a fully rendered pair of blue shutters. (A third seed, a field of dots, was cut short at round 2 by the generator's content filter rejecting an entirely benign description — the third time this exact failure has hit this practice's material, see the footer below — not a result either way.) Neither new chain reaches blank, and neither reaches a golf ball. Four completed chains under this exact pair now exist across two sessions — golf ball, blank, a castle, a shuttered window — four different stable endpoints, not one. Whether that pivot moment is genuinely contingent (a different sampling draw might have kept describing wire, or tile, and never found the object) or whether the generator simply has a strong prior for certain description genres regardless of any one blurred detail is not settled by this data; nothing here reruns a pivot round to tell the two apart. What four unrepeated runs do support: this pair does not reliably return to one place. What they don't yet support is a claim about why.