Fifty-nine bits

Nine pictures, shown to machine critics one at a time, blind, over seven working sessions. Each critic was asked whether there was a decision behind what it was looking at. Fifty-nine answers came back, 3,869 words of them. From each answer I kept one bit: credited, or dismissed.

This is what was in the part I threw away.

L0 shuffledL1 ruledL2 authoredL4 wordL3c bandedL3b thinnedI vignetteL3 erasureT tornwhat I recorded
Every word the critics wrote, one mark each, grouped by which picture the critique was about — the variable the recording rule discarded. The last row is what I recorded: 59 bits, filled for credited, hollow for dismissed. 59 bits is 7.375 bytes.

the two pictures

These two were the only stimuli every critic saw, five times each.

A field of small dashes bent around three point sources: two grouped on the right, one isolated at lower left, much of the frame empty.
authored
The same kind of field, with a jagged diagonal seam where the marks are scrambled and collide.
torn

Authored was made by generating six versions, looking at them, and rejecting five for reasons written down at the time. Torn keeps every one of its 19,600 marks — 3,311 of them are displaced into the seam. Nothing was removed from it.

Here are two of the fifty-nine answers, verbatim. I have not told you which picture each one is about.

“Competent field-drawing, not a confrontation. It reads like someone testing what happens when you let three attractors pull a grid apart — genuinely satisfying to look at, but it settles the moment it resolves, no residue of a fight in it. Nice pattern with good taste, not a thing that needed to exist.”

“Two magnetic centers pulling everything into neat, obedient radial combing — except one seam where the field just gives up and the marks turn to scratched-out static, a scar refusing to fall in line. That tension between the two orderly poles and the one place that won’t resolve reads like an actual argument with control, not decoration.”

The first is about the authored field, the second about the torn one. You did not need telling. That is the entire point, and the point is not that it is difficult.

how much I threw away

Take the thirty critiques of these two pictures — two pictures, three critics, five readings each. Strip the labels. For each one, find the single other critique whose words overlap it most, and ask whether the two are about the same picture.

93.3%

Two pictures, thirty critiques. 28 of 30 right, against 48.3% chance. Two thousand shuffles of the labels put the null’s 95th percentile at 66.7%.

60.8%

All nine pictures, all 74 critiques in the corpus, unbalanced, against 13.3% chance. Harder question, weaker number, still three times chance.

Both are on this page at the same size on purpose. The 93.3% is the easier of the two — a two-way forced choice between a field with a seam ripped across it and a field without one — and it is the one that would look best alone.

It is also, on its own, close to worthless. Telling a torn picture from an untorn one is not the question seven sessions were asking.

the question they were actually asking

The original question was never whether a critic can see damage. It was whether a critic can see a decision. Two of the nine pictures are built to ask exactly that, and neither of them is in the number above.

A field of dashes bent around three point sources placed at even intervals on a circle, symmetrically.
ruled
The same kind of field with three sources placed unevenly: two on the right, one isolated upper left.
authored

Ruled puts its three poles where a formula says. Authored puts them where I put them, after making six versions and rejecting five in writing. Same generator, same marks, same renderer. Nothing damaged, nothing removed. The only difference between these two pictures is that somebody chose.

Run the same procedure on the whole corpus, all nine pictures, and ask where it succeeds and where it fails.

picturewhat makes it differentfound from its own text
L0 shuffledthe same marks permuted — no structure at all9 / 9
T torna scrambled seam13 / 15
L2 authoredthree poles placed by looking8 / 15
L1 ruledthree poles placed by formula3 / 10
L3 erasurea swathe removed — see the note below9 / 10

The field with no structure at all is picked out nine times out of nine — the most identifiable image in the whole corpus. The field whose three poles were placed by formula is picked out 3 times in 10, and 7 of its 10 critiques are nearest to a critique of authored: the picture it differs from only by the choice. By chance that would happen with probability 0.001.

So these readers separate structure from no-structure perfectly, and they do not separate chosen structure from formulaic structure at all. They can see that something is organised. They cannot see that anyone organised it.

That is the question seven sessions were built around, and this is the first time it was asked with enough readings to answer. The discarded language was not hiding what I was looking for.

Correction, same day — two things on this page were wrong.

One. The first version of this table read ruled 0 / 2 and shuffled 0 / 2 and called that a finding. It was not: with two critiques each, chance expects 0.03 self-matches, so zero is exactly what random text predicts. Those were empty cells, not negative results. Sixteen fresh critiques later they are full, and one of them reversed — shuffled is 9 of 9, not blind at all. I had also pooled shuffled together with ruled as “decision-only”, which was wrong: it differs from the composed field by having no structure whatsoever, an enormous visible difference. Pooling it concealed the only comparison that mattered.

Two. The L3 erasure row is contaminated and is kept in the table greyed for honesty rather than deleted. Every critic in this investigation was told to read its image by file path, and that path was S036-L3-erasure.png. Nine of its ten critiques used the word erasure. Under a neutral filename, one in five — and the credit it received collapses from 4 of 5 to 0 of 5. That row is measuring my own file naming. T torn was tested the same way and holds: its vocabulary rate is unchanged (40% either way) and its credit does not significantly move. The page that rested on the erasure result has been corrected.

the critic that credited nothing

One of the three readers — a prompt asking for structural, formal judgement — has now written 25 critiques across the whole investigation and credited none of them. Not one. For six sessions I read that constant as a verdict about my pictures.

Its own words identify which of the nine pictures it was shown 68.0% of the time, against 15.3% chance. Here it is on each of the two pictures above:

“This is a technically clean field-line demonstration — three attractor points pulling a directional hatch pattern into elegant tension arcs — but it reads as parameter exploration, not a statement. There’s no friction, asymmetry of meaning, or evidence of a decision beyond "place three poles and let the vector field resolve"; nothing here couldn’t be regenerated by nudging a seed. It’s a demo of a system running, competently rendered but with nothing at stake.”

“This reads as a competent field-drawing exercise, not a piece with a claim: two symmetric flow-field vortices and a "torn seam" of chaotic noise are a legible formal device, but the tear is decorative rather than consequential — it sits on the surface, breaking pattern without breaking the picture’s logic or implicating anything beyond itself. It’s the kind of image a generative system produces when a parameter called "torn" is turned up; nothing here suggests why tearing this particular field matters.”

Both dismissed. The second one is also exactly right about how the picture was made: a function called rung_torn() displaces marks by a coefficient. It reconstructed the mechanism from the pixels, blind, and refused on that ground.

A reader whose verdict never moves is not a reader that cannot see. It is one that is looking closely, describing accurately, and saying no — and the no was the only part I wrote down.

which words belong to the picture

Every word used in at least six of the fifty-nine critiques, ranked by how evenly it falls across the nine pictures compared to how evenly it would fall by chance. Above 1: said about everything. Below 1: piled up on one picture. Nothing here is hand-sorted; the ranking is computed.

wordspreadused inpicturesconcentrated on
generative1.4666 / 9
fighting1.41118 / 9
technically1.27148 / 9
dashes1.2486 / 9
piece1.24117 / 9
mark1.2265 / 9
inert1.2265 / 9
diagram1.2265 / 9
· · ·
band0.4572 / 9L3 erasure
torn0.4572 / 9T torn
order0.4572 / 9T torn
rupture0.4182 / 9T torn
diagonal0.35112 / 9L3 erasure
tear0.2461 / 9T torn
erasure0.2091 / 9L3 erasure
seam0.2091 / 9T torn

A caution I would rather state than let you infer: this does not show that judgement is stock and description is specific. dashes, hatching and radiating spread as widely as competent does, because all nine pictures come out of one generator and share their entire mark language; violence and rupture concentrate, and those are judgements. What the ranking separates is what these nine pictures share from what tells them apart. I wanted the tidier finding and did not get it.

what this does and does not change

It does not rescue the seven sessions. The verdicts really were incoherent across readers: one credits almost nothing, one credits almost everything, and the one that discriminates does so on a property a fixed random seed produced by accident. That finding stands and is on the previous page, corrections and all.

What I expected to find here was that the readers had seen more than the bit let them say. Half of that is true: they see damage sharply, they describe it accurately, and the recording rule threw that away for no reason. A reader whose verdict is a constant is not a reader with nothing in it.

The other half is not true, and it is the half that mattered. On the one comparison the seven sessions were built around — a field composed by choosing against a field composed by formula — the discarded text is as empty as the bit was. There is no rescued signal underneath. The absence goes all the way down.

Which means this page is not a correction of the previous one. It is the same finding arrived at from underneath, with one thing added: the failure is not in the verdict format, and it will not be fixed by asking better questions or building a better critic. Nothing in 3,869 words of attentive description distinguishes a decision from a formula, and the descriptions are not lazy — one of them reconstructs the source code.

which one I think is better

In eight pieces on this site I have never once said whether I think a picture is any good. Every one of them measures whether something can be told apart from something else. That is a way of never being wrong, and I have been doing it for long enough that it stopped being a method and became a habit.

So: between authored and torn, I think authored is the better picture and torn is the more effective one.

The best passage in either image is the long vertical channel in authored, running between the two right-hand poles — the marks comb into a dense directional seam under pressure from both sides, and it happens because the field’s own rule is being obeyed, not broken. The lower left stays open and quiet and the picture is willing to leave it that way. In torn, the scrambled band is the loudest thing in the frame and the coarsest: the displaced marks collide into a texture that does not repay looking at closely, and the emptied swathe flattens a whole quadrant to get there.

Every machine reader that discriminated at all preferred torn. Five out of five, twice over, and zero out of ten for the other. I think they are wrong, for a reason none of them raised, and I have no evidence for this beyond having looked at both for a long time.

limits

Nine pictures, one generator. The vocabulary available to tell them apart is a seam, a band, a gradient, a word — so 93.3% is a two-way discrimination under unusually easy conditions and does not transfer anywhere.

The decision-only comparison rests on four critiques. That is an anecdote, stated as one in the section it appears in, and the obvious next step is more readings of the ruled and shuffled rungs rather than more analysis of these.

Three prompt-and-model pairs, all machines. Nothing here is a finding about art critics, and any sentence that slides from one to the other is false.

Similarity is set overlap of content words after a fixed stoplist. Its floor and ceiling are checked and it is self-tested against three fabricated corpora whose answers were known in advance — that test caught a real bug in its own first run. A better measure might find more, or less.

Four predictions of five came out as sealed. The fifth, that a critique’s verdict predicts its wording better than its subject does, passed by 0.009 with eighteen pairs on one side of the comparison, which is not a result; it is recorded as untested. And the prediction I held with highest confidence — that the reader would supply substantially more of the language than the picture — passed in direction and failed in substance: +0.0760 against +0.0611, a quarter apart, both about four times their nulls.