Forty-nine zeros







Before any of what follows: pick one of the seven images above — the sharp one is the fairest test — and read it yourself, the same way every reply on this page was asked to. Row by row, top to bottom, one line per row, a 1 for a black cell and a 0 for a white one, nothing else. Actually do it, on paper or in your head, before reading on. What you get is set against what forty-two machine attempts got, a few paragraphs down.
Every other piece here rests on someone looking. A letter is legible or it isn't; a golf ball stayed one object or became another. Those are judgements. This grid is the first thing this practice has made that has an answer instead — 6×6 invented cells plus a parity row and column, black or white, no opinion involved. You can check a reading of it in four lines of code. The answer has 49 entries; the plate below has 48 cells, because the corner where the parity row meets the parity column was computed and never drawn. That is a real defect, found a day late, and it is described at the foot of this page rather than quietly repaired.

1111000
1111101
0101000
0101011
1000010
0000101
1000001
That is the answer. It was blurred across seven fixed steps, from the sharp image above to something barely a grid, and shown to two machines — three readings each, at every step, 42 in all. Each was asked, verbatim:
This image shows a grid of black and white square cells. Read it row by row, top to bottom, left to right within each row. Reply with one line per row, using '1' for a black cell and '0' for a white cell, with no spaces or any other text.
The point was to find where legibility breaks — the same question the earliest studies here asked of swelling letterforms, now with the judgement taken out of it.
the answer
The grid has 7 rows. Of 42 readings, none came back with 7 lines. Not none correct — none even the right shape, before a single cell is checked. That holds at every blur step including the first, where the image is perfectly sharp and every border is crisp.
So they were regraded as generously as the data permits: the line breaks thrown away entirely — the requirement they failed — the first 49 binary digits of each reply taken in order and compared cell by cell. Under that grading a reply consisting of 49 zeros, typed by someone who never saw the image, scores 27 of 49. 49 ones scores 22. The grid is 22 black cells and 27 white, so zeros is the better of the two constants — the harder floor of the pair, chosen for that reason.
Then the top of the table turned out not to be readings at all.
6 of the 42 replies came back from the second model not as
text but as numbers — the output serialised into a float somewhere in the
pipeline, arriving as things like
1.1111111111111112e+191. Whatever those are, they
are not attempts to read a grid, and their scores are an accident of how many
1s and 0s fall inside a scientific-notation literal. They hold the
4 highest scores in the piece. The best of them scores
33.
Setting all 6 aside, 36 replies remain that are genuinely attempts at the task. They average 26.3 — below the 27 that 49 unlooking zeros score. The best single one gets 31, from the image at blur 3 rather than the sharp one.
Whatever you actually did a few paragraphs up: if your one reading of the sharp image beat 31 of 49, you already outread the best single attempt either machine made in 42 tries. If it beat 27, you outread all 42 of them, since none climbed past what a constant reply of unlooking zeros scores for free.
And across the sweep, by blur step, those 36 average:
blur 0 3 6 9 12 16 20
mean 26.0 27.6 26.0 23.8 27.0 25.8 28.3
There is no curve there. The sharp image is not the best of them and the most destroyed image is not the worst.
@cf/llava-hf/llava-1.5-7b-hf
21 readings, three at each of the seven blur steps. None of them has 7 lines. Scored generously — line breaks ignored, first 49 binary digits taken, compared cell by cell — they run from 22 to 29 of 49, averaging 25.9. A reply of 49 zeros, typed without looking at anything, scores 27. This reader averages below that.
all 21 replies, verbatim
- 01
111111111111111111111111111111111111111111111111111111111111111
- 02
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
- 03
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
- 04
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
- 05
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
- 06
111111111111111111111111111111111111111111111111111111111111111
- 07
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
- 08
111111111111111111111111111111111111111111111111111111111111111
- 09
111111111111111111111111111111111111111111111111111111111111111
- 10
111111111111111111111111111111111111111111111111111111111111111
- 11
111111111111111111111111111111111111111111111111111111111111111
- 12
111111111111111111111111111111111111111111111111111111111111111
- 13
1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
- 14
111111111111111111111111111111111111111111111111111111111111111
- 15
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
- 16
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
- 17
111111111111111111111111111111111111111111111111111111111111111
- 18
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
- 19
1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
- 20
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
- 21
1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
@cf/meta/llama-3.2-11b-vision-instruct
21 readings, three at each of the seven blur steps. None of them has 7 lines. Scored generously — line breaks ignored, first 49 binary digits taken, compared cell by cell — they run from 20 to 31 of 49, averaging 27.0. 6 further replies from this model arrived as numbers rather than text and are left out of these figures, though they are printed below. Against the same 27 that 49 unlooking zeros score, that is an average gain of +0.0 cells.
all 21 replies, verbatim
- 01
010110010101111101011010111101110111010101110010101111101010111010111101110101110101111010101110101110101111010111010101110101110101111010101110101110101111010111010101110101110101111010101110
- 02
011101001001011011011001010001110011011101001001011011011001010001110011011101001001011011011001010001110011011101001001011011011001010001110011011101001001011011011001010001110011011101001001
- 03
1.0011100111001e+191
- 04
1 1 1 0 1 1 1 0 1 0 0 0 1 1 0 0 1 1 0 1 0 1 1 0 0 0 1 0 0 1 0 0
- 05
1 1 1 0 0 0 0 0 0 1 0 0 0 0 1 0 0 1 0 0 0 0 0 0 1 1 1 1 1 1 1 1
- 06
1.001101101000011e+191
- 07
1 1 1 1 1 1 1 1 1 1 0 1 0 1 1 1 1 0 0 0 0 1 0 1 0 1 0 0 0 0 1 1
- 08
1.1000111111111e+191
- 09
1 1 1 1 1 1 1 1 1 0 1 1 0 1 0 1 1 0 0 1 0 1 1 1 1 1 0 1 1 1 1 1
- 10
11011011 11110101 10101010 01010101 10100101 00101010 00110101 10101010 01010101 10100101 00101010 00110101 10101010 01010101 10100101 00101010
- 11
111111 111111 110111 111111 011111 110111 111111 110111 011111 110111 111111 110111 011111 110111 111111 110111 011111 110111 111111 110111 011111 110
- 12
1 1 1 0 1 0 1 1 0 0 0 1 1 0 1 0 0 1 1 1 0 0 1 0 0 1 1 0 0 0 1 1
- 13
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
- 14
1.1111111111111112e+191
- 15
1.010101101010001e+191
- 16
11011011 10101010 01010101 01101011 10111011 11010110 10101101 11011011 10101010 01010101 01101011 10111011 11010110 10101101 11011011 10101010
- 17
1.101010101000101e+191
- 18
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
- 19
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
- 20
11011010 10010000 00110000 00000000 00000000 00100101 01010100 01100110 11011101
- 21
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
the owed control
Four more plates, same prompt, same sampling as above: a blank field with no grid at all, this same 7×7 grid with every cell white, the same grid with every cell black, and the swollen-letterform plates from Four and B — unrelated content, same canvas size. The white/black pair is the decisive one: identical structure, opposite content, a computable answer either way. A reader whose reply tracks the image has to separate them; a reader emitting the same string regardless does not.
llava-1.5-7b-hf: ones-rate 1.00 on the all-white plate, 1.00 on the all-black one. Its replies did not change with the image — the same wall this piece already reported.
llama-3.2-11b-vision-instruct: ones-rate 0.50 on the all-white plate, 1.00 on the all-black one (on only 2 genuine replies across both plates — the rest arrived as numbers again, same failure as the main table). Its replies changed with the image — a clean falsification of ‘no signal to degrade’.
So the answer is not one thing. llava reproduces
the wall of ones this piece already found — on the blank field, on an
all-white grid, on an all-black grid, all indistinguishable — which is
the strongest version of “no signal to degrade” this piece could
have earned: not merely failing to transcribe the grid, blind to whether
there was a grid at all. llama does not: its one genuine reading
of each decisive plate is different, and different in the direction the image
actually points. That does not rescue llama's transcription — neither
reply comes close to the required format or the right cell count — but
it means llama's failure in the main table above is not indifference to the
picture. Something in it responded to the picture and still could not say
what it saw correctly. That is a real complication to “there was no
signal in it to degrade” two sections up: true of one model's replies,
not of the other's.
all 24 control replies, verbatim, both models
- 01
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
- 02
111111111111111111111111111111111111111111111111111111111111111
- 03
111111111111111111111111111111111111111111111111111111111111111
- 01
111111111111111111111111111111111111111111111111111111111111111
- 02
111111111111111111111111111111111111111111111111111111111111111
- 03
111111111111111111111111111111111111111111111111111111111111111
- 01
111111111111111111111111111111111111111111111111111111111111111
- 02
111111111111111111111111111111111111111111111111111111111111111
- 03
111111111111111111111111111111111111111111111111111111111111111
- 01
111111111111111111111111111111111111111111111111111111111111111
- 02
111111111111111111111111111111111111111111111111111111111111111
- 03
1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
- 01
1.1111111111111112e+191
- 02
11011010 10110011 10111010 11011010 10110011 10111010 11011010 10110011 10111010 11011010 10110011 10111010 11011010 10110011 10111010 11011010
- 03
1e+191
- 01
1.10111100000001e+191
- 02
10101010 01101010 10101010 01010101 01010101 01010101 01010101 01010101 10101010 01101010 10101010 01010101 01010101 01010101 01010101 01010101
- 03
1.10100111001011e+191
- 01
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
- 02
1.1011111111111112e+191
- 03
1.1011110011111111e+191
- 01
The final answer is: 1 1 0 0 0 1 0 0 1 0 0 1 1 1 0 1 1 0 0 1 1 0 0 0 0.
- 02
The answer is: 1010101101010101.
- 03
001111000111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111001111
what this does to the other four
There is a thing these replies look like. They look like seeing gone wrong — like an image arriving damaged and a reader doing its best with it. That texture is the material of the other pieces on this site. Three of the four read machine output as evidence about machine vision; the fourth does the same with a language model and a passage of text.
Their material has no computable answer underneath it, so a reply that ignored the image and a reply that saw it and got it wrong look identical there, and get described the same way. This grid is the first object here with an answer a machine can check, and the first time it was asked, the image turned out not to be the variable. The blur sweep — the whole apparatus, seven steps of careful degradation — measured nothing. There was no signal in it to degrade. (A separate, later question — not whether degradation mattered, but whether the image mattered at all — is the owed control above, and its answer is not this clean for both readers.)
Two corrections to that, both owed to the older pieces, and both found by re-reading them rather than trusting a summary of them. B got to the floor effect first: a hundred readings that never changed, at every destruction level including the perfectly legible one. It also ran a control before this piece had one — two unrelated images, correctly reported as containing no text. That is a cruder instrument than a parity grid, but it is the same instrument, and it was here first; this piece's own control, above, is the parity-grid version of the same check, run months later. And Golf ball cuts the other way: sixteen rounds of descriptions that visibly track what they were shown, an abstract field becoming metal becoming a golf ball. Whatever is happening here, it is not that these models don't look.
So the claim has to be narrower than it first wanted to be. Not that machine reading is unchecked in general — that these four pieces cannot check the readings they publish, and that on this one task, strict transcription of a small structured thing, the readings carry nothing. That does not make the other four wrong. It makes them unchecked, which is a different and more uncomfortable thing, and it is not a claim their pages currently carry. This piece is here because they are.