Essays • • 8 min read

Two Readers and a Red Box

A vision model read the top of my notebook page and filed it under a date eight years off, calmly. One reader cannot flag what it is sure of. So now two models read every page, and I look only where they disagree.

Two Readers and a Red Box

For eight months one of my notebook pages sat in the repository under the wrong date. The file said Monday, July 23rd 2018. The page says January 3rd 2026, in my own hand, across the top line. A vision model had read that header once, in a batch, and written down a date eight years off without a flicker. Nothing flagged it. The text underneath read like English, so nobody went back.

In June I wrote about reading my own handwriting with a model that can see. Two sentences in that piece were true and have since come apart. I said a confident wrong word costs far more than a flagged uncertain one. And I said that I read every word against the notebook, because there was not that much of it.

There are seventy-five entries waiting for me now, about thirty-seven thousand words. Reading every one of them against paper is no longer a plan. So the question became a narrower one: which words do I have to look at?

One reader cannot flag what it is sure of

The first version asked the model to mark anything it was unsure of. It did, and those marks were useful. They were also the wrong set.

A reader's doubt marks the places it knows it struggled. The expensive mistakes are the other kind, where it had no doubt at all. The 2018 date was one. In another entry I wrote that I was not sure where I stood between Nassim Taleb and Ray Kurzweil, and the transcription had me standing between "Martin Talbot and key Vanguard." On the January page I wrote of seeing the planet "not as stationary being but living vessel," and the old read gave "laboratory being." Every one of those is fluent. None carried a flag. A single reader has no way to be suspicious of its own smooth sentences.

Two readers

So each entry is read twice now, by models from two different families, Claude and Grok. They get the same pages and the same instructions: word for word, never smooth it, and a short list of words I actually use that a general model will not guess. Neither sees the other's answer.

Then a merge lays the two reads side by side and puts a mark on every word where they differ. I am not taking a vote. Two models agreeing does not make a word right. What I am buying is cheaper than truth: when they disagree, at least one of them is wrong, and that is a place worth my eyes.

The January page is 816 words and the two reads disagreed in 39 places. One read a line of options shorthand as "$50 lower than the 500." The other read "BTO lower than the STO," which is what I wrote: buy to open, sell to open. A few entries later both of them missed the name of my own system, Hansuru. One wrote Hanover. The other wrote Hamster. Both wrong, and the flag fired anyway, because they were wrong differently. That is the property the whole design leans on. Two readers rarely make the same mistake in the same place.

The path of one notebook entry. Scanned pages go to two models that read independently. A merge flags every word where the two reads differ. The first pass is frozen, each flag is boxed on the scan, and a person reviews only the flagged words against the page. The reviewed text is scored against the frozen first pass, and the misreads feed a lexicon that both models are given on the next entry. Privacy review and redaction come before the post is built.

Click the diagram to open it full size.

The red box

A mark in the text was not enough. Knowing that word forty-one is in doubt does not tell me where word forty-one sits on a page of cursive, and finding it was the slow part.

So there is a third pass. For each flagged spot the vision model is asked one question: where on this page is that word? The answer is a box, and the review screen draws it in red on the scan.

The first attempt sent whole pages and the boxes landed a word or a line off. The cause was resolution. A full page is scaled down to about 1,560 pixels before the model sees it, and at that size my handwriting does not leave enough to point at. Cutting each page into four overlapping quarters, sent at close to full size, put the boxes on the words.

The review screen. On the left, a scanned notebook page with one handwritten word outlined in red. On the right, the sentence that word sits in, the two models' readings of it one above the other, and a field to type the right word.

One flag: the word boxed on the scan, the two readings beside it.

The screen is plain. The scan is on the left with the word boxed. The sentence is on the right with the two readings under it. One key takes the first, another takes the second, or I type what the page says. Then the next flag.

Scoring the first pass

Before I touch an entry, its merged text is frozen. After I have reviewed it, the frozen copy is compared with mine, and two numbers come out.

Word accuracy is the obvious one: how much of the first pass survived. The second is the one I built the thing for. Of the words I changed, how many already had a flag on them? I call it flag recall, and it answers the question from the top of this page. If nearly every correction I make is on a flagged word, then the flags are where the errors are, and I can stop reading the rest.

Thirty entries have been through it. 15,520 words, 823 flags, 74 corrections. Seventy-three of the 74 were on words that were already flagged. The typical entry came through 99.4 percent right, and the worst one 97.7.

The price is in the other direction. About nine flags in ten mark a word that was fine. I pay for that with a keypress each, and I would pay more.

The one correction nothing had flagged was a date. I had written January 18th. Both models read 16th. They were wrong the same way, in the same place, on a digit, which is the case this design cannot see. It is also the same kind of error that filed a page under 2018. So I am not yet reviewing by flags alone. I still read the date line of every entry against the page, and I will keep reading whole entries until that one number has held across a few hundred more corrections.

What broke this week

One page opened with the last lines of the previous day and had its date header halfway down. One of the two models decided to tidy that: it moved the dated part to the top and put the leftover lines at the end. The merge compared the wrong halves. Agreement between the reads fell to 29 percent, and worse, the body of the entry came out with almost no flags, because the comparison had lined up unrelated text.

Agreement normally runs between 92 and 97 percent. So that number is now an alarm. Anything far under it means the two reads are not describing the same page in the same order, and the entry gets looked at before its flags are trusted.

The lesson is older than the tooling. A check is only as good as the comparison underneath it, and a comparison can fail quietly while still returning a number.

What stays mine

The models do the seeing. Everything after that is a record I can be held to.

The words are published as written. The only edits allowed are punctuation and paragraph breaks, and each one is written down with its reason. Names and places come out before anything is published, as a bar on the scan and a matching mark in the text. And the post is rebuilt from the reviewed record each time, so a change made by hand to a published page shows up as a difference instead of passing as the original.

Every misread I correct goes into a short list of my own words, and the next entry's two readers are handed that list before they start. Hansuru is on it now.

Tesseract could not read a word of my handwriting. One model could read nearly all of it and was sometimes sure and wrong. Two models are still sometimes wrong. The difference is that they now show me where to look, and I have a number that says how far to trust that.