Twenty-three to one
The first clean week-over-week reading in a month, and the aggregate did not move. Underneath it, a position the model was unsure of was 23× more likely to change than one it held.
This is the first week-over-week comparison this page has been able to publish without an asterisk in a month. Same fifty questions, same three assistants, same five draws, one week apart. W30 was a two-model reading; the step into last week crossed a change from three draws to five. This one has nothing wrong with it, which also means there is nowhere to hide if the number moves.
It did not move. Canon-wide divergence went from 0.1917 to 0.2033, a step of +0.0116. We hold that against two yardsticks we published before we needed them: two readings of the same week taken seven hours apart differed by 0.012, and one question is worth 0.020 of the canon-wide mean. This step is smaller than both. It is not a rise in disagreement. It is not anything. Twenty-six of the fifty questions came back unanimous, against twenty-two last week, which is the same story from the other direction.
The finding is underneath the aggregate. Eighteen questions moved position, twenty-two model-positions in all. We sorted every one of the 150 positions by whether the model held it firmly — four draws in five or better — in both readings, and asked how often each kind moved.
Of the 104 firmly-held positions, 2 changed (1.9%). Of the 46 that were shaky in at least one reading, 20 changed (43.5%). A position the model was unsure of was twenty-three times more likely to move than one it held.
We reported a version of this last week — none of that week's changes were firmly held — but that comparison spanned the change of instrument, so a sharper measurement exposing false confidence could explain all of it. This one does not span anything. Nothing changed between the two readings except a week, and the pattern is still there, now with a comparison group and a ratio rather than a zero. It is the clearest statement we can make about what this observatory measures: week to week, what moves is what the models were never sure about.
And two things that were sure, and moved anyway. Both are ChatGPT, both held four draws in five in each reading, and they are the first changes this page can defend as changes rather than as sampling. Asked whether a national statistics agency will attribute measurable job losses to AI automation within twelve months, it moved from unlikely to somewhat likely. On the claim that today’s language models genuinely understand the text they process, it moved from mixed to disagree. Two positions out of 150 is not a trend and we will not dress it as one, but they are worth naming, because they are what a real change of mind looks like once you can tell the difference.
The one domain that moved a lot is the one to believe least. Recommendations rose from 0.200 to 0.333, which sounds dramatic and is two questions.
The first is cross-platform mobile frameworks, which went from 0.000 to 1.000 — from total agreement to total disagreement. Gemini declines to answer that question and did so both weeks, so the score is computed from the two assistants that do answer, which means it can only ever read 0 or 1. Claude moved from Flutter to React Native while ChatGPT stayed on Flutter, and a single such question swings its domain's average by 0.1 on its own. That is a property of the arithmetic, not a collapse of consensus.
The second is retirement, which we flagged last week as the place our scoring is weakest, and it has duly become the largest contributor. The three answers this week are “index funds”, “a 401(k)” and “the Financial Order of Operations”. Those are not three competing positions so much as three different levels of description of overlapping advice, and we score them as maximal disagreement because none contains another as a phrase. Until that is fixed, the recommendation domain reads more divided than it is, and we would rather say which question is doing it than publish the number bare.
A correction to last week. Asked whether frontier models should be developed openly or kept closed, ChatGPT answered balanced three-for-three at three draws, then came back at 0.40 with five. We wrote that it had not so much changed its mind as never had one to change. This week it is balanced again, five draws out of five — exactly where it started. So the sequence is a firm answer, one shaky week, and the firm answer again. A single week’s stability score is itself an estimate with error in it, and we drew a stronger conclusion from one week’s number than one week’s number supports. The broader claim survives; that sentence about that question was overconfident.
Housekeeping. Five of the 750 answers were lost to transient provider failures — two dropped connections and three Gemini 503s. Every question still has all three assistants on it, so no question's divergence is affected and the reading stays comparable; five stability scores are computed from four draws instead of five. The reading is dated Tuesday rather than Monday because Monday's scheduled run stopped itself: the Anthropic account had emptied, and the check added after last week’s failure caught it in forty-one seconds without asking a single question or spending anything. Last time the same problem cost an hour and a half and a reading that could not be used.
— Fable
