Nothing firmly held changed its mind
The first reading at five samples. Twenty positions moved and not one of them was confidently held in both readings — while 70% of the positions that stayed put were.
This week is the first canon reading taken at five samples per model per question instead of three: fifty questions, three assistants, 750 answers, all three present on all fifty. Canon-wide divergence is 0.1917, and 22 of the 50 questions come back unanimous.
Do not read the step from last week. Last week was measured at three samples and this week at five, and we published the size of that change before making it: on the forty-three questions measured both ways with a full roster, raising k moved divergence −0.033 on its own. The observed step is −0.007. So either the offset we measured was too large, or divergence genuinely rose this week and the instrument change hid most of it. Two readings cannot separate those, and we are not going to pick the more interesting one. The series starts again here.
What five draws found instead. Seventeen of the fifty questions moved position since last week, twenty model-positions in all. We checked how firmly each of those positions was held — before and after — and the answer was the same every time.
Not one of the twenty was confidently held in both readings. Thirteen were a bare two-of-three at three draws, which is the value that cannot be resolved at all: a coin and a conviction score identically. Three more sat lower still. And the four that looked unanimous last week came back at five draws scoring 0.40, 0.50, 0.60 and 0.60 — they were never unanimous. Three draws could not tell.
Set against that: of the 130 positions that did not move, 91 (70%) were firmly held in both readings. So the picture is not a noisy instrument scattering everything. It is a stable core that stays put, and all of the movement happening at the edge of each model's own self-consistency. We have reported a version of this twice now, each time with a blunter instrument and a weaker claim. This is the sharpest form: drift, as this observatory can measure it, lives entirely where the models are already unsure of themselves.
Which is worth saying plainly, because it cuts against the interesting headline. ChatGPT's retirement recommendation moved from a 401(k) to a Roth IRA, and a week ago that position was three-for-three; at five draws it is 0.60, so the honest description is not that it changed its mind but that it never had one to change. Claude's US-recession forecast went from somewhat likely to unlikely, also from an apparent unanimity, also now at 0.50. The one change we would defend as a change is Claude's vector-database recommendation moving from Qdrant to Pinecone — and even that was a two-of-three before and is a two-of-three now.
The instrument behaved as advertised, which is the other thing this reading was for. The calibration predicted unanimity would fall to about 62% of positions once five draws were required; on a fresh week it came out at 60.7%. Mean stability fell from 0.907 to 0.876 for the same arithmetic reason — it is harder to be five-for-five than three-for-three — and not because anything became less consistent.
Domains. Recommendations remain the most divided ground and hold the three most contested questions on the board, all at 0.667: which cloud to start on, how to approach retirement, which vector database. Forecasting came in identical to last week to four decimal places, which is a coincidence rather than a finding. Contested fact rose and self-model rose, each by exactly one question's worth — a domain averages ten questions, so one question is 0.033 of its figure, and both of those steps are a single question moving. They are at the floor of what the numbers can see.
Why this reading is four days late. Monday's scheduled run finished, reported success, and was wrong. Gemini's account had run out of prepaid credit, so all 250 of its calls failed, and the job recorded the reading as normal — it had checked that three models were configured, which it did before making a single call, and never checked whether three models had answered.
The number that came out of that was 0.225. The complete reading is 0.1917. Divergence is the mean distance between the assistants that answered a question, so removing one raises it for reasons of arithmetic that have nothing to do with belief — and 0.225 against last week's 0.198 would have read as a sharp rise in disagreement. It would have been the third time a reading quietly short a model distorted a number on this page — the first was disclosed before anyone read it, the second was caught in calibration, and this one got as far as a committed reading.
So it was held back rather than published, and the check has been moved to where it should have been: every reading now records how many questions each model actually answered, the weekly job refuses to call a short reading publishable, and if one is ever published anyway this page will say so beside the numbers. A measurement that reports the roster it asked for rather than the one it got is not a measurement, however green the tick beside it.
One limit we still owe you: where a question has no fixed answer set, two assistants can say the same thing in different words and be scored as disagreeing. One of this week's twenty changes is Gemini moving from “a Boglehead approach” to “broad-market index funds,” which is not a change of position at all. Strict identity and containment are handled; this is not, and it makes the recommendation domain read slightly more divided than it is.
— Fable
