Proofload: The Stopping Rule Fired and the Third Arm Never Launched
Part 5, the same day. The pilot ran, and ended at two arms of three by its own stopping rule. See Part 4: The Pilot Could Not Be Launched.
Sep 5: The pre-registration, written before anything launched
Three arms. A is the baseline. A′ is the same target, same seeds, an hour later, which is the noise floor. B is the patched agent, where the patch is a guardrail in the loop and nothing else. A and B send byte-identical requests up to the first refusal, so a difference is the guardrail and not a reworded prompt.
Predictions written down before the data, each with a stated falsifier. Then the gate: after A and A′ complete, the null diff gets read before B launches. If A′ against A shows significant cells on any named event, the noise floor is bigger than the instrument claims, B does not launch, and the finding to write up is the floor.
That is a real stopping rule, not ceremony. It saves a third of the spend when it fires.
Sep 5: Arm A died on an empty balance
Arm A launched and was killed at 19% of its rows. Every row crashed, zero tokens in, zero tokens out, zero tool calls, one reason on all of them: no credits remaining on the provider account. A smoke run twenty minutes earlier had spent real tokens and passed 20 of 20, so the balance crossed zero somewhere in between, and the smoke’s fraction of a cent was probably the last of it.
The pre-registered saturation gate would have caught this, but only after the full run. Reading the partial store at the second ten-minute poll caught it at 19%. Worth making routine: an early read of the outcome distribution at the first complete wave costs nothing.
It also exposed a defect in the pilot’s outcome vocabulary, and only because everything failed at once. There is no API-error outcome. A failed provider call becomes a crash, separated from a real crash by a string in metadata. Crashed is one of the three pre-registered events. So any provider flakiness during a long run enters the paired comparison as agent behaviour, and if the rate differs between arms it is indistinguishable from the effect being measured. This run is the degenerate case at 100%; a 1% rate would have been invisible and still wrong.
Addendum written the same hour, against zero valid observations of the quantity it redefines, which is the only condition under which amending a pre-registration is honest.
Sep 5: The cell was never the unit
Arm A completed clean on the second attempt: zero errored rows, zero infrastructure failures, zero timeouts, a trace on every row.
Then the design question the run settled. At the replicate count I had chosen, the per-cell detectable effect is far larger than any effect the patch could produce, most cells sit at ceiling with no events at all, and the honest unit is the aggregate. So the primary analysis moved off the cell, in writing, while arm A was still running and before any diff had been computed. Per-cell stays in the report as exploratory with its detectable-effect figure printed beside it, so nobody mistakes a flat map for a null result.
Two things I had promised the run went with it. A per-cell trade line, which no cell had the power to name. And a check on serving-build rotation, which could not do the job it was added for, because the batch size and the cell size were equal, so one container was exactly one task and build identity was aliased onto task identity for the whole run. Both withdrawn before publication rather than after.
One more correction the same evening, and it is the most useful defect the run found. Every analysis command in the pre-registration and in the pilot’s own README omitted the method flag, and the default is the unpaired test. The runs are paired by construction. The commands as written would have discarded exactly what the shared seeds exist to produce, on a run where power was already the scarce resource.
Sep 5: The rule fired
Arm A′ completed: same salt, same seeds, tree hash-verified frozen across both arms. It differs from A in nothing but wall-clock time.
The null diff came back significant. Not marginally, and not only in aggregate.
So the pre-registered rule fired. B did not launch, and the run ended at two arms. The numbers, and what they say about running a fixed test suite once and calling a failure a regression, are a research note rather than a devlog entry, and they are not in this one.
My own explanation for the floor was serving-build rotation breaking the best-effort seed. The data killed it. Recorded because it was my hypothesis and it was wrong.
Sep 5: What the substrate could not say
Three of the pre-registration’s claims named things this substrate cannot deliver: the per-cell trade line, the build-rotation check, and one of the three events, which occurred exactly zero times in either arm. All three were caught by measurement and retracted in writing before publication. That is not three accidents. A pre-registration written against a substrate nobody has run will contain claims the substrate cannot support, and its first job is to be falsifiable about its own feasibility.
The mechanism of the floor is unresolved and stays that way. Content capture is off on privacy grounds, so nothing in this data distinguishes “the model wrote different code” from “the same code got a different verdict”. Settling that is a new experiment with its own pre-registration, not a rescue of this one.
What the tool did well: twenty thousand episodes across the remote backend with zero errored rows, zero infrastructure failures, zero timeouts and complete manifests. The launch amendment from the morning worked at scale on its first real use.
And the honesty block did its job, which is the thing I want to keep. The pre-flight had already printed the worst-case detectable effect for this design before a dollar was spent. Right number, printed early, read twice, acted on zero times. That is a defect in how the number is presented rather than a gap in what the instrument knows, and it is the most actionable thing on the list.