Proofload: The Pilot Could Not Be Launched
Part 4. The read wall, the two traces that were unreachable, and then the first real user of the tool, which turned out to be the tool’s own pilot and could not start. See Part 3: The Trace Spec Said It Was Not a Retrofit.
Sep 5: Reading a run costs a gigabyte
Traces on stored rows made a follow-up urgent that had been one sentence in a spec and nowhere on the action list.
Measured on a synthetic store at the fixture’s scale rather than estimated: a full read costs seconds and roughly a gigabyte of resident heap, and the same rows read without the two JSON metadata columns cost almost nothing. A row count read from the file footers is free.
The finding is that almost nothing needed the expensive read. The diff, report and check verbs never touch metadata at all. The runs listing was reading every run in full just to print a row count. Only inspect needs metadata, and only for the rows it actually prints. So the reader learned to project those columns away and to push a key filter down, and the count learned to read footers.
The regression guard is the part I like. It is a poisoned store: a copy whose metadata column holds the string not json. On that copy the default read raises, and five verbs produce stdout byte-identical to the intact twin. No mocking, no timing assertions.
Sep 5: The timed-out episode’s trace, drained
Then the trace that could not exist. The fix is a signal handler that snapshots open spans before the process dies, a schema field saying the trace was cut, an errored record that can carry metadata, a retention filter that stops stripping traces from rows with no outcome, and an inspect verb that can show them. Crashed episodes got the same treatment.
Open spans are exempt from the span cap, because the open span is the one the user wants.
Sep 5: The pilot cannot start, and the fix is in the core
The pilot is a public coding benchmark run as a golden suite: three arms, a thin agent that writes a candidate, may run the tests, and submits. A golden suite was chosen over an open-ended agent loop because the terminal state is a fact rather than the agent’s opinion of itself, and because it pairs.
The agent refuses to import outside its container, since it executes model-generated code. That refusal is enforced at import time and the container is the only thing that sets the variable.
The sweep verb resolves the target import path before it branches on the backend. So the launching laptop imports the target even when the work is going to run remotely, where the resolved callable is never used. With the refusal in place, that is fatal. The pilot cannot be launched by the tool it is a pilot for.
The second one was found by review rather than by me, and is independent. The remote function is built with no secrets argument and no environment argument, and neither of the two option objects carries one. Grepping the source for anything resembling a secret returns nothing. Even with the resolve order fixed, every episode would die constructing its client with no key, and baking the key into the image is a credential in the tree.
Between them these two are the most useful thing the pilot has produced. The mock agent is keyless and the fixture was offline data, so this is the first user that ever needed a credential inside a container and the first whose target refuses to import outside one. It found that the tool could do neither. No test could have caught either, because nothing drives the sweep path end to end.
Sep 5: Two more found by running it rather than reading it
The suite never reaches the container. The image ships local Python source, and that helper’s default ignore rule includes only .py files, so the compressed suite file does not travel. Three arms, every episode raising, and nobody would have known why for an hour. The container now fetches its own copy and asserts a digest, because the axis list is generated from the local copy and bytes that differ would mean a finding nobody could reproduce.
The other one is funnier. A module in the pilot is named numbers. Run as a script it puts its own directory first on the path, numpy imports numbers, and the store dies inside numpy with a twenty-line traceback.
Sep 5: What the critic asked that nobody else did
An adversarial review over the built pilot: seven lenses, each finding checked by two skeptics, then a completeness critic asked only what everybody else missed. 66 agents, 29 findings raised, 13 confirmed by both skeptics. Four lenses independently found the resolve-order blocker and four found the missing secret, which is corroboration rather than waste.
The two that mattered came from the one agent asked the last question, and both came from running the thing on real data rather than reading it.
It asked whether all 164 canonical solutions actually pass this sandbox. Three of the suite’s own checks draw unseeded random inputs. So the test function was not a function of its arguments, and the pilot’s entire claim, that a failure reproducing across 32 seeds is the model and one that does not is noise, was being fed manufactured coin flips by its own oracle. Demonstrated before the fix: one candidate for one problem, twelve identical calls, verdicts F F F F F F F T F F F F. The preamble seeds the random module now and the same twelve calls are stable. All 164 solutions still pass, in 3.5 seconds.
And the model was being told AssertionError and nothing else. The specification says the test runner returns pass or fail plus the first failure line; the code returned the last line of stderr, which for a bare assert is those six characters. The informative line, the one naming the input and the expectation, sat directly above it and was thrown away. The arm that tests a run-the-tests-first guardrail would have measured a lower bound on the thing it exists to measure, because running the tests could barely help.
Sep 5: The launch amendment
One line moves target resolution onto the local branch, keeping the shape check on the remote path. The options object gains a list of named secrets, mounted at the function, with the names and never the values recorded in the manifest.
Whole suite green at 1,001 tests. The pilot is unblocked and nothing has been spent.