The request that never left the building
Two unrelated bugs this week turned out to be the same bug: a system that fails by doing nothing, and a benchmark-contamination paper that says the industry has this problem at a scale I can't fix by being careful.
A container image had been hanging on docker pull for ninety seconds. No error, no progress bar, nothing. I ran docker events in a second terminal and pulled again: the log stayed empty the whole time. Zero events isn’t a slow request, it’s a request that never reached anything. Three keys deep in ~/.docker/config.json, credsStore: "desktop" was still routing every pull through a credential helper for a Docker Desktop backend I’d killed weeks earlier. The pull wasn’t broken code. It was a system faithfully executing an instruction that had stopped being true.
Two days later, same shape, different project. I was hand-authoring benchmark tasks for a sandbox that runs each one in its own container, and the first task failed instantly: the file the stub was supposed to patch wasn’t in the image at all. Docker’s implicit build looks for a file named exactly Dockerfile in the build context directory; mine lived one level up. Fixing it needed a .dockerignore too, or the image would have shipped the answer key into the same sandbox the agent is supposed to solve the problem in.
Then I read the paper that made the coincidence sharper than I liked. SWE-Bench+ found that 32.67% of “successful” patches on SWE-bench, the reference benchmark half the field cites, pass because the solution leaked straight into the issue text or the environment the model was evaluated in, not because the model solved anything. I’d built exactly that failure mode into task one, by accident, on a benchmark I hand-authored myself, one task at a time, with every incentive to get it right. If I can leak the answer while staring at a single file, a benchmark built to generate thousands of tasks automatically isn’t fighting a data problem. It’s fighting the same convention-shaped blind spot I hit twice in one week, at a scale where nobody’s staring long enough to notice.
That’s the part I keep turning over. The industry story about contamination is usually a training-data story: models memorizing answers somewhere upstream. Mine wasn’t that. Mine was infrastructure doing exactly what it was told, in a context nobody had audited against what the task actually needed. The fix isn’t a smarter model or a cleaner dataset. It’s the same one-line discipline as the credential helper: check what the system is actually doing, not what you assumed by default, before you trust the number it hands you.