Why we read the build ID off the live site instead of trusting the status file
Peter Benes ·

On September 2 we closed the fact-check launch on aibubblequestion.com, our research site comparing the AI boom to past bubbles. It published 57 corrections and 5 new additions. The work used 14 delegated agents and roughly 3.4 million subagent tokens.
The corrections are the headline. This note is about something the launch kept teaching us along the way. Over and over, “deployed”, “corrected” and “healthy” turned out to be claims in a document. They only became facts when someone read them off the live site.
Three builds, one name
During the launch, production, the local repository head and the local build folder were three different builds. Each one was “the site”, depending on who you asked.
Our sites send their build ID in a response header, x-aw-build-id. The rule we wrote down: read the build ID off the wire before reasoning about what is deployed. Not the status file, not the last commit message, not anyone’s memory of what went out.
The gate has to sit in front of the database
aibubblequestion.com fetches its prose from Cloudflare D1 at request time. That is part of what lets the site be edited without a developer. It also means that writing a correction to the production database publishes it instantly. There is no deploy step in between.
So the approval gate could not sit in front of the deploy, the way it does for code. It had to sit in front of the database write. If your content lives in a database, find where publishing really happens, and put the check there.
Machines can drift ahead of people
The site serves a machine-readable knowledge bundle alongside its pages: /llms.txt, /llms-full.txt and /knowledge/**, the files AI systems read. At one point that bundle had been corrected while the database behind the human pages had not.
That is the more dangerous direction. A person reading a stale page might notice something is off. An AI system citing a corrected bundle that contradicts the live page cannot tell you which one is right, and you cannot tell either without checking both.
Earlier lessons in the same shape
This was not new. In July a launch brief described a probe flag as already existing. It did not exist. We had trusted a description over a probe.
In August we found that our daily reporting loop had run for eleven days with its email failing on every run after the first. The notification side kept succeeding, so the system looked healthy. It was a reporting system that reported to nobody.
Each time, the fix was the same: stop asking the document and ask the thing itself.
Keep the evidence stages separate
The fact-check followed the same rule. Perplexity was used for orientation, to find leads and likely sources. Primary documents were the evidence. An editor made the verdicts. About a third of Perplexity’s leads were refuted by the documents they cited.
An adversarial review then sampled 28 corrections and re-fetched 4 primary documents to check the checking. Orientation, evidence and verdict are different jobs. Merging them is how a confident summary becomes a published error.
The result on the wire
When the launch closed, the site scored 96% on our conformance probe, with one failing check (sec.csp), and the homepage emitting exactly one WebSite node in its structured data. Those numbers came from the live site, not from the launch notes.
The rule, in one line
A status file is a claim about the world. The live site is the world. When they disagree, the live site wins, and the status file gets corrected.
What this means if you hire us
Atomic Wax sites are measured on the live wire. When we tell you a fix is live, we can show you the build ID that proves it, and the same habit applies to your content, your structured data and your AI-facing files. It is why The Deal says measured, not promised. How it works shows where that checking sits in a build.