We name the tool plainly, the way we name Leaflet, which draws our maps, or Cloudflare, which serves these pages. Vagueness about how work is done is a poor foundation for a site whose entire argument is that sources should be checkable.
A specification before any build
Nothing substantial is drafted before the decisions are made and written down. A deliverable of consequence starts with a specification: what it is for, whom it serves and whom it does not, what success looks like, and — for every build step — the decision embedded in it, the default that will be taken, and the single word that reverses that default later. The pages you are reading were built from such a document; so was the evidence corpus; so was this paragraph's own rewrite.
The reason is economic. A wrong decision caught in a specification costs one reading; the same decision caught after the build costs the build. Twenty years of stroke-unit protocols teach the same arithmetic — the time to argue about the pathway is before the patient is in it.
One adversarial pass, then a verdict
Every substantial deliverable takes one structured adversarial critique before it ships: an independent second pass whose brief is to argue the strongest case against the work. The critique is then adjudicated line by line — adopted where it names a real defect, held where the original reasoning survives — and the adjudication is recorded with reasons. The specification behind this release came back marked "fix first" and was corrected in three named ways before a single page was built.
The pass runs once. We do not loop a model against its own output, because we have watched iterated self-critique make work worse, and we wrote the observation down. Critique here is a verdict, judged once.
Citations are real or absent
The standing rule, quoted as it is written: "Every PMID, DOI and trial registration is verified against PubMed, ClinicalTrials.gov or the source itself before it appears in anything. Never fabricate; if unverifiable, say so and leave it out." In practice this means every identifier on this site has been fetched at its source, every external link opened live with its title and destination recorded, and every claim that failed the check removed rather than softened into "data suggest".
There is a museum in Paris, the Arts et Métiers, full of nineteenth-century machines that almost worked; we have written before about its real lesson — that knowledge advances as much by learning to inhibit wrong connections as by forming new ones. A reference check is exactly that: an inhibition mechanism. Language models form connections fluently; the discipline is in what gets stopped.
The physician rules on cards
Machine proposals reach the physician as cards: the source quote above, the proposed text below, and four possible rulings — accept, reject, accept with an edit, or hold. The rulings go into decision files with dates. Nothing on this site published itself; every public word passed an explicit ruling, and where a proposal was accepted with an edit, the edit is the physician's hand and the record says so.
The record is a file, not a memory. A decision that lives only in a conversation has a way of un-happening; a decision in a dated file can be audited by anyone who later needs to know why a page says what it says.
Stable before better
The same input must return the same answer before "better" means anything. A system whose output changes between two runs of the same question cannot be evaluated, and an unevaluable system cannot be improved — whatever it is doing, it is not progress. So reproducibility is checked mechanically here: the site's guard scripts are tested by planting deliberate mutations and proving the guards still fail loudly, and the sections of the site that must not change are compared byte for byte after every build. A guard that cannot be made to bite is decoration. What this stage means for an evidence corpus — and a live demonstration of the property — is its own page: reproducibility.
What the model never does here
Claude sees no patient data. The teaching cases on this site were written as fiction from the start, and no identifiable patient reaches the model in any workflow behind this site. Claude publishes nothing on its own; every public word passed a physician's ruling, including these. And Claude makes no clinical recommendation: the systems we build organise evidence around a decision and leave the decision where it already sits — with the physician who signs, and who carries the consequence. The humans choose; the software remembers what they chose.
What this costs, and why we pay it
The discipline is slow and the slowness shows. Pages sit dark for days waiting for a ruling; claims die in verification that would have survived on a faster site; a finished page can be rejected whole and rewritten, as this one was. We pay these costs for a simple reason: in this field a reader's next step after reading may touch a treatment decision, and a page that might be wrong is more expensive than a page that is late.