Independent review and gates
An independent review puts an artifact in front of a reviewer who did not write it and asks
one question about it. The optional REVIEW-LEDGER.md records each review and
turns the faults reviews find into gates: questions to ask before the next
experiment is written or submitted, each founded on a run that failed that way.
Over time the ledger shows which checks have paid for themselves and how often review finds nothing. The review ledger reference specifies the file's tables and gate format. Reviews complement the human checkpoints in the experiment workflow; they do not replace a person's approval of a design or an interpretation.
Requested reviews
A person asks for a review, or the project's instructions require one at a named transition, such as before a script's first paid submission or before a plan begins execution. An agent does not start reviews on its own initiative, and the ledger does not schedule them. It records the reviews that happened and what they found.
Name the transitions in the project's instructions when a review should never be skipped. An agent session does not see a transition that is not written in the project's instructions.
What to review
| Artifact | When | What the reviewer judges |
|---|---|---|
| Plan | Before execution, and at completion | Whether dependencies, gates, decision rules, and the completion condition are stated and testable. |
| Experiment design | Before code is written against it | Whether the registered design can answer its question, and whether the decision rule computes the quantity the design measures. |
| Experiment script | Before its first costly run, and after any change | Whether the code implements the design, and whether its statistics, artifacts, and provenance are sound. |
| Results | Before a finding or claim relies on them | Whether the numbers and interpretation follow from the retained evidence. |
A design review does not establish that the code implements the design. Neither review establishes that the code runs or that the instrument can move the construct it measures. Those questions need execution: a bounded smoke run, and a pilot with a predeclared positive control.
Reviewer independence
Record how the reviewer relates to the author. A fresh reviewer receives the artifact and the record it implements, and nothing of the session that produced them.
| Relation | Meaning |
|---|---|
self | The author reviewed their own work. |
same model, no shared context | A fresh session of the same model, given only the artifact and its record. |
different model | A reviewer from a different model family or provider. |
human | A person other than the author. |
Consecutive repairs of one artifact benefit from alternating authors as well as reviewers. A reviewer reports an author's blind spot once per round, and the next repair by the same author tends to reproduce it.
Log every review, including the clean ones
The review register gets one row per review. Clean reviews are the
denominator: without them, neither the practice's value nor any single gate's value can be
measured. Many similar reviews may share one
summary row, such as R040–R061: 22 script reviews, 2 findings.
The Target names the exact bytes reviewed, as a path with a content digest or revision. A review of earlier bytes does not cover a later revision, so a script changed after its review needs another one before its costly run.
Findings
A finding is one fault a review reported. Its severity is
silent when the fault would have produced a wrong number with no error, and
loud when it would have failed visibly.
The author's own passing tests show that a repair runs; they do not show that it is right. Record that the repair landed, and close the finding when a review of the repaired bytes settles it. Until then the finding is visible to anyone reading the ledger, and the run it blocks waits.
Gates
A gate is a question with a founding incident. Add one when a run fails, or a review finds a fault, for a reason no existing gate covers. When a later finding is another instance of an existing gate, attach it and increment the gate's tally instead of adding a near duplicate. Resemblance between two faults is not by itself evidence that they are the same; read the gate before attaching to it.
The ledger splits gates into two sections. Gates decidable from a script and its output go first, and a fresh-context reviewer receives only that section. Gates that need the project's history, such as whether a sibling experiment already solved this problem, stay with reviewers who have that history. Never ask a reviewer to work out which gates it can reach.
Consult the gates before a costly submission
Before submitting a new or modified experiment script, read the gate titles and open the detail only for gates that plausibly apply. The file is written to be skimmed in under a minute. After the run, log the review that preceded it, whatever it found.
Edit the ledger through a tool that allocates IDs and serializes concurrent writers when one is available. Without one, append rows by hand and never renumber or delete them.
A script review that founds a gate
Two reviews of a synthetic comparison script
Project instructions require a script review before any first paid submission. On
2026-09-10 the owner requests one for scripts/exp_002_comparison.py at
revision 3f1c9a2e. A fresh session of the author's model receives the script,
the EXP-002 record, and the artifact-checkable gates. It reports two faults. F001: the
summary averages over every question cell, including a pilot cell the design excludes;
the result would be wrong with no error, so it is silent. F002: the retry path writes to an
output directory that already exists and exits with an error, which is loud. No existing
gate covers F001, so it founds G1.
A different author repairs both faults. The script's tests pass, and both findings are marked as repaired but still open. The run waits.
REVIEW-LEDGER.md after the repair lands## Review register
| ID | Date | Kind | Target | Reviewer | Relation | Findings |
| --- | --- | --- | --- | --- | --- | --- |
| R001 | 2026-09-10 | review-script | scripts/exp_002_comparison.py @ 3f1c9a2e | fresh session | same model, no shared context | F001, F002 |
## Open findings
| ID | Fault | Caught by | Severity | Status |
| --- | --- | --- | --- | --- |
| F001 | Summary averages over all cells, including the excluded pilot cell | R001 | silent | repair landed; awaiting review |
| F002 | Retry path fails on an existing output directory | R001 | loud | repair landed; awaiting review |
## Gates — checkable from the artifact
#### G1 — Does the summary aggregate over the intended slice?
**Added** 2026-09-10, at EXP-002. **Severity:** silent. **Tally: 1** — F001 (founding). **Slug:** `aggregation-slice`.
On 2026-09-11 the owner requests a second review of revision 8b04d17c from a
reviewer on another provider. It confirms both repairs and finds nothing new. The register
gains R002 with relation different model and findings
none, and F001 and F002 close on R002's verification. The clean review is
logged like any other. Only then is the script submitted.
Request a script review and log the pass
Replace SCRIPT_PATH and EXPERIMENT_PATH and send this prompt to the
coding agent. It requests a review; it does not authorize a repair or a submission.
Use research-lab-notebook to arrange an independent review of SCRIPT_PATH at its current revision against the design registered in EXPERIMENT_PATH. Give the reviewer only the script, the experiment record, and the gates in REVIEW-LEDGER.md that are checkable from the artifact, and nothing from this session. Ask whether the code implements the registered design and whether its statistics, artifacts, and provenance are sound. Log the pass in the review register with the exact bytes reviewed as a path plus digest or revision, the reviewer, and its relation to the author, even if it finds nothing. Record each fault as an open finding with its severity; attach it to an existing gate only after reading that gate, and found a new gate when none covers it. Do not repair the script, close findings, or submit the run.
Promotion and retirement
A gate with a high tally is a candidate for promotion: into a template, a shared checklist, or a mechanical check that makes the fault impossible to write. Record the promotion in the ledger's promoted section and say what it does not cover; a template reaches only the records written from it.
A gate that stops firing is a candidate for retirement. Its tally, read against the clean reviews in the register, is the evidence for either move.