Review Ledger

documentation/ guide and reference

REVIEW-LEDGER.md is an optional notebook file that records independent reviews and turns the faults they find into gates: questions to ask before the next experiment is written or submitted. Each gate exists because a run failed that way. Reviews are requested by a person or required by project instructions at a named transition; the ledger records them and does not schedule them.

The independent review guide covers when to request a review, reviewer independence, and a worked example.

review ledger / layout · source

File layout

# Review ledger

## Review register

| ID | Date | Kind | Target | Reviewer | Relation | Findings |
| --- | --- | --- | --- | --- | --- | --- |
| R001 | 2026-09-10 | review-script | scripts/exp_002_comparison.py @ 3f1c9a2e | fresh session | same model, no shared context | F001, F002 |
| R002 | 2026-09-11 | review-script | scripts/exp_002_comparison.py @ 8b04d17c | second provider | different model | none |

## Open findings

| ID | Fault | Caught by | Severity | Status |
| --- | --- | --- | --- | --- |
| F002 | Summary averages over all cells, including the excluded pilot cell | R001 | silent | repair landed; awaiting review |

## Gates — checkable from the artifact

#### G1 — Does the summary aggregate over the intended slice?
**Added** 2026-09-10, at EXP-002. **Severity:** silent. **Tally: 1** — F001 (founding). **Slug:** `aggregation-slice`.

What failed, in two or three sentences, with the experiment it cost.

**Check:** the one question a reviewer answers from the artifact.

## Gates — need project history

## Promoted

The validator requires a ## Review register section when the file exists, and checks the ## Open findings table's columns when that section is present. Edit the ledger through a tool that allocates IDs and serializes concurrent writers when one is available. Without one, append rows by hand and never renumber or delete them.

review ledger / review register

Review register

Every review gets a row, including clean ones; similar reviews may share a summary row. Columns, in order: ID | Date | Kind | Target | Reviewer | Relation | Findings.

ColumnMeaning
IDR-prefixed and unique. Many similar reviews may share one summary row, such as R040–R061: 22 script reviews, 2 findings.
KindThe kind of review, such as review-script: a plan, experiment design, experiment script, or results review.
TargetThe exact bytes reviewed: a path with a content digest or revision. A review of earlier bytes does not cover a revision.
Reviewer, RelationWho reviewed, and how the reviewer relates to the author.
FindingsThe finding IDs the review raised, or none.

Clean reviews are the denominator: without them, neither the practice's value nor any single gate's value can be measured. The Relation column takes one of self, same model, no shared context, different model, human.

RelationMeaning
selfThe author reviewed their own work.
same model, no shared contextA fresh session of the same model, given only the artifact and its record.
different modelA reviewer from a different model family or provider.
humanA person other than the author.

review ledger / open findings

Findings

Columns, in order: ID | Fault | Caught by | Severity | Status. Severity is one of silent or loud: silent when the fault would have produced a wrong number with no error, and loud when it would have failed visibly.

Record that a repair landed, and close the finding when a review of the repaired bytes settles it.

review ledger / gates

Gates

Each gate entry has a heading of the form #### G1 — Question?, a metadata line with the date and experiment where it was added, its severity, its tally with the findings attached to it, and a slug; two or three sentences on what failed and the experiment it cost; and a Check line stating the one question a reviewer answers.

Add a gate when a run fails for a reason no existing gate covers. When a later finding is another instance of an existing gate, attach it and increment the tally instead of adding a near duplicate. Resemblance between two faults is not by itself evidence that they are the same.

SectionContainsGiven to
Gates checkable from the artifactQuestions decidable from a script and its outputAny reviewer, including a fresh-context one
Gates that need project historyQuestions such as whether a sibling experiment already solved this problemOnly reviewers who have that history

A reviewer is never asked to work out which gates it can reach. Before submitting a new or modified experiment script, read the gate titles and open the detail only for gates that plausibly apply.

review ledger / promoted

Promotion

A gate with a high tally is a candidate for promotion: into a template, a shared checklist, or a mechanical check that makes the fault impossible to write. Record the promotion under ## Promoted and say what it does not cover; a template reaches only records written from it. A gate that stops firing is a candidate for retirement.

Independent review guide Experiment records Reference