Research Notebook System

Keep the reasoning with the experiments

Research Notebook connects questions, experimental evidence, and the decisions that follow. It keeps this account in linked Markdown files alongside the project's code.

You can adopt the notebook a piece at a time, without learning its layout first. Describe what you want to your coding agent in ordinary sentences, and it creates and updates the files. Setup creates a few core files; plans, spend tracking, and manuscripts are added the first time you ask for them. The guide and reference describe what the agent writes, for when you want to check it.

Continuity across people and sessions
Researchers and agents can resume from recorded predictions, results, and next actions.
Claims traceable to evidence
A claim identifies the experiment, measured quantity, and design choices that support its stated scope.

An example conversation with the agent

Five requests take the example project from setup to a supported claim. Its job IDs and results are illustrative.

  1. YouSet up a research notebook. Our question is whether gradient accumulation reproduces true large-batch training. Jobs run through Slurm.

    AgentCreates lab-notebook/ and records the question as RQ1.

    STATUS.mdQUESTIONS.md

  2. YouDesign a cheap pilot for RQ1. Don't run it yet.

    AgentWrites the hypothesis, a prediction that the two losses differ by less than 0.02 at one seed, and a decision rule, then asks you to review them.

    EXP-001-accumulation-pilot.md

  3. YouLooks good. Run it.

    AgentSubmits Slurm job 48152. When it finishes, records a difference of 0.004 and applies the decision rule: proceed to the three-seed comparison.

    EXP-001-accumulation-pilot.md

  4. You, in a new sessionWhere are we? Continue.

    AgentReads the status file, finds the pilot result and its next step, and drafts EXP-002, a three-seed comparison, for your review.

    STATUS.mdEXP-002-accumulation-comparison.md

  5. YouApproved. Run it, then write up what we found.

    AgentAfter the run, writes a finding that combines both experiments and adds claim C1, citing EXP-001:E1, EXP-002:E1, and EXP-002:E2.

    findings/2026-08-16-…CLAIMS.md

Each file name opens that file in the explorer below, except the pilot record, which opens on GitHub. Install the skill to start a notebook in your own project.

Notebook files

Explore the notebook folder, then open an experiment or plan file to see its sections. The two files come from the example notebook, which tests whether gradient accumulation reproduces true large-batch training.

lab-notebook/ Select a file to see its role.
plans/
experiments/
findings/
papers/
reports/
references/

STATUS.md

The project's current state: what's established, what's blocked, and what to do next. Links point to the supporting records.

Written byresearcher + agent

Used byresearchers and agents starting a session

Indexes and dashboards draw from these files. Add an optional record when the project needs it.

System features

FeatureBenefit
Portable, version-controlled files Markdown files work offline, support Git/jj diffs, and are searchable with ordinary tools. Obsidian provides linked navigation. The notebook can have its own version-control repository.
Research questions organize the work Each experiment links to its research questions and keeps its hypothesis, method, runs, results, and interpretation together. Findings combine evidence from several experiments.
One home for each kind of information STATUS describes the project's current state; QUESTIONS tracks inquiry; PRIORITIES identifies next work; PUBLICATION tracks submission readiness. Each links to the supporting records, so a result can be updated in one place.
Handoffs between sessions and agents One session can design and launch an experiment, another process its results, and a third continue the work. Recorded predictions, decision rules, and follow-ups let each recover the reasoning behind the next step.
Registration for each measured quantity An estimand is the quantity an experiment seeks to estimate. The notebook records what was fixed in advance for each estimand, plus what earlier evidence informed the design. This distinguishes planned tests from discoveries within the same run and makes post-hoc reinterpretation easier to detect.
Traceable claims and provenance A claim cites a specific experiment estimand, such as EXP-042:E2. Links connect plans, experiments, jobs, claims, and papers. A reader can trace a claim's support and identify what needs reconsideration when evidence changes.
A result-processing workflow After a run finishes, validate the output and instrument, record the result, apply the decision rule, update dependencies, preserve the artifacts, and report. Process failed and null results too, so their evidence and follow-ups remain available.
Cost estimation and spend tracking Plans estimate work in units that stay reproducible: GPU-hours, API calls, tokens, elapsed time. Experiment records log usage for each attempt. All money is kept in plans/spend/: the owner's approved ceiling in AUTHORITY.json, and an append-only ledger of what each attempt committed and cost. Incurred spend, committed spend, and remaining headroom are calculated from the ledger. A run gets its budget as seconds or calls, never dollars, and the validator flags any currency amount in an experiment record.
Mechanical checks and review tracking Validators check defined formats and references. Human review decisions stay with the experiments, plans, and claims they govern. Content-hash review tracking can identify changed artifacts that need another review.
A path from evidence to publication CLAIMS.md links paper claims to their evidence; PUBLICATION.md records paper scope, venues, and blockers. Manuscripts and supporting notes stay linked to the experiments they draw on.

Session handoffs

  1. Design and launch

    Write the hypothesis, predictions, and decision branches before inspecting results. Record what earlier evidence informed the design. Estimate the work in units and get its spend approved.

  2. Process the result

    Check the output and instrument, record the outcome, apply the decision rule, update dependencies, and preserve the artifacts.

  3. Apply the recorded decision

    The next session follows the decision rule and follow-ups. Failed and null results can close a branch as well as open one.

Check, interpret, and record results before moving on to the next experiment. The same requirement applies to laptop runs and remote jobs.

Installation

The public agent skill supplies the notebook conventions and workflows.

1. Install in a terminal

Run this command from your research project's directory:

npx skills add osteele/research-notebook -s research-lab-notebook -y

2. Ask your coding agent to set up the notebook

After installation, send this prompt to the agent working in your project:

Use $research-lab-notebook to add a research notebook to this project. Jobs run through Slurm.

Name your project's runner or explain how you run scripts locally. The skill supplies instructions for working with those tools; it does not install a scheduler or research loop.

Start with the records your project needs, and keep them current as the research changes.

For designing a follow-up or assessing a claim, use the evidence and data-reuse workflow. It covers experiment records, confirmation reserves, and the claims ledger.

Integration with other research tools

Experiment files connect hypotheses and interpretations to outputs from metric trackers, computational notebooks, version control, and job runners. The spend ledger connects each attempt to the provider's charges for it.

Connections through file links and workflow conventions
ToolWhat it managesConnection to the notebook
Weights & Biases / MLflowMetrics, run comparisons, and hyperparameters.Experiment records link to runs and recorded metrics, then document the hypothesis, interpretation, and resulting claims.
JupyterInteractive exploration and computational analysis.Experiment records reference computational notebooks and their outputs, preserving the design and conclusions alongside the analysis.
Git / jjCode history and versioned changes.Notebook files are versioned; experiment records identify the code revision that produced a result.
Slurm / SkyPilot / WeftJob execution, status, logs, and output retrieval.Run records retain job IDs and artifact locations. Result processing records whether the output was validated, interpreted, and acted on.
Provider billingCharges for compute and API calls, from cloud consoles and API usage dashboards.Spend-ledger rows are keyed to job IDs. When a final bill differs from the recorded charge, the difference goes in as a correction row in plans/spend/LEDGER.md.

These connections use file links and project-specific commands. API integration requires a separate connector.

Read and edit the files in Obsidian, a text editor, or a coding agent.