Spend and Units
Money has one home in a notebook: plans/spend/. Its
AUTHORITY.json holds the ceilings and allocations the owner approved,
and its LEDGER.md holds money committed and spent per authorized
attempt. Experiment records, plans, and experiment scripts state scientific quantities and
engineering units, and carry no currency amount of any kind.
The cost estimation and tracking guide follows a synthetic campaign through estimate, authorization, tracking, and settlement.
spend / surfaces · source
Surfaces
| Surface | Carries | Never carries |
|---|---|---|
| Experiment record | Scientific quantities; compute quantities such as GPU-hours, device class, forward passes, API calls, tokens, and retained bytes; elapsed time; Backend and Job ID for each attempt | Currency amounts of any kind |
| Plan | Scope, packages, the experiments in each package, and the resource estimate in units | Ceilings, allocations, rates, spend figures |
| Experiment script | The same units, and a runtime cap received as seconds, forward passes, calls, or tokens | Currency constants, rates, rate × time arithmetic, currency flags |
plans/spend/AUTHORITY.json | The ceilings and allocations the owner approved, who approved them, when, and their history | — |
plans/spend/LEDGER.md | Money committed and spent per authorized attempt | — |
Scientific budgets, such as a retry budget, a strike budget, or an alpha budget, are design quantities and stay in the record. The validator reports a currency amount in an experiment record as an error.
A price is set when a job is placed and can change later, so a rate copied into a record goes stale; the unit counts stay fixed. A ceiling copied into a script or record becomes a second home for it, and the copy goes stale: a script that still checks an old figure refuses the correct authorization. Records are published as preregistration artifacts and paper supplements, while rates and an owner's ceilings are private operating detail.
spend / units in records and plans
Units
- Compute quantity, with its unit and device count, such as 4 GPU-hours on one 48 GB device, or 24,300 forward passes.
- Calls and tokens for a provider API.
- Elapsed time, including queueing and parallelism assumptions.
- Human effort, in person-hours.
- Retained data, in bytes.
Keep them separate. Parallel execution shortens elapsed time without reducing GPU-hours; a queue wait lengthens elapsed time without consuming compute. A projection across providers or devices keeps the configuration and drops the rate: "two 8-GPU nodes for 3 hours" stays true when the price changes.
spend / prices as data
Prices as data
The rule governs spend on the research itself. A price can also be the object of study, or
an input to the model under study. Those dollars would mean the same thing if someone else had
paid for the compute, so they are scientific quantities. A record keeps them between an
opening usd: marker and a closing /usd marker:
<!-- usd: measured — simulator objective: total cost per job -->
| Policy | Cost per job (USD) |
|---|---|
| greedy | $1.20 |
<!-- /usd --> | Label | Use |
|---|---|
measured | Dollars a simulator or model outputs as its objective. They are experimental measurements. |
parameter | An external market input such as a spot price, a total-cost-of-ownership rate, or a market-derived cost-efficiency figure. The reason names the source or the model it feeds. |
The marker wraps lines, not files. One record often holds both a simulator table and a line about what the rental cost, and a file-level exemption would license the second under cover of the first. Nothing is inferred from position, because scientific dollars appear in prose as well as tables. The validator skips marked lines, reports a marker that is unclosed, nested, stray, unlabeled, or missing its reason, and adds a note with each record's marked line count so growth in exemptions stays visible.
The marker covers figures that remain once spend is out of the record, never spend itself.
A study of cost-effectiveness reports the units that determine the price and marks the
per-unit price as a parameter.
spend / LEDGER.md
Ledger
The ledger is append-only and includes failed and voided attempts: a run that billed before
it failed still spent the money. Its table uses these columns, in order:
Row | Date | Plan | Experiment | Attempt | Job | Phase | Committed USD | Actual USD | Outcome.
| Row | Date | Plan | Experiment | Attempt | Job | Phase | Committed USD | Actual USD | Outcome |
|---|---|---|---|---|---|---|---|---|---|
| S001 | 2026-09-10 | 2026-09-01-retrieval-comparison | EXP-002 | pilot-r1 | demo:SYN-PILOT-2 | pilot | 3.00 | 2.00 | complete | | Column | Meaning |
|---|---|
Row | An S-prefixed number, unique across the whole file. |
Plan | A plan filename stem, or - for spend that belongs to no plan, such as an owner-directed evaluation with no ceiling. A - row needs no authority entry. |
Attempt, Job | The stable attempt identity the experiment record uses, and Backend:Job ID when the attempt ran on a runner. |
Committed USD | The amount the authorization reserved. |
Actual USD | The best current figure for what the attempt cost, or - while unknown. |
Outcome | An open or terminal outcome, described below. |
An attempt holding its commitment has an open outcome:
authorized or running.
Afterward it takes a terminal outcome:
complete, failed, canceled, operational-void, correction.
Append /provisional to a terminal value when the actual is
usage-estimated or not yet billed.
A wrong or late figure is corrected by a new correction row naming the same
Attempt and carrying the difference, so the column sum stays the total.
A final bill that adds 0.50 for idle time appends a 0.50 correction; it does not restate the
attempt. A row with the wrong number of cells is an error, never skipped, because a skipped
money row understates spend.
spend / deriving the account
Deriving the account
A, C, and U are computed when needed and never stored. For one plan at one as-of time, taking only rows whose Plan is that stem:
| Quantity | Derived from |
|---|---|
| A: incurred | The sum of Actual. |
| C: committed | The sum of Committed on rows whose Outcome is authorized or running. |
| U: planned | Not stored. The plan's remaining work in units, priced at a current rate when an authorization is built. |
Headroom after commitments = shared ceiling − A − C, and it must still cover U.
Neither headroom nor an underspend grants approval for more work. Unknown spend renders as
unknown, never as zero, and an incomplete subtotal is not a settled total.
spend / run authorization
How a run gets its budget
Tooling that runs where the notebook is readable builds each authorization. It reads the authority and the ledger, checks the requested ceiling against headroom, converts the ceiling and the placement's rate into seconds (or calls, or tokens), writes the commitment as a ledger row, and hands the script an authorization in those units.
The script enforces the units it was given and refuses an authorization that carries any currency field, since a money field could carry a ceiling that has since changed. It never reads the notebook, which often does not reach the execution host, and holds no plan-level number of its own.
Cost estimation and tracking guide Plan resources and spend Experiment resources Reference