Estimating and tracking costs
Cost-aware planning chooses the cheapest instrument sufficient for a scientific decision.
Plans and experiment records state what that instrument needs in engineering units such as
GPU-hours, forward passes, API calls, and elapsed time. Money lives in one place,
plans/spend/, where the owner's ceiling and the ledger of attempts sit together.
Keep the scientific design and its resource estimate in the plan, per-attempt usage in experiment records, and every currency amount in the spend directory. The spend reference specifies both files. The execution guide distinguishes a read-only preview from permission to run work. Discussing a cheaper option grants neither execution nor permission to change the records.
Choose an instrument that can settle the decision
An instrument is the analysis, check, or experiment that supplies evidence for a decision. Start with the decision and the smallest observation that could change it. Compare reusing saved data, checking the measurement, running a pilot, collecting the full comparison, and obtaining independent confirmation. Reject an option that cannot distinguish the alternatives, however cheap.
Controls establish whether the measurement is interpretable. Valid uncertainty describes what the observations can support. Independent confirmation tests a frozen claim using evidence that was not used to select it. Cost reduction must preserve these requirements; dropping controls or reusing selection data as confirmation undermines the intended inference.
A fictional paired retrieval campaign
A team wants to decide whether a revised retrieval method improves answer quality over its baseline. A paired comparison evaluates both methods on the same questions. Existing outputs can support preliminary analysis, but they do not cover the frozen revision. Scoring checks and a pilot precede the full comparison. Independent confirmation uses held-out questions after the design is frozen.
Every identifier, date, usage record, rate, and billing source below is synthetic. The campaign illustrates resource decisions, not scientific findings. No actual experiment records or provider quotations are represented.
| Candidate | Decision it can support |
|---|---|
| Existing-data analysis | Inspect saved paired outputs and uncertainty before collecting more. Preserve the selection history; reused evidence cannot independently confirm a claim selected from it. |
| Measurement checks | Use known-answer cases to test the scorer. A failure stops interpretation and can avoid an expensive invalid comparison. |
| Pilot | Check feasibility, resource use, and whether the instrument reaches the phenomenon. A small pilot does not replace the planned uncertainty analysis. |
| Full comparison | Collect the paired observations needed for the stated decision rule, with its controls and uncertainty requirements intact. |
| Independent confirmation | Test the frozen claim on untouched evidence when confirmation is required. Reserve the evidence and resources before exploratory choices consume them. |
Compare total resources, including data preparation, analysis, and human review. Existing-data analysis can still use compute and labor. A failed check may make stopping the cheapest sufficient action. Record what the cheaper option cannot answer before choosing a larger run.
Units in records
A unit here is an engineering quantity that determines what a run consumes: GPU-hours on a stated device class, forward passes, provider calls and tokens, elapsed time, person-hours, and retained bytes. The same pilot can bill differently on two devices a day apart while its forward count, seconds per forward, and device class stay fixed. A price is set when the job is placed and can change later, so a rate copied into a record goes stale.
Projections keep the configuration and drop the rate. "Two 8-GPU nodes for three hours" stays true when prices move, and it records what scale of hardware the arm requires, which is often part of the scientific argument.
Experiment records are also published. They serve as preregistration artifacts and paper supplements, and compute-reporting checklists ask for hardware and GPU-hours. Rates and an owner's ceilings are private operating detail. A record that names its Backend and Job ID loses nothing by omitting cost: the runner or provider holds the charge per job, and the spend ledger holds it per plan.
Scientific budgets, such as a retry budget or an alpha budget, are design quantities and stay in the record. So do dollars that are data rather than spend, such as a simulator's cost objective or a market price fed to a cost model, inside a marked region that names what they are.
Estimate setup, attempts, and conditional work in units
A baseline is the dated estimate for the intended campaign under stated assumptions. An attempt is one execution, including an execution that fails. Estimate setup, successful attempts, failures and retries, teardown, and any conditional branches separately, each in the units that determine it.
In rental accounting, one GPU provisioned for one hour consumes one GPU-hour; two GPUs consume two GPU-hours. Keep compute quantity, elapsed time, and human effort separate. Parallel execution can shorten elapsed time without reducing GPU-hours. Queue waits can lengthen elapsed time without consuming compute. Human review time is neither.
The campaign baseline
The synthetic plan 2026-09-09-paired-retrieval records this estimate on
2026-09-09. Each attempt runs on one GPU of a single synthetic device class, and the
estimate counts provisioned time, including idle time. CPU work outside the rental,
storage, network transfer, and human labor are outside this estimate; they are unknown,
not zero.
| Planned work | Experiment | GPU-hours |
|---|---|---|
| Setup and scoring checks | EXP-001 | 0.5 |
| Pilot | EXP-002 | 1 |
| Comparison: four attempts | EXP-003 | 4 × 1 = 4 |
| Independent confirmation | EXP-004 | 3 |
| Baseline | 0.5 + 1 + 4 + 3 = 8.5 | |
| Retry policy | EXP-002 | one pilot retry, at most 1 |
The baseline assumes the pilot succeeds on its first attempt and the scientific gates permit comparison and confirmation. Confirmation is planned but conditional on those gates. A stop after a failed gate is a different branch with its own estimate; no outcome probability has been assigned to either branch.
The plan contains no rate and no dollar figure. The held-out questions form its scientific confirmation reserve, which holds back evidence. Money held back for retries and idle time is a separate matter, recorded as an allocation in the authority file.
Estimate elapsed time separately using dependencies, concurrency, queue delays, and review waits. Estimate human effort for setup, diagnosis, analysis, and review in person-hours. Neither is estimated numerically in this GPU-rental example, so 8.5 GPU-hours is not a completion-time or labor promise.
When attempt duration is uncertain, retain a range or a conservative bound and explain its basis. A pilot measurement is evidence for a revision, not a guarantee that every full run will use the same resources. List conditional paths without inventing probabilities to make a precise-looking expected total.
Preview resources and spend without changing records
Replace PLAN_PATH with the plan path and send this prompt to the coding agent.
The preview may read the spend files and calculate a proposed rollup in its reply. It
cannot append ledger rows or launch work.
Use research-lab-notebook to preview the resources and spend for PLAN_PATH using existing records and saved outputs only. Compare existing-data analysis, measurement checks, a pilot, the full comparison, and required independent confirmation against the scientific decision. Preserve controls and valid uncertainty. Report the estimate in units: compute quantity with device class and count, forward passes or calls and tokens, elapsed time, human effort, and retained bytes, with setup, attempt counts, retries, and conditional branches kept separate. Read plans/spend/AUTHORITY.json and plans/spend/LEDGER.md for this plan's filename stem and derive incurred A and committed C as of a stated time. Price the remaining work U only at a stated current rate, name its source, and compare headroom with it. Treat unknown spend as unknown, never as zero. Do not modify files, append ledger rows, launch or retry jobs, or change execution authority.
Record usage in experiments and money in the ledger
Experiment records own the usage evidence for each attempt. Record it beside
Runs, optionally under a Resources
heading, keyed by the existing Backend plus Job ID. A manual
run needs a stable attempt identity too. A retry is another attempt linked to the failed one.
Record units, device, elapsed time, the usage source, and its retrieval timestamp; record
unknown usage as unknown, with the reason and the next check.
Money for each attempt is a row in plans/spend/LEDGER.md, keyed by the same
attempt identity. Committed is the amount the authorization reserved.
Actual is the best current figure for what the attempt cost, or
- while unknown. An attempt holds its commitment with outcome
authorized or running, then takes a terminal outcome.
A /provisional suffix marks an actual that is usage-estimated or not yet billed.
Failed attempts keep their rows: a run that billed before it failed still spent the money.
Derive A, C, and U from the ledger at one as-of time
For one plan at one as-of time, A (incurred) is the sum of Actual on the
plan's rows. C (committed) is the sum of Committed on rows whose outcome is
authorized or running. U (planned) is not stored:
it is the plan's remaining work in units, priced at a current rate when an authorization is
built. Headroom after commitments is the shared ceiling minus A minus C, and it must still
cover U. Unknown spend renders as unknown, and an incomplete subtotal is not a settled total.
The campaign at its checkpoint
At synthetic time 2026-09-10T12:00Z, setup is complete. The first pilot failed after
0.25 GPU-hour; its authorized retry completed in 1 GPU-hour. Two comparison attempts
finished. Two more are running, and each has consumed 0.25 GPU-hour. Confirmation has not
been submitted. Usage comes from synthetic runtime snapshot SYN-USAGE-01,
which does not yet report provisioned idle time.
## Resources
| Attempt | Job | Device | GPU-hours | State | Source |
|---|---|---|---|---|---|
| comp-1 | demo:SYN-COMP-1 | 1 × synthetic GPU | 1 | complete | SYN-USAGE-01 |
| comp-2 | demo:SYN-COMP-2 | 1 × synthetic GPU | 1 | complete | SYN-USAGE-01 |
| comp-3 | demo:SYN-COMP-3 | 1 × synthetic GPU | 0.25 so far | running | SYN-USAGE-01 |
| comp-4 | demo:SYN-COMP-4 | 1 × synthetic GPU | 0.25 so far | running | SYN-USAGE-01 |
SYN-USAGE-01 retrieved 2026-09-10T12:00Z; provisioned idle time not reported. plans/spend/LEDGER.md at the checkpoint| Row | Date | Plan | Experiment | Attempt | Job | Phase | Committed USD | Actual USD | Outcome |
|---|---|---|---|---|---|---|---|---|---|
| S001 | 2026-09-09 | 2026-09-09-paired-retrieval | EXP-001 | setup | demo:SYN-SETUP | setup | 1.00 | 1.00 | complete/provisional |
| S002 | 2026-09-09 | 2026-09-09-paired-retrieval | EXP-002 | pilot-r1 | demo:SYN-PILOT-1 | pilot | 2.00 | 0.50 | failed/provisional |
| S003 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-002 | pilot-r2 | demo:SYN-PILOT-2 | pilot | 2.00 | 2.00 | complete/provisional |
| S004 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-1 | demo:SYN-COMP-1 | comparison | 2.00 | 2.00 | complete/provisional |
| S005 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-2 | demo:SYN-COMP-2 | comparison | 2.00 | 2.00 | complete/provisional |
| S006 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-3 | demo:SYN-COMP-3 | comparison | 2.00 | - | running |
| S007 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-4 | demo:SYN-COMP-4 | comparison | 2.00 | - | running | | Quantity | Derivation | USD |
|---|---|---|
| A: incurred | Actual on S001 to S005 | 1.00 + 0.50 + 2.00 + 2.00 + 2.00 = 7.50 |
| C: committed | Committed on the running rows S006 and S007 | 2.00 + 2.00 = 4.00 |
| Headroom | Ceiling − A − C | 21.00 − 7.50 − 4.00 = 9.50 |
| U: planned | Confirmation, 3 GPU-hours at the teaching rate | 6.00 |
| Headroom after U | 9.50 − 6.00 | 3.50 |
The running attempts' 0.25 GPU-hour each stays in the experiment record; their money stays in the commitment until they end, so it is counted once. The failed pilot drew on the contingency allocation. Every actual so far is usage-estimated, and idle time is still unreported, so admitting confirmation needs a conservative allowance for that exposure.
When an attempt ends, its row takes a terminal outcome and its actual, and its commitment leaves C. When approved planned work is submitted, its authorization appends a new row that moves the work from U into C. A cancellation request releases a commitment only after cancellation is confirmed, with any consumption before shutdown kept in the actual. Headroom and underspend never grant scope approval.
A charge shared by several experiments needs an attribution policy and a source identity. Allocate it once in the ledger, or leave it visibly unallocated, and link each experiment to the evidence without charging each for its full value.
Reconcile final billing without reopening execution
Reconciliation matches later usage and billing evidence to the attempts and
resolves discrepancies. Financial settlement means the plan's ledger rows
have their final figures. Scientific processing may finish earlier, provided the missing
charges and a follow-up owner are recorded. A wrong or late figure is corrected by a new
correction row naming the same attempt and carrying the difference, so the
Actual column still sums to the total. The attempt's own row is never restated.
The campaign's provisional close and final settlement
The two running comparisons finish, each using its remaining 0.75 GPU-hour. After the
required gate and within the existing approval, confirmation is authorized as
demo:SYN-CONFIRM and completes in 3 GPU-hours. The experiment records now sum
to 0.5 + 0.25 + 1 + 4 + 3 = 8.75 GPU-hours. No scientific result is asserted here.
Synthetic final bill SYN-BILL-01, retrieved on 2026-09-12T12:00Z and covering
the campaign's full rental interval, reports 9 GPU-hours. Its line for
demo:SYN-SETUP includes 0.25 GPU-hour of provisioned idle time missing from
the runtime snapshots, billed at USD 0.50. The other attempts agree with their estimates.
EXP-001 records setup as 0.75 GPU-hours, keeping the earlier 0.5 as dated history
superseded by the bill. The ledger takes the difference as row S009.
plans/spend/LEDGER.md after settlement| Row | Date | Plan | Experiment | Attempt | Job | Phase | Committed USD | Actual USD | Outcome |
|---|---|---|---|---|---|---|---|---|---|
| S001 | 2026-09-09 | 2026-09-09-paired-retrieval | EXP-001 | setup | demo:SYN-SETUP | setup | 1.00 | 1.00 | complete |
| S002 | 2026-09-09 | 2026-09-09-paired-retrieval | EXP-002 | pilot-r1 | demo:SYN-PILOT-1 | pilot | 2.00 | 0.50 | failed |
| S003 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-002 | pilot-r2 | demo:SYN-PILOT-2 | pilot | 2.00 | 2.00 | complete |
| S004 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-1 | demo:SYN-COMP-1 | comparison | 2.00 | 2.00 | complete |
| S005 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-2 | demo:SYN-COMP-2 | comparison | 2.00 | 2.00 | complete |
| S006 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-3 | demo:SYN-COMP-3 | comparison | 2.00 | 2.00 | complete |
| S007 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-4 | demo:SYN-COMP-4 | comparison | 2.00 | 2.00 | complete |
| S008 | 2026-09-11 | 2026-09-09-paired-retrieval | EXP-004 | confirm | demo:SYN-CONFIRM | confirmation | 6.00 | 6.00 | complete |
| S009 | 2026-09-12 | 2026-09-09-paired-retrieval | EXP-001 | setup | demo:SYN-SETUP | setup | 0.00 | 0.50 | correction | | Quantity | Units, from records | USD, from the ledger |
|---|---|---|
| Baseline | 8.5 GPU-hours | 17.00 when the authority was built |
| A: incurred | 9 GPU-hours | 7.50 + 2.00 + 2.00 + 6.00 + 0.50 = 18.00 |
| C: committed | none running | 0.00 |
| U: planned | none remaining | 0.00 |
| Variance from baseline | +0.5 GPU-hours | +1.00 |
| Unused ceiling | 21.00 − 18.00 = 3.00 |
The bill settles the usage-estimated actuals, so their provisional marks come off. The setup figure differs, so it gets a correction row instead of an edit to S001. The variance is the failed pilot's 0.25 GPU-hour plus the missing idle time's 0.25 GPU-hour, or USD 0.50 each. The unused USD 3 authorizes no further runs. A successor estimate should include provisioned idle overhead and the observed retry exposure, without treating one failure as an estimated failure probability.
At closure, the plan compares resources used with the baseline estimate in units and explains the variance. The ledger holds the money; the plan names any pending charges and the owner of their follow-up. A scientifically completed plan can have financial settlement pending, and its completion report should say both. If a provider later revises a bill, append another correction row; scientific closure does not make an earlier charge immutable.
Authorize an accounting-only update
Use this separate request when you want the records changed. Replace PLAN_PATH,
EXPERIMENT_PATHS, SOURCE_PATHS, and AS_OF with the
plan, owning experiment files, saved usage or billing evidence, and cutoff timestamp.
This authorizes accounting edits, not experiment execution or a larger ceiling.
Use research-lab-notebook to update only the accounting for PLAN_PATH as of AS_OF, using the saved usage and billing evidence in SOURCE_PATHS. In EXPERIMENT_PATHS, record each attempt's usage in units under its Backend plus Job ID or stable manual-attempt identity, with source identity, coverage, and retrieval timestamp. Keep superseded usage figures as dated history, and write no currency amounts into records. In plans/spend/LEDGER.md, give each finished attempt's open row its terminal outcome and actual, marking usage-estimated or unbilled actuals /provisional. When evidence changes a terminal figure, append a correction row that names the same attempt and carries the difference; never restate, renumber, or delete a row. Apply the recorded shared-cost attribution policy once and disclose unallocated or unknown charges. Report A, C, headroom against plans/spend/AUTHORITY.json, variance from the baseline in units, and pending settlement with its follow-up owner, or flag that one is missing. Do not edit AUTHORITY.json, launch, submit, retry, cancel, or rerun jobs, collect new scientific evidence, approve new work, or treat this update as execution approval.
A read-only results walkthrough may explain these figures from existing records. Writing corrections requires separate authorization such as the accounting request above. Any proposed follow-up returns to the plan workflow for a scientific and resource decision.