Estimating and tracking costs

documentation/ guide and reference

Cost-aware planning chooses the cheapest instrument sufficient for a scientific decision. Plans and experiment records state what that instrument needs in engineering units such as GPU-hours, forward passes, API calls, and elapsed time. Money lives in one place, plans/spend/, where the owner's ceiling and the ledger of attempts sit together.

Keep the scientific design and its resource estimate in the plan, per-attempt usage in experiment records, and every currency amount in the spend directory. The spend reference specifies both files. The execution guide distinguishes a read-only preview from permission to run work. Discussing a cheaper option grants neither execution nor permission to change the records.

Choose an instrument that can settle the decision

An instrument is the analysis, check, or experiment that supplies evidence for a decision. Start with the decision and the smallest observation that could change it. Compare reusing saved data, checking the measurement, running a pilot, collecting the full comparison, and obtaining independent confirmation. Reject an option that cannot distinguish the alternatives, however cheap.

Controls establish whether the measurement is interpretable. Valid uncertainty describes what the observations can support. Independent confirmation tests a frozen claim using evidence that was not used to select it. Cost reduction must preserve these requirements; dropping controls or reusing selection data as confirmation undermines the intended inference.

A fictional paired retrieval campaign

A team wants to decide whether a revised retrieval method improves answer quality over its baseline. A paired comparison evaluates both methods on the same questions. Existing outputs can support preliminary analysis, but they do not cover the frozen revision. Scoring checks and a pilot precede the full comparison. Independent confirmation uses held-out questions after the design is frozen.

Every identifier, date, usage record, rate, and billing source below is synthetic. The campaign illustrates resource decisions, not scientific findings. No actual experiment records or provider quotations are represented.

CandidateDecision it can support
Existing-data analysisInspect saved paired outputs and uncertainty before collecting more. Preserve the selection history; reused evidence cannot independently confirm a claim selected from it.
Measurement checksUse known-answer cases to test the scorer. A failure stops interpretation and can avoid an expensive invalid comparison.
PilotCheck feasibility, resource use, and whether the instrument reaches the phenomenon. A small pilot does not replace the planned uncertainty analysis.
Full comparisonCollect the paired observations needed for the stated decision rule, with its controls and uncertainty requirements intact.
Independent confirmationTest the frozen claim on untouched evidence when confirmation is required. Reserve the evidence and resources before exploratory choices consume them.

Compare total resources, including data preparation, analysis, and human review. Existing-data analysis can still use compute and labor. A failed check may make stopping the cheapest sufficient action. Record what the cheaper option cannot answer before choosing a larger run.

Units in records

A unit here is an engineering quantity that determines what a run consumes: GPU-hours on a stated device class, forward passes, provider calls and tokens, elapsed time, person-hours, and retained bytes. The same pilot can bill differently on two devices a day apart while its forward count, seconds per forward, and device class stay fixed. A price is set when the job is placed and can change later, so a rate copied into a record goes stale.

Projections keep the configuration and drop the rate. "Two 8-GPU nodes for three hours" stays true when prices move, and it records what scale of hardware the arm requires, which is often part of the scientific argument.

Experiment records are also published. They serve as preregistration artifacts and paper supplements, and compute-reporting checklists ask for hardware and GPU-hours. Rates and an owner's ceilings are private operating detail. A record that names its Backend and Job ID loses nothing by omitting cost: the runner or provider holds the charge per job, and the spend ledger holds it per plan.

Scientific budgets, such as a retry budget or an alpha budget, are design quantities and stay in the record. So do dollars that are data rather than spend, such as a simulator's cost objective or a market price fed to a cost model, inside a marked region that names what they are.

Estimate setup, attempts, and conditional work in units

A baseline is the dated estimate for the intended campaign under stated assumptions. An attempt is one execution, including an execution that fails. Estimate setup, successful attempts, failures and retries, teardown, and any conditional branches separately, each in the units that determine it.

In rental accounting, one GPU provisioned for one hour consumes one GPU-hour; two GPUs consume two GPU-hours. Keep compute quantity, elapsed time, and human effort separate. Parallel execution can shorten elapsed time without reducing GPU-hours. Queue waits can lengthen elapsed time without consuming compute. Human review time is neither.

The campaign baseline

The synthetic plan 2026-09-09-paired-retrieval records this estimate on 2026-09-09. Each attempt runs on one GPU of a single synthetic device class, and the estimate counts provisioned time, including idle time. CPU work outside the rental, storage, network transfer, and human labor are outside this estimate; they are unknown, not zero.

Planned workExperimentGPU-hours
Setup and scoring checksEXP-0010.5
PilotEXP-0021
Comparison: four attemptsEXP-0034 × 1 = 4
Independent confirmationEXP-0043
Baseline0.5 + 1 + 4 + 3 = 8.5
Retry policyEXP-002one pilot retry, at most 1

The baseline assumes the pilot succeeds on its first attempt and the scientific gates permit comparison and confirmation. Confirmation is planned but conditional on those gates. A stop after a failed gate is a different branch with its own estimate; no outcome probability has been assigned to either branch.

The plan contains no rate and no dollar figure. The held-out questions form its scientific confirmation reserve, which holds back evidence. Money held back for retries and idle time is a separate matter, recorded as an allocation in the authority file.

Estimate elapsed time separately using dependencies, concurrency, queue delays, and review waits. Estimate human effort for setup, diagnosis, analysis, and review in person-hours. Neither is estimated numerically in this GPU-rental example, so 8.5 GPU-hours is not a completion-time or labor promise.

When attempt duration is uncertain, retain a range or a conservative bound and explain its basis. A pilot measurement is evidence for a revision, not a guarantee that every full run will use the same resources. List conditional paths without inventing probabilities to make a precise-looking expected total.

Preview resources and spend without changing records

Replace PLAN_PATH with the plan path and send this prompt to the coding agent. The preview may read the spend files and calculate a proposed rollup in its reply. It cannot append ledger rows or launch work.

Use research-lab-notebook to preview the resources and spend for PLAN_PATH using existing records and saved outputs only. Compare existing-data analysis, measurement checks, a pilot, the full comparison, and required independent confirmation against the scientific decision. Preserve controls and valid uncertainty. Report the estimate in units: compute quantity with device class and count, forward passes or calls and tokens, elapsed time, human effort, and retained bytes, with setup, attempt counts, retries, and conditional branches kept separate. Read plans/spend/AUTHORITY.json and plans/spend/LEDGER.md for this plan's filename stem and derive incurred A and committed C as of a stated time. Price the remaining work U only at a stated current rate, name its source, and compare headroom with it. Treat unknown spend as unknown, never as zero. Do not modify files, append ledger rows, launch or retry jobs, or change execution authority.

Authorize a scope in the plan and a ceiling in the authority file

An approved ceiling is the most money the owner allows a plan to spend, inclusive of everything already spent. Approval also names the work, permitted retries, scientific gates, and who may decide to proceed. The plan records the scope, retry policy, and authorized branches in units. The ceiling, any sub-ceilings, who approved them, and when go in plans/spend/AUTHORITY.json under the plan's filename stem, so an ordinary status move of the plan file cannot orphan its authority.

An allocation is an optional sub-ceiling inside the shared ceiling; it does not add to it. Contingency is one such allocation: capacity for specified uncertainty or failure, not money already spent. A raise or reallocation updates the current fields and appends the previous state to history with a dated reason. A larger ceiling does not authorize a different study.

The campaign's bounded approval

To build the authorization, the owner prices the 8.5 GPU-hour baseline and a 2 GPU-hour contingency at an invented teaching rate of USD 2 per provisioned GPU-hour. This rate is synthetic, not a provider quote. The baseline prices at USD 17 and the contingency at USD 4, so the synthetic approval on 2026-09-09 sets a shared ceiling of USD 21. The rate does not appear in the plan or any record.

Synthetic plans/spend/AUTHORITY.json
{
  "schema": "plan-spend-authority/v1",
  "plans": {
    "2026-09-09-paired-retrieval": {
      "shared_ceiling_usd": 21.0,
      "authorized_by": "Synthetic owner",
      "recorded_utc": "2026-09-09T00:00:00+00:00",
      "allocations": {
        "contingency": {
          "ceiling_usd": 4.0,
          "note": "Retry and provisioned-idle uncertainty, including one pilot retry."
        }
      },
      "history": []
    }
  }
}

The plan covers setup, the pilot, four comparison attempts, and independent confirmation after their scientific gates. A single pilot retry of at most 1 GPU-hour is permitted after an execution failure is diagnosed and recorded. Another retry, a changed design, or work beyond the ceiling requires a new decision. These statements authorize no real job.

Each run gets its budget from tooling that can read the notebook. It reads the authority and the ledger, checks the requested amount against headroom, converts that amount and the placement's rate into seconds, calls, or tokens, writes the commitment as a ledger row, and hands the script an authorization in those units. A comparison attempt authorized for USD 2 at the teaching rate on one GPU becomes a cap of 3,600 seconds. The script enforces the cap it was given and refuses any authorization carrying a currency field, because a money field could carry a ceiling that has since changed. The script never reads the notebook, which often does not reach the execution host.

Before submitting parallel jobs, admit them against conservative upper bounds, including permitted retries, provisioning overhead, billing increments, and the delay before a stop takes effect. Several agents cannot each spend the same headroom. A written limit does not stop a process: describe the duration, concurrency, and spending controls the runner actually enforces in RUNNER.md, and use explicit checkpoints where hard enforcement is unavailable.

Record usage in experiments and money in the ledger

Experiment records own the usage evidence for each attempt. Record it beside Runs, optionally under a Resources heading, keyed by the existing Backend plus Job ID. A manual run needs a stable attempt identity too. A retry is another attempt linked to the failed one. Record units, device, elapsed time, the usage source, and its retrieval timestamp; record unknown usage as unknown, with the reason and the next check.

Money for each attempt is a row in plans/spend/LEDGER.md, keyed by the same attempt identity. Committed is the amount the authorization reserved. Actual is the best current figure for what the attempt cost, or - while unknown. An attempt holds its commitment with outcome authorized or running, then takes a terminal outcome. A /provisional suffix marks an actual that is usage-estimated or not yet billed. Failed attempts keep their rows: a run that billed before it failed still spent the money.

Derive A, C, and U from the ledger at one as-of time

For one plan at one as-of time, A (incurred) is the sum of Actual on the plan's rows. C (committed) is the sum of Committed on rows whose outcome is authorized or running. U (planned) is not stored: it is the plan's remaining work in units, priced at a current rate when an authorization is built. Headroom after commitments is the shared ceiling minus A minus C, and it must still cover U. Unknown spend renders as unknown, and an incomplete subtotal is not a settled total.

The campaign at its checkpoint

At synthetic time 2026-09-10T12:00Z, setup is complete. The first pilot failed after 0.25 GPU-hour; its authorized retry completed in 1 GPU-hour. Two comparison attempts finished. Two more are running, and each has consumed 0.25 GPU-hour. Confirmation has not been submitted. Usage comes from synthetic runtime snapshot SYN-USAGE-01, which does not yet report provisioned idle time.

Synthetic excerpt from EXP-003 at the checkpoint
## Resources

| Attempt | Job | Device | GPU-hours | State | Source |
|---|---|---|---|---|---|
| comp-1 | demo:SYN-COMP-1 | 1 × synthetic GPU | 1 | complete | SYN-USAGE-01 |
| comp-2 | demo:SYN-COMP-2 | 1 × synthetic GPU | 1 | complete | SYN-USAGE-01 |
| comp-3 | demo:SYN-COMP-3 | 1 × synthetic GPU | 0.25 so far | running | SYN-USAGE-01 |
| comp-4 | demo:SYN-COMP-4 | 1 × synthetic GPU | 0.25 so far | running | SYN-USAGE-01 |

SYN-USAGE-01 retrieved 2026-09-10T12:00Z; provisioned idle time not reported.
Synthetic plans/spend/LEDGER.md at the checkpoint
| Row | Date | Plan | Experiment | Attempt | Job | Phase | Committed USD | Actual USD | Outcome |
|---|---|---|---|---|---|---|---|---|---|
| S001 | 2026-09-09 | 2026-09-09-paired-retrieval | EXP-001 | setup | demo:SYN-SETUP | setup | 1.00 | 1.00 | complete/provisional |
| S002 | 2026-09-09 | 2026-09-09-paired-retrieval | EXP-002 | pilot-r1 | demo:SYN-PILOT-1 | pilot | 2.00 | 0.50 | failed/provisional |
| S003 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-002 | pilot-r2 | demo:SYN-PILOT-2 | pilot | 2.00 | 2.00 | complete/provisional |
| S004 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-1 | demo:SYN-COMP-1 | comparison | 2.00 | 2.00 | complete/provisional |
| S005 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-2 | demo:SYN-COMP-2 | comparison | 2.00 | 2.00 | complete/provisional |
| S006 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-3 | demo:SYN-COMP-3 | comparison | 2.00 | - | running |
| S007 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-4 | demo:SYN-COMP-4 | comparison | 2.00 | - | running |
QuantityDerivationUSD
A: incurredActual on S001 to S0051.00 + 0.50 + 2.00 + 2.00 + 2.00 = 7.50
C: committedCommitted on the running rows S006 and S0072.00 + 2.00 = 4.00
HeadroomCeiling − A − C21.00 − 7.50 − 4.00 = 9.50
U: plannedConfirmation, 3 GPU-hours at the teaching rate6.00
Headroom after U9.50 − 6.003.50

The running attempts' 0.25 GPU-hour each stays in the experiment record; their money stays in the commitment until they end, so it is counted once. The failed pilot drew on the contingency allocation. Every actual so far is usage-estimated, and idle time is still unreported, so admitting confirmation needs a conservative allowance for that exposure.

When an attempt ends, its row takes a terminal outcome and its actual, and its commitment leaves C. When approved planned work is submitted, its authorization appends a new row that moves the work from U into C. A cancellation request releases a commitment only after cancellation is confirmed, with any consumption before shutdown kept in the actual. Headroom and underspend never grant scope approval.

A charge shared by several experiments needs an attribution policy and a source identity. Allocate it once in the ledger, or leave it visibly unallocated, and link each experiment to the evidence without charging each for its full value.

Reconcile final billing without reopening execution

Reconciliation matches later usage and billing evidence to the attempts and resolves discrepancies. Financial settlement means the plan's ledger rows have their final figures. Scientific processing may finish earlier, provided the missing charges and a follow-up owner are recorded. A wrong or late figure is corrected by a new correction row naming the same attempt and carrying the difference, so the Actual column still sums to the total. The attempt's own row is never restated.

The campaign's provisional close and final settlement

The two running comparisons finish, each using its remaining 0.75 GPU-hour. After the required gate and within the existing approval, confirmation is authorized as demo:SYN-CONFIRM and completes in 3 GPU-hours. The experiment records now sum to 0.5 + 0.25 + 1 + 4 + 3 = 8.75 GPU-hours. No scientific result is asserted here.

Synthetic final bill SYN-BILL-01, retrieved on 2026-09-12T12:00Z and covering the campaign's full rental interval, reports 9 GPU-hours. Its line for demo:SYN-SETUP includes 0.25 GPU-hour of provisioned idle time missing from the runtime snapshots, billed at USD 0.50. The other attempts agree with their estimates. EXP-001 records setup as 0.75 GPU-hours, keeping the earlier 0.5 as dated history superseded by the bill. The ledger takes the difference as row S009.

Synthetic plans/spend/LEDGER.md after settlement
| Row | Date | Plan | Experiment | Attempt | Job | Phase | Committed USD | Actual USD | Outcome |
|---|---|---|---|---|---|---|---|---|---|
| S001 | 2026-09-09 | 2026-09-09-paired-retrieval | EXP-001 | setup | demo:SYN-SETUP | setup | 1.00 | 1.00 | complete |
| S002 | 2026-09-09 | 2026-09-09-paired-retrieval | EXP-002 | pilot-r1 | demo:SYN-PILOT-1 | pilot | 2.00 | 0.50 | failed |
| S003 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-002 | pilot-r2 | demo:SYN-PILOT-2 | pilot | 2.00 | 2.00 | complete |
| S004 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-1 | demo:SYN-COMP-1 | comparison | 2.00 | 2.00 | complete |
| S005 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-2 | demo:SYN-COMP-2 | comparison | 2.00 | 2.00 | complete |
| S006 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-3 | demo:SYN-COMP-3 | comparison | 2.00 | 2.00 | complete |
| S007 | 2026-09-10 | 2026-09-09-paired-retrieval | EXP-003 | comp-4 | demo:SYN-COMP-4 | comparison | 2.00 | 2.00 | complete |
| S008 | 2026-09-11 | 2026-09-09-paired-retrieval | EXP-004 | confirm | demo:SYN-CONFIRM | confirmation | 6.00 | 6.00 | complete |
| S009 | 2026-09-12 | 2026-09-09-paired-retrieval | EXP-001 | setup | demo:SYN-SETUP | setup | 0.00 | 0.50 | correction |
QuantityUnits, from recordsUSD, from the ledger
Baseline8.5 GPU-hours17.00 when the authority was built
A: incurred9 GPU-hours7.50 + 2.00 + 2.00 + 6.00 + 0.50 = 18.00
C: committednone running0.00
U: plannednone remaining0.00
Variance from baseline+0.5 GPU-hours+1.00
Unused ceiling21.00 − 18.00 = 3.00

The bill settles the usage-estimated actuals, so their provisional marks come off. The setup figure differs, so it gets a correction row instead of an edit to S001. The variance is the failed pilot's 0.25 GPU-hour plus the missing idle time's 0.25 GPU-hour, or USD 0.50 each. The unused USD 3 authorizes no further runs. A successor estimate should include provisioned idle overhead and the observed retry exposure, without treating one failure as an estimated failure probability.

At closure, the plan compares resources used with the baseline estimate in units and explains the variance. The ledger holds the money; the plan names any pending charges and the owner of their follow-up. A scientifically completed plan can have financial settlement pending, and its completion report should say both. If a provider later revises a bill, append another correction row; scientific closure does not make an earlier charge immutable.

Authorize an accounting-only update

Use this separate request when you want the records changed. Replace PLAN_PATH, EXPERIMENT_PATHS, SOURCE_PATHS, and AS_OF with the plan, owning experiment files, saved usage or billing evidence, and cutoff timestamp. This authorizes accounting edits, not experiment execution or a larger ceiling.

Use research-lab-notebook to update only the accounting for PLAN_PATH as of AS_OF, using the saved usage and billing evidence in SOURCE_PATHS. In EXPERIMENT_PATHS, record each attempt's usage in units under its Backend plus Job ID or stable manual-attempt identity, with source identity, coverage, and retrieval timestamp. Keep superseded usage figures as dated history, and write no currency amounts into records. In plans/spend/LEDGER.md, give each finished attempt's open row its terminal outcome and actual, marking usage-estimated or unbilled actuals /provisional. When evidence changes a terminal figure, append a correction row that names the same attempt and carries the difference; never restate, renumber, or delete a row. Apply the recorded shared-cost attribution policy once and disclose unallocated or unknown charges. Report A, C, headroom against plans/spend/AUTHORITY.json, variance from the baseline in units, and pending settlement with its follow-up owner, or flag that one is missing. Do not edit AUTHORITY.json, launch, submit, retry, cancel, or rerun jobs, collect new scientific evidence, approve new work, or treat this update as execution approval.

A read-only results walkthrough may explain these figures from existing records. Writing corrections requires separate authorization such as the accounting request above. Any proposed follow-up returns to the plan workflow for a scientific and resource decision.