Repo2RLEnv

Experiment accounting: models, compute and stage costs

Edit on GitHub

Audited 2026-09-30 from saved receipts and frozen reports. The yield and cost overview gives the comparable scope definitions; this page explains what the experiments actually consumed. No paid reruns were needed for this audit.

How the total is reconciled

$1,586.71 accounted, plus $31.72 unresolved/reserved, across the disjoint scopes below. The combined accounting exposure is $1,618.43. This is a recovered experiment subtotal, not an invoice or a complete lifetime project bill. Older native-pipeline costs and interactive assistant usage are outside it.

Measurement scopeAccountedHeld separatelyWhat it paid for
Initial research pilots + Tasksmith development/evaluation$972.45$21.07291 recipe exports + 50 final Tasksmith tasks; includes failures, shared work and $22.43 prior spend
Research-recipe expansion$445.94$8.50989 new exports across 14 recipes; earlier retained tasks excluded
CodeMidas campaign$99.91$1.47100 curated tasks; includes historical pilots, generation and evaluation
FrontierSmith campaign$68.42$0.67100 selected tasks; includes pilot, expansion and sample evaluation

Do not add the Tasksmith expansion, original recipe pilots, or per-stage tables again: they are subsets of these rows. Parent-to-child budget transfers were excluded. The accounting detail explains scope, attribution and evidence.

The initial program ledger contains 5,004 operations and $950.02 accounted, plus $22.43 recorded as prior external spend. It includes the initial recipe pilots, Tasksmith, independent quality work and shared costs. The extra $22.43 has no recovered stage or provider split and is retained as an external prior amount.

Research expansion adds $445.94 in separate leaf ledgers. The repository controller also carries a $21.29 transfer for terminal finishing; the audit excluded that parent row and counted the child costs once. Operation IDs in those expansion ledgers were checked for overlap with the frozen initial ledger. CodeMidas and FrontierSmith have separate campaign ledgers; FrontierSmith's pilot is included exactly once.

This subtotal excludes unledgered interactive assistance, historical native pipeline expenses without a complete ledger, and any provider charge not represented by the selected records. Invoice-level reconciliation has not been performed. Do not divide it by the total published inventory: its scope includes development, discarded work, retained-task repair and different evaluation coverage.

Initial program cost allocation

Attributed populationAccountedHeldInterpretation
Final 50 Tasksmith tasks$649.56$9.53Directly attributable operations only
Other Tasksmith PRs$56.02$0.79Attempts outside the final 50
Recipe candidates$58.51$8.00Initial generation/evaluation attributable to tasks
Shared or unattributed$185.92$2.75Cannot safely allocate to a task or recipe

The direct-allocation and stage views below partition the same $950.02 ledger. Neither includes the separate $22.43 prior amount.

Initial program stageAccountedHeldOperations
authoring and investigation$99.57$2.322,211
blind solver model$36.41$12.00176
compute estimates$413.15$3.00593
other unattributed$142.16$2.751,142
quality review and repair$258.73$1.00882

“Other/unattributed” includes recipe authoring and unclassified operations; it is not all compute. The initial ledger spans multiple experiments, so its compute/model ratio is not Tasksmith's isolated generation ratio.

Tasksmith development and quality work

The archived inventory contains 56 distinct PRs and 50 final accepted tasks. Six other PRs and earlier prototypes remain outside that cohort. The archive does not establish a complete pre-screening denominator, so this is not a claim of 50/56 unattended yield. All 50 final blind solver configurations used anthropic/claude-sonnet-4-6; 19 reached full reward.

Model role in initial programRecorded modelAccountedHeldOperations
authoring and investigationanthropic/claude-sonnet-4-6$99.57$2.322211
other unattributedanthropic/claude-opus-4-6$104.85$2.25816
other unattributedanthropic/claude-sonnet-4-6$19.29$0.00301
other unattributedopenai/gpt-5-mini$0.02$0.506
quality review and repairanthropic/claude-opus-4-6$65.00$0.00245
quality review and repairanthropic/claude-sonnet-4-6$52.51$0.00323
quality review and repairopenai/gpt-6-astra$141.21$1.00314

This model table covers explicitly model-labelled ledger rows. Blind Harbor trial rows are a separate stage in the program table and are not re-added here.

Final task requirementTasksCPUMemory MiBGPU
CPU task3912048None
GPU task1281922 × L4
GPU task64163841 × L4
GPU task44163842 × L4

These are learner task requirements, not the controller sizes. Tasksmith used Modal CPU controllers and native Modal L4 execution for GPU tasks. Model inference was an API charge. The recovered cost scope does not provide a defensible per-task split of CPU, GPU, image build and shared idle costs.

Recorded assistance in final 50Tasks
design and bootstrap preparation4
no confirmed task specific intervention15
operational or review continuation1
task or verifier edit28
validation control repair2

Automated repair evidence exists for 35 tasks. This overlaps the assistance categories; it is not 35 additional tasks. “No confirmed intervention” means history is incomplete, not proven hands-off generation. Per-task direct costs, resource requirements, solver rewards and assistance categories are retained in the sanitized accounting data.

Cloud resource measurements

Each row below counts saved worker records with their configured resources. Worker-hours include setup, builds, execution, waiting and idle time. Concurrent worker-hours add together; they are neither wall-clock completion time nor CPU utilization. Receipts with no stop timestamp are excluded from hours and exposed in the coverage count.

Initial recipe/development workers

ProviderCPUsMemory MiBDisk GBWorker recordsWorker-hours
daytona2409610411.84 (3/4 timed)
daytona24096n/a1n/a (0/1 timed)
modal240961047.45 (4/4 timed)
modal24096n/a10.93 (1/1 timed)
modal48192101122.51 (11/11 timed)

repository-expansion

ProviderCPUsMemory MiBDisk GBWorker recordsWorker-hours
daytona48192102628.23 (26/26 timed)
modal481921062.27 (6/6 timed)

terminal-expansion

ProviderCPUsMemory MiBDisk GBWorker recordsWorker-hours
daytona4819210126.68 (12/12 timed)

swe-flow-expansion

ProviderCPUsMemory MiBDisk GBWorker recordsWorker-hours
daytona4819210182.59 (18/18 timed)

seta-seed2synth-expansion

ProviderCPUsMemory MiBDisk GBWorker recordsWorker-hours
daytona4819210108.73 (10/10 timed)

seta-evol-expansion

ProviderCPUsMemory MiBDisk GBWorker recordsWorker-hours
daytona4819210125.94 (12/12 timed)

tmax-expansion

ProviderCPUsMemory MiBDisk GBWorker recordsWorker-hours
daytona48192101111.09 (11/11 timed)

terminalworld-expansion

ProviderCPUsMemory MiBDisk GBWorker recordsWorker-hours
daytona4819210439.44 (43/43 timed)

dataarc-expansion

ProviderCPUsMemory MiBDisk GBWorker recordsWorker-hours
daytona4819210113.12 (11/11 timed)

CodeMidas

ProviderCPUsMemory MiBDisk GBWorker recordsWorker-hours
daytona4819210411.62 (4/4 timed)
daytona48192301n/a (0/1 timed)

FrontierSmith, pilot + expansion

ProviderCPUsMemory MiBDisk GBWorker recordsWorker-hours
daytona24096101322.56 (13/13 timed)

How compute was priced

Historical recipe receipts generally used a 4-CPU / 8-GiB / 10-GB Daytona rate of $0.00009215 per second ($0.33174/hour) plus an explicit $1 image-build allowance per worker. The allowance is a conservative accounting choice, not a measured provider build charge. Short or failed workers can therefore have high cost per task. Modal estimates used their recorded resource/lifetime basis; they are not silently repriced with Daytona rates.

CodeMidas used the same nominal Daytona hourly rate with a 2× safety factor and $1 build allowance on the shown cost receipt. FrontierSmith used 2-CPU / 4-GiB workers and lifecycle estimates without a free-tier deduction. These assumptions differ, so compute totals are not a hardware benchmark. Stored example cost receipts and their hashes are included in the accounting data; no claim is made about today's provider prices.

Repository expansion control coverage

The completion audit matched executable bundle hashes to baseline/reference receipts for the new exports. Retained pilot tasks were outside this check. A missing match means evidence was not established by this audit, not that the task necessarily failed.

RecipeNew exportsMatching baseline/reference evidence
r2e8080
r2e-gym8080
scaler800
swe-gen8080
swe-next8080
swe-smith7668

Per-recipe generation detail

The initial pilot API costs below are already inside the initial program total. Expansion costs are disjoint from those pilot costs. Full-lifecycle compute cannot be allocated per recipe from the shared initial workers; no equal-share estimate is presented as measured spend. Stage buckets include retries. Exact historical prompt text and source behavior may differ from the current release.

swe-smith

swe-smith pipeline. Earlier pilot: 24 exports, $0.65 generation API, $0.00 held. Expansion: 76 new exports, $2.41 API and n/a attributable compute. Final inventory: 100. Compute pool: repository-expansion.

Combined recorded pilot + expansion authoring API: $3.05, or $0.031 per final task. This excludes independent quality work and all compute.

Pilot author models: anthropic/claude-opus-4-6 (1 calls), anthropic/claude-sonnet-4-6 (27 calls).

Expansion stageAccounted APIHeldOperations
Issue/instruction authoring$2.41$0.00119
Expansion modelAPI costCalls with token receipts / settled callsInput tokensOutput tokens
anthropic/claude-sonnet-4-6$2.41119 / 119366,95587,107

Earlier quality pilot: 5/5 baseline/reference control pairs passed; 5/5 original Sonnet attempts scored one; 5/5 selected after review/revision. Those are different outcomes and apply only to the five-task pilot, not the expanded dataset.

Mutation and instruction generation are cheap on cached small repositories. Not every expansion export retained matching nop/reference receipts: 68 of 76 did in the completion audit. Five earlier quality-pilot tasks are not a gate for the final 100.

r2e

r2e pipeline. Earlier pilot: 20 exports, $1.46 generation API, $0.00 held. Expansion: 80 new exports, $17.73 API and n/a attributable compute. Final inventory: 100. Compute pool: repository-expansion.

Combined recorded pilot + expansion authoring API: $19.19, or $0.192 per final task. This excludes independent quality work and all compute.

Pilot author models: anthropic/claude-sonnet-4-6 (47 calls).

Expansion stageAccounted APIHeldOperations
Public specification$2.13$0.0080
Equivalence tests and repairs$15.60$1.50287
Expansion modelAPI costCalls with token receipts / settled callsInput tokensOutput tokens
anthropic/claude-sonnet-4-6$17.73365 / 3651,724,935836,703

Earlier quality pilot: 5/5 baseline/reference control pairs passed; 5/5 original Sonnet attempts scored one; 5/5 selected after review/revision. Those are different outcomes and apply only to the five-task pilot, not the expanded dataset.

Equivalence-test authoring and retries dominate API spend. Passing a frozen reference only establishes behavior on the generated tests; sampled solver success does not prove exhaustive equivalence.

swe-gen

swe-gen pipeline. Earlier pilot: 20 exports, $0.38 generation API, $0.00 held. Expansion: 80 new exports, $1.63 API and n/a attributable compute. Final inventory: 100. Compute pool: repository-expansion.

Combined recorded pilot + expansion authoring API: $2.01, or $0.020 per final task. This excludes independent quality work and all compute.

Pilot author models: anthropic/claude-sonnet-4-6 (20 calls).

Expansion stageAccounted APIHeldOperations
Instruction authoring$1.63$0.0081
Expansion modelAPI costCalls with token receipts / settled callsInput tokensOutput tokens
anthropic/claude-sonnet-4-6$1.6381 / 81280,79252,817

Earlier quality pilot: 5/5 baseline/reference control pairs passed; 5/5 original Sonnet attempts scored one; 5/5 selected after review/revision. Those are different outcomes and apply only to the five-task pilot, not the expanded dataset.

Reuses existing PR fixes and tests, so the measured API work is mostly instructions. Repository discovery, dependency setup and shared worker costs prevent treating the API figure as an all-in price.

swe-next

swe-next pipeline. Earlier pilot: 20 exports, $1.22 generation API, $0.00 held. Expansion: 80 new exports, $7.51 API and n/a attributable compute. Final inventory: 100. Compute pool: repository-expansion.

Combined recorded pilot + expansion authoring API: $8.73, or $0.087 per final task. This excludes independent quality work and all compute.

Pilot author models: anthropic/claude-sonnet-4-6 (20 calls).

Expansion stageAccounted APIHeldOperations
PR task instructions$7.51$0.0085
Expansion modelAPI costCalls with token receipts / settled callsInput tokensOutput tokens
anthropic/claude-sonnet-4-6$7.5185 / 852,122,46176,468

Earlier quality pilot: 5/5 baseline/reference control pairs passed; 4/5 original Sonnet attempts scored one; 5/5 selected after review/revision. Those are different outcomes and apply only to the five-task pilot, not the expanded dataset.

Uses existing PR tests and fixes. Input context is large relative to instruction output; test discovery and offline setup can fail before export.

r2e-gym

r2e-gym pipeline. Earlier pilot: 20 exports, $1.31 generation API, $0.00 held. Expansion: 80 new exports, $3.52 API and n/a attributable compute. Final inventory: 100. Compute pool: repository-expansion.

Combined recorded pilot + expansion authoring API: $4.83, or $0.048 per final task. This excludes independent quality work and all compute.

Pilot author models: anthropic/claude-sonnet-4-6 (20 calls).

Expansion stageAccounted APIHeldOperations
Commit task instructions$3.52$0.7581
Expansion modelAPI costCalls with token receipts / settled callsInput tokensOutput tokens
anthropic/claude-sonnet-4-6$3.5280 / 80863,44262,213

Earlier quality pilot: 5/5 baseline/reference control pairs passed; 5/5 original Sonnet attempts scored one; 5/5 selected after review/revision. Those are different outcomes and apply only to the five-task pilot, not the expanded dataset.

Commit mining supplies the implementation and tests. The missing-response reservation is preserved; do not count an unknown response as free.

scaler

scaler pipeline. Earlier pilot: 20 exports, $0.00 generation API, $0.00 held. Expansion: 80 new exports, $0.00 API and n/a attributable compute. Final inventory: 100. Compute pool: repository-expansion.

Combined recorded pilot + expansion authoring API: $0.00, or $0.000 per final task. This excludes independent quality work and all compute.

Pilot author models: none; deterministic generation.

Expansion stageAccounted APIHeldOperations
No authoring calls$0.00$0.000
Expansion modelAPI costCalls with token receipts / settled callsInput tokensOutput tokens
None$0.000 / 000

Earlier quality pilot: 5/5 baseline/reference control pairs passed; 4/5 original Sonnet attempts scored one; 4/5 selected after review/revision. Those are different outcomes and apply only to the five-task pilot, not the expanded dataset.

Deterministic generators require no authoring LLM. Shared compute and later solver evaluation still cost money; construction format checks are not an independent semantic audit.

endless-terminals

endless-terminals pipeline. Earlier pilot: 20 exports, $10.77 generation API, $0.00 held. Expansion: 80 new exports, $29.47 API and $12.05 attributable compute. Final inventory: 100. Compute pool: terminal-expansion.

Combined recorded pilot + expansion authoring API: $40.23, or $0.402 per final task. This excludes independent quality work and all compute.

Pilot author models: anthropic/claude-opus-4-6 (105 calls).

Expansion stageAccounted APIHeldOperations
Artifact/test building and repair$18.26$0.00162
Design and suitability screening$11.21$0.00289
Expansion modelAPI costCalls with token receipts / settled callsInput tokensOutput tokens
anthropic/claude-sonnet-4-6$29.47451 / 4512,588,2431,446,891

Earlier quality pilot: 5/5 baseline/reference control pairs passed; 5/5 original Sonnet attempts scored one; 5/5 selected after review/revision. Those are different outcomes and apply only to the five-task pilot, not the expanded dataset.

Terminal design and environment building include rejected attempts and repair. The initial pilot used Opus, the expansion Sonnet, so pooling them hides a model change.

cli-gym

cli-gym pipeline. Earlier pilot: 20 exports, $22.99 generation API, $1.00 held. Expansion: 5 new exports, $3.54 API and $2.17 attributable compute. Final inventory: 25. Compute pool: terminal-expansion.

Combined recorded pilot + expansion authoring API: $26.53, or $1.061 per final task. This excludes independent quality work and all compute.

Pilot author models: anthropic/claude-opus-4-6 (175 calls).

Expansion stageAccounted APIHeldOperations
Environment break/repair proposals$3.54$0.0048
Expansion modelAPI costCalls with token receipts / settled callsInput tokensOutput tokens
anthropic/claude-sonnet-4-6$3.5448 / 48874,79761,281

Earlier quality pilot: 5/5 baseline/reference control pairs passed; 0/5 original Sonnet attempts scored one; 5/5 selected after review/revision. Those are different outcomes and apply only to the five-task pilot, not the expanded dataset.

Only five expansion exports were measured, added to 20 retained pilot tasks. The first five sampled Sonnet rollouts scored zero before repairs; later repairs and reruns are separate evidence, not initial successes.

swe-flow

swe-flow pipeline. Earlier pilot: 24 exports, $1.12 generation API, $0.00 held. Expansion: 76 new exports, $4.05 API and $18.86 attributable compute. Final inventory: 100. Compute pool: swe-flow-expansion.

Combined recorded pilot + expansion authoring API: $5.17, or $0.052 per final task. This excludes independent quality work and all compute.

Pilot author models: anthropic/claude-sonnet-4-6 (48 calls).

Expansion stageAccounted APIHeldOperations
Docstring reconstruction$0.86$0.0076
Specification reconstruction$3.19$0.0076
Expansion modelAPI costCalls with token receipts / settled callsInput tokensOutput tokens
anthropic/claude-sonnet-4-6$4.05152 / 152667,339136,484

Earlier quality pilot: 5/5 baseline/reference control pairs passed; 4/5 original Sonnet attempts scored one; 3/5 selected after review/revision. Those are different outcomes and apply only to the five-task pilot, not the expanded dataset.

Reuses source for reconstruction and generates docstrings/specifications. Compute dominates the expansion bill; two published instruction issues remain labelled needs_repair.

seta-seed2synth

seta-seed2synth pipeline. Earlier pilot: 23 exports, $20.99 generation API, $0.00 held. Expansion: 77 new exports, $41.74 API and $12.90 attributable compute. Final inventory: 100. Compute pool: seta-seed2synth-expansion.

Combined recorded pilot + expansion authoring API: $62.73, or $0.627 per final task. This excludes independent quality work and all compute.

Pilot author models: anthropic/claude-opus-4-6 (57 calls), anthropic/claude-sonnet-4-6 (74 calls).

Expansion stageAccounted APIHeldOperations
Artifact/test building and repair$28.58$0.00241
Design and suitability screening$6.83$0.00112
Infrastructure/contract review$6.33$0.00264
Expansion modelAPI costCalls with token receipts / settled callsInput tokensOutput tokens
anthropic/claude-sonnet-4-6$41.74617 / 6175,478,9641,686,904

Earlier quality pilot: 5/5 baseline/reference control pairs passed; 4/5 original Sonnet attempts scored one; 3/5 selected after review/revision. Those are different outcomes and apply only to the five-task pilot, not the expanded dataset.

Design, build and review may repeat before export. Three of five earlier quality-pilot tasks were selected after revision; the published 100 are not all independently validated.

seta-evol

seta-evol pipeline. Earlier pilot: 20 exports, $9.77 generation API, $0.00 held. Expansion: 80 new exports, $30.75 API and $13.97 attributable compute. Final inventory: 100. Compute pool: seta-evol-expansion.

Combined recorded pilot + expansion authoring API: $40.52, or $0.405 per final task. This excludes independent quality work and all compute.

Pilot author models: anthropic/claude-opus-4-6 (25 calls), anthropic/claude-sonnet-4-6 (43 calls).

Expansion stageAccounted APIHeldOperations
Artifact/test building and repair$20.46$1.25154
Design and suitability screening$6.55$0.0092
Infrastructure/contract review$3.75$0.00146
Expansion modelAPI costCalls with token receipts / settled callsInput tokensOutput tokens
anthropic/claude-sonnet-4-6$30.75391 / 3913,825,3291,284,989

Earlier quality pilot: 5/5 baseline/reference control pairs passed; 5/5 original Sonnet attempts scored one; 3/5 selected after review/revision. Those are different outcomes and apply only to the five-task pilot, not the expanded dataset.

Evolution can produce ambiguous instructions or weak state checks. One candidate was unfinished and one model charge remained unresolved at stop.

tmax

tmax pipeline. Earlier pilot: 20 exports, $22.54 generation API, $0.00 held. Expansion: 35 new exports, $58.83 API and $14.68 attributable compute. Final inventory: 55. Compute pool: tmax-expansion.

Combined recorded pilot + expansion authoring API: $81.37, or $1.479 per final task. This excludes independent quality work and all compute.

Pilot author models: anthropic/claude-opus-4-6 (150 calls).

Expansion stageAccounted APIHeldOperations
Artifact/test building and repair$38.22$5.00241
Design and suitability screening$13.99$0.00309
Infrastructure/contract review$6.62$0.00214
Expansion modelAPI costCalls with token receipts / settled callsInput tokensOutput tokens
anthropic/claude-sonnet-4-6$58.53733 / 7336,698,3492,562,447
openai/gpt-5.4-mini$0.3027 / 27146,47349,404

Earlier quality pilot: 5/5 baseline/reference control pairs passed; 4/5 original Sonnet attempts scored one; 1/5 selected after review/revision. Those are different outcomes and apply only to the five-task pilot, not the expanded dataset.

Multi-pass design/build/review makes failed attempts expensive. The final target was reduced to 55; eight expansion candidates were unfinished. A small unsuccessful GPT-5.4-mini comparison is included in cost.

terminalworld

terminalworld pipeline. Earlier pilot: 20 exports, $12.31 generation API, $1.25 held. Expansion: 80 new exports, $54.34 API and $46.13 attributable compute. Final inventory: 100. Compute pool: terminalworld-expansion.

Combined recorded pilot + expansion authoring API: $66.66, or $0.667 per final task. This excludes independent quality work and all compute.

Pilot author models: anthropic/claude-opus-4-6 (239 calls).

Expansion stageAccounted APIHeldOperations
Artifact/test building and repair$20.60$0.00498
Design and suitability screening$29.92$0.001679
Infrastructure/contract review$3.81$0.00203
Additional tests$0.02$0.001
Expansion modelAPI costCalls with token receipts / settled callsInput tokensOutput tokens
anthropic/claude-sonnet-4-6$54.342381 / 238112,131,6991,196,456

Earlier quality pilot: 5/5 baseline/reference control pairs passed; 4/5 original Sonnet attempts scored one; 1/5 selected after review/revision. Those are different outcomes and apply only to the five-task pilot, not the expanded dataset.

Recorded sessions often fail suitability screening: 1,145 of 1,293 candidates were screened out. That explains much of the 6.2% raw yield. Three published verifier gaps remain labelled needs_repair.

dataarc

dataarc pipeline. Earlier pilot: 20 exports, $18.55 generation API, $0.00 held. Expansion: 80 new exports, $14.53 API and $12.04 attributable compute. Final inventory: 100. Compute pool: dataarc-expansion.

Combined recorded pilot + expansion authoring API: $33.08, or $0.331 per final task. This excludes independent quality work and all compute.

Pilot author models: anthropic/claude-opus-4-6 (63 calls).

Expansion stageAccounted APIHeldOperations
Artifact/test building and repair$11.80$0.00113
Infrastructure/contract review$2.73$0.00111
Expansion modelAPI costCalls with token receipts / settled callsInput tokensOutput tokens
anthropic/claude-sonnet-4-6$14.53224 / 2241,607,392647,246

Earlier quality pilot: 5/5 baseline/reference control pairs passed; 5/5 original Sonnet attempts scored one; 0/5 selected after review/revision. Those are different outcomes and apply only to the five-task pilot, not the expanded dataset.

High construction yield does not establish task quality. The published cohort used recipe v1; current v2 differs. None of five earlier quality-pilot tasks was selected under that pilot review despite five initial solver rewards of one.

CodeMidas stage accounting

GPT-6 Luna handled source/task construction and ordinary solving; GPT-6 Sol handled construction review, independent rollout review and difficulty screening. Daytona workers ran CPU tasks. Costs include 233 construction attempts with historical revisions, while the current-collection yield uses 213 attempts and 128 unique exports. Reporting the narrower yield does not remove historical costs.

StageAccountedHeldLedger operations
adversarial attempt$0.39$0.008
compute estimate$11.71$0.005
construction design$2.58$0.001229
construction review$28.32$0.89386
construction tests$4.14$0.031367
frontier screen$23.14$0.00428
ordinary solver$2.13$0.00546
other$0.00$0.005
rollout review$27.49$0.551368

The adversarial-attempt spend does not mean those checks succeeded: all 100 curated tasks remain blocked for full-method acceptance. Review found concrete defects in 26/128 exports. CodeMidas results give the solver and source breakdown. Token totals are not reconciled across every author and agent trace here, so no whole-campaign token figure is claimed.

FrontierSmith stage accounting

All authoring, seed creation, reviews and sampled rollouts used gpt-6-sol in separate contexts. The 100-task selection cost includes the original ten-task pilot, original seed authoring, 153 candidate attempts, two post-construction generator repairs and collection review. It is not the cost of a single successful generation run.

Model stageAPI cost
baseline and sampled programs$25.52
blind rollouts$2.86
collection contract review$1.58
collection diversity review$0.51
formulation and review$5.70
infrastructure and bounded repair$24.10
other development calls$0.56
post construction repair$0.15
seed authoring$0.54
semantic diversity review$3.14

Add $3.76 estimated compute to $64.66 API usage for $68.42 accounted; keep $0.67 held separately. Baseline/sample generation and test infrastructure dominate the cost. All 100 are construction-checked; 24 final tasks had blind rollout attempts and 21 completed feasibly. None has full optimization-aware quality acceptance. References are sampled feasible programs, not proven optima. Collection details retain those limits. A unified token total across author and Harbor-agent receipts is not claimed.

Historical native pipeline accounting

The native datasets predate the campaign ledgers above. A fresh read of archived task metadata recovered the generation model for three 100-task cohorts. This names the recorded model; it does not recover missing token, retry or compute charges.

Native pipelineCohort metadataRecorded synthesis API total
commit_runtimeanthropic/claude-sonnet-4-6 on 100 tasksn/a
code_instructanthropic/claude-sonnet-4-6 on 100 tasks$3.775929
equivalence_testsanthropic/claude-sonnet-4-6 on 100 tasks$2.512044

Code-instruct's total sums the last run-cumulative counter for each of five repositories. Equivalence tests include only productive-run counters, so its total is a lower bound. Commit-runtime has no complete synthesis ledger despite its model stamp. These partial amounts are excluded from the program subtotal to preserve its stated scope.

PR-diff, PR-runtime and CVE-patch costs and generation-model coverage remain unavailable for their exact selected inventories. Do not infer an old run's model from today's example configuration. Historical solver samples include Sonnet 4.6, GPT-5.3-Codex and Qwen3.6-35B-A3B, on different tasks; their small model-cost samples and missing setup/compute charges are documented in native results.

Evidence and refresh

The portable audit data records SHA-256 fingerprints of the saved reports, ledgers and worker receipts, grouped costs and selected sanitized task metadata. Paths are relative evidence identifiers, not required checkout files. The pipeline summary supplies the published sample definitions and native history; the FrontierSmith report supplies the final collection identity.

To refresh: identify the exact output cohort; freeze the source reports; exclude transfer envelopes and duplicate/revision rows; retain unknown-call holds; reconcile each API stage and resource group; then update the portable summaries. Do not sum run-cumulative per-task counters or multiply an old unit cost by a newer dataset size.

The generator validates totals, per-model/stage sums, token coverage, source references and transfer exclusions before writing pages. A clean clone can rebuild the documentation using only the committed summaries:

python3 docs/_tools/generate_metrics.py
python3 docs/_tools/generate_metrics.py --check

Token counts are reported input/output counts on recovered settled receipts, not a new price calculation. Cache fields are retained separately and must not be added twice. Uncertain requests have no usable token totals. The recipe expansion receipts used recorded LiteLLM estimates; other campaigns used their configured rate tables or agent-reported costs. None is presented as an invoice.

On this page