Repo2RLEnv

Yield and cost per task

Edit on GitHub
989new recipe exports

14 recipes · generation sample

$445.94generation accounted

API + compute · excludes held charges

6 / 14recipes with measured yield

A missing denominator stays unknown

Research recipes · September 14, 2026

Generation cost per export

Same accounting scope, different tasks and difficulty. These are observed generation results, not a quality or price ranking. Independent evaluation and unresolved charges are excluded.

Model API Estimated computeUSD per new export · bars start at zero

  • 476 new exports · API $0.069 + compute $0.091

  • 80 new exports · API $0.368 + compute $0.151

  • cli-gym$1.142

    5 new exports · API $0.709 + compute $0.434

  • swe-flow$0.301

    76 new exports · API $0.053 + compute $0.248

  • 77 new exports · API $0.542 + compute $0.167

  • seta-evol$0.559

    80 new exports · API $0.384 + compute $0.175

  • tmax$2.100

    35 new exports · API $1.681 + compute $0.419

  • 80 new exports · API $0.679 + compute $0.577

  • dataarc$0.332

    80 new exports · API $0.182 + compute $0.150

The repository pool groups swe-smith, r2e, swe-gen, swe-next, r2e-gym and scaler. Its shared worker bill cannot be allocated reliably to individual recipes. Use “API only” to compare their recorded model costs.

Keep the denominator in view. These charts cover generation only. Tasksmith acceptance and the CodeMidas / FrontierSmith campaigns include different review and rollout work; their separate cost tables follow below. Explore models, tokens and stage costs →

Evidence audited 2026-09-30. Recipe/Tasksmith samples were measured September 14, CodeMidas September 25, and FrontierSmith September 29, 2026. These are historical experiments on different inputs and checks, not a controlled price or quality ranking.

Read the stage costs, models, tokens and compute for each recipe, or the published inventory for dataset versions and quality labels.

Measurement definitions

TermDefinition
Candidate yieldExports divided by distinct recorded candidates in the same sample; retries do not become new candidates. The recipe can start counting before design screening.
Attempt yieldUsed explicitly for CodeMidas/FrontierSmith: exports or selections divided by recorded attempts; historical revisions and repeated seeds are identified.
Export / selected / acceptedA written Harbor bundle / a curated subset / a task that passed its named quality profile. These are different denominators.
Accounted costRecorded API-usage estimates plus attributable compute estimates, including unsuccessful attempts and repairs within the stated scope.
Held costUnresolved or reserved charges, reported separately from accounted spend; not evidence of a paid invoice.
Per-task costThe stated sample cost divided by its new outputs. Retained tasks and later validation must not silently enter the denominator.
n/aEvidence unavailable or not safely attributable. It never means free.

Recorded experiment spend

$1,586.71 accounted, plus $31.72 unresolved/reserved, across the disjoint scopes below. The combined accounting exposure is $1,618.43. This is a recovered experiment subtotal, not an invoice or a complete lifetime project bill. Older native-pipeline costs and interactive assistant usage are outside it.

Measurement scopeAccountedHeld separatelyWhat it paid for
Initial research pilots + Tasksmith development/evaluation$972.45$21.07291 recipe exports + 50 final Tasksmith tasks; includes failures, shared work and $22.43 prior spend
Research-recipe expansion$445.94$8.50989 new exports across 14 recipes; earlier retained tasks excluded
CodeMidas campaign$99.91$1.47100 curated tasks; includes historical pilots, generation and evaluation
FrontierSmith campaign$68.42$0.67100 selected tasks; includes pilot, expansion and sample evaluation

Do not add the Tasksmith expansion, original recipe pilots, or per-stage tables again: they are subsets of these rows. Parent-to-child budget transfers were excluded. The accounting detail explains scope, attribution and evidence.

Research recipes and Tasksmith generation

These 14 generation samples added 989 tasks to 291 retained recipe tasks. The final published recipe inventory is 1,280 tasks. Costs include failed generation and bounded repairs, but independent quality-pilot costs belong to the earlier program and are reported separately.

RecipeCandidate denominatorNew / final tasksExport yield
swe-smithn/a76 / 100n/a
r2en/a80 / 100n/a
swe-genn/a80 / 100n/a
swe-nextn/a80 / 100n/a
r2e-gymn/a80 / 100n/a
scalern/a80 / 100n/a
endless-terminals9980 / 10080.8%
cli-gymn/a5 / 25n/a
swe-flown/a76 / 100n/a
seta-seed2synth11277 / 10068.8%
seta-evol9280 / 10087.0%
tmax10435 / 5533.7%
terminalworld129380 / 1006.2%
dataarc8380 / 10096.4%

The first six repository recipes and CLI-Gym/SWE-flow lack a reliable deduplicated attempt denominator for these cost samples. Hitting a 100-task target is not 100% yield. TerminalWorld counted 1,293 recordings before suitability screening; 1,145 were screened out, leaving 148 candidates and 80 exports (54.1% after screening). SETA Evol includes one unfinished candidate; TMax includes eight.

RecipeAPI totalCompute totalHeldAPI / new taskCombined / new task
swe-smith$2.41n/a$0.00$0.032n/a
r2e$17.73n/a$1.50$0.222n/a
swe-gen$1.63n/a$0.00$0.020n/a
swe-next$7.51n/a$0.00$0.094n/a
r2e-gym$3.52n/a$0.75$0.044n/a
scaler$0.00n/a$0.00$0.000n/a
endless-terminals$29.47$12.05$0.00$0.368$0.519
cli-gym$3.54$2.17$0.00$0.709$1.142
swe-flow$4.05$18.86$0.00$0.053$0.301
seta-seed2synth$41.74$12.90$0.00$0.542$0.710
seta-evol$30.75$13.97$1.25$0.384$0.559
tmax$58.83$14.68$5.00$1.681$2.100
terminalworld$54.34$46.13$0.00$0.679$1.256
dataarc$14.53$12.04$0.00$0.182$0.332

The six repository recipes share $43.09 compute, without a defensible per-recipe allocation. Their combined 476 new exports cost $75.90, or $0.159/task. SCALER has zero generation API spend, not zero runtime cost.

Most expansion calls used claude-sonnet-4-6; TMax also tried gpt-5.4-mini. The earlier pilots used Sonnet and Opus in different proportions. Repository expansion started with six Modal workers, then used 26 Daytona workers; terminal/reconstruction expansion used Daytona. All were CPU workers. See the per-recipe detail for stages, exact model identifiers, calls and tokens.

Tasksmith and optional evaluation

The final cohort contains 50 verified tasks and 19 full Sonnet solves. It came from assisted development; neither the archived PR inventory nor the 50-task target establishes unattended conversion yield.

Cost scopeAccountedDenominatorCost per task
Direct operations attributed to the final 50$649.5650 accepted tasks$12.991
Expansion, including retained-task repair/revalidation$392.0126 additional accepted tasks$15.077

These are overlapping views, not additive bills. Direct attribution excludes shared/unattributed costs and rejected PRs; expansion cost includes unsuccessful attempts and work on previously retained tasks. Do not call either a complete marginal production price.

Expansion componentTotalPer added accepted task
Authoring model$63.91$2.458
Review and repair models$86.06$3.310
Blind solver model$21.20$0.815
Compute, including validation$220.83$8.494

A further $4.98 is held in that expansion scope. Authoring and blind solving used Sonnet 4.6; the larger program also used Opus 4.6 and GPT-6 Astra for review/repair. The final task requirements comprise 39 CPU tasks and 11 L4 GPU tasks. Cloud controllers, learner resources and model API inference are separate costs. See Tasksmith accounting and interventions.

CodeMidas generation and evaluation

Measured 2026-09-25 using GPT-6 Luna/Sol and Daytona. This whole-campaign sample includes historical pilots, failed construction, independent review, rollouts and compute. It is not comparable to generation-only prices above; interactive assistant usage is excluded.

MeasureResult
Construction yield128/213 (60.1%)
Ordinary review yield101/128 (78.9%)
Recorded API usage$88.20
Conservative compute estimate$11.71
Combined accounted$99.91
Unknown API billing reserved$1.47
Accounted per export$0.78
Accounted per reviewed task$0.99
Accounted per curated task$1.00

Compute is an estimate, not an invoice. The construction denominator includes two candidates stopped after the goal was met. The 100 curated tasks remain adversarial-blocked; there is no cost per fully accepted task. See stage costs and limitations.

FrontierSmith optimization synthesis

Measured 2026-09-29 with OpenAI gpt-6-sol and Daytona CPU workers. 153 candidate attempts across 152 distinct seeds produced 101 initial exports; 100 were selected after construction and collection review. Initial export yield was 66.0%; final selection was 65.4%.

Cost componentWhole collectionPer selected task
Recorded API usage$64.66$0.647
Estimated compute$3.76$0.038
Accounted combined$68.42$0.684
Unknown API charges reserved separately$0.67$0.007
Measurement scopeAttemptsInitial exportsSelectedCombined cost
Development pilot151110$5.61
Expansion and collection review1389090$62.81
Model-call stageRecorded API cost
Baseline and sampled programs$25.52
Blind agent rollouts$2.86
Finished-task contract review$1.58
Collection diversity review$0.51
Task formulation and review$5.70
Test infrastructure and bounded repair$24.10
Other development calls$0.56
Post-construction generator repair$0.15
Original seed descriptions$0.54
Sample algorithm diversity review$3.14

Costs include original seed authoring, failed candidates, bounded repair, construction trials, collection reviews and sample rollouts. The ten-task development pilot used evolving checks; the expansion used the recorded fixed recipe. This is a measured assisted campaign, not a guarantee of future yield. 21/100 selected bundles have blind rollout evidence; full quality acceptance remains pending.

Interactive assistant usage is excluded. Model costs use recorded usage and the configured rate table. Compute uses worker lifecycle duration and the Daytona resource rates, with no free-tier deduction; neither amount is an invoice reconciliation. All workers were terminated. Lost API responses retain their conservative reservations rather than being counted as free or silently retried. See the collection audit and machine-readable results.

Native pipeline measurements

These May–July 2026 runs have less complete accounting. Recorded synthesis cost excludes bootstrap, compute and solver evaluation; it is not comparable to the total generation costs above. n/a means unavailable. See historical results for the evidence and sample boundaries.

PipelineRetained tasksMeasured generation yieldRecorded synthesis / taskScope
pr_diff181n/an/aNo complete generation-cost ledger recovered. Unavailable does not mean zero.
pr_runtime100n/an/aNo complete generation-cost ledger recovered. Unavailable does not mean zero.
commit_runtime100n/an/aNo complete generation-cost ledger recovered. Unavailable does not mean zero.
code_instruct100100/136 (73.5%)$0.038Run-cumulative synthesis counters: sum one final maximum per repo, not every task. Includes retries through the last export; excludes bootstrap, compute, rollouts and earlier development.
equivalence_tests100n/a≥ $0.025Lower bound from seven productive runs only. Excludes zero-output runs, later failed attempts, bootstrap, compute and rollouts; not an all-in task price.
cve_patches19n/an/aNo complete generation-cost ledger recovered. Unavailable does not mean zero.

Code-instruct's complete generation log records 136 candidates, correcting the earlier 132-candidate claim. Equivalence-test logs contain at least 200 candidates, including zero-output runs, but several runs lack a final summary; its overall yield is unavailable. Its $0.025/task figure is only a lower bound from productive-run counters.

Measurement source

Tables are generated from pipeline measurements, reconciled experiment evidence, and the FrontierSmith collection report. The accounting guide describes the source fingerprints and refresh procedure. Raw generated tasks, receipts and campaign scripts remain outside Git.

python3 docs/_tools/generate_metrics.py
python3 docs/_tools/generate_metrics.py --check

On this page