commit_runtime¶
R2E-Gym SWE-GEN-style PR mining: walk commits, not PRs. Trades signal quality for yield, and catches drive-by fixes that never went through a PR — repos that squash-merge or commit directly to main aren't reachable from pr_runtime at all.
As of v0.8.4 the problem statement is LLM-synthesized (synthesize_with_llm, default on): the commit/issue text is rewritten into a clean, symptom-focused problem statement with the solution stripped out. This fixed the two failure modes of raw commit text — leakage (changelog bullets naming the fix → gameable) and thinness (title-only subjects → unsolvable). A min_problem_statement_words floor drops near-empty commits, and max_pass_to_pass (default 50) caps the regression set so whole-suite P2P doesn't inflate flaky-reward risk. A 100-env audit went from ~33% → 100% clean instructions; Opus solves the sampled tasks (non-gameable: tasks that scored 1.0 only when the fix leaked now require a real solve).
Reference dataset: AdithyaSK/repo2rlenv-commit-runtime — 100 oracle-verified envs (Python + Go). The original …commit-runtime-test (52 envs, pre-synthesis) is kept for comparison.
| Status | stable (v0.8.4) |
| Sandbox at generation | Yes — reuses the bootstrap image from pr_runtime |
| LLM use | bootstrap-time (one-time env build, cached) + one synthesis call per emitted task to write the leak-free problem statement (synthesize_with_llm). |
| Reward | Graded F2P/P2P (reward = f2p_rate × p2p_rate) via the in-container verifier, identical to pr_runtime. Tracked / command_resolved / eval_grade split documented in pr_runtime. |
| Languages | Any (commit_runtime inherits language-agnostic bootstrap from pr_runtime) |
| Inspired by | R2E-Gym (SWE-GEN) (Jain et al., COLM '25) |
Why commits, not PRs¶
R2E-Gym's headline finding: "instead of using human-written PRs, good-quality execution environments can directly be curated from commits." Commit-based curation:
- No PR-review bottleneck. Works on any repo with commit history, including ones that never use PRs (research / internal / solo-maintained).
- Larger candidate pool. 3-10× bigger than the PR list for most repos.
- Noisier signal. No reviewer signed off; filters have to do all the work.
commit_runtime is a sibling of pr_runtime, not a replacement. Each works best on different repo shapes:
| Repo style | Better fit |
|---|---|
| Squash-merge + PR-first (pallets, Django, etc.) | pr_runtime — every fix is a PR; commits look like a wall of merge commits |
| Direct-commit + squash-merge (Go projects, single-maintainer crates, internal repos) | commit_runtime — fixes land as plain commits, not behind PR merges |
Arc 3's 52-env dataset is Go-heavy for exactly this reason: urfave/cli (15), gin (8), mux (6) are commit-friendly; pallets/click (7) is the Python exception with enough non-PR commits to mine.
Algorithm¶
git clone --depth N(defaultclone_depth=200; bump for monthly mining).git log --first-parent <since>..<until>⇒ candidate commit list.- Metadata filter (
_metadata_filter): - Drop merge commits (
skip_merge_commits) - Drop excluded authors (bots)
- Drop short messages (<
min_message_words, default 5) - Reject non-bugfix conventional-commit types (
chore:/docs:/feat:/refactor:/style:/test:/ci:/build:/perf:/revert:) - Require a bugfix-positive signal:
fix:prefix ORCloses #Nissue trailer OR a bugfix keyword in the subject (fix/bug/regression/crash/broken/incorrect/wrong/fail/ …) git show <sha>⇒ split into(source_patch, test_patch)using_split_patch_and_test_patchfrompr_runtime.- Structural filter (
_structural_filter): skip CI-only diffs, sweeping refactors (file count >max_source_files_per_commit), and commits without ≥1 new test function (require_new_test_funcs). - Validate inside the bootstrap sandbox (
pr_runtime'svalidate_prharness): pre-fix run discoversFAIL_TO_PASS+PASS_TO_PASSsets; post-fix run confirms the flip. - Build instruction (
build_instruction_from_commit): - If
Closes #Nis present ⇒ fetch the GitHub issue body viagithub.fetch_issueand use it as the problem statement. - Otherwise ⇒ commit subject + body, run through
_strip_info_leak+_reflow_pr_bodyto remove cross-refs (SHAs, fix-PR links,(#NNNN)squash trailers,repo#Ncross-repo refs) and trim template noise. - Emit a Harbor task with the same shape as
pr_runtime:environment/Dockerfile,tests/test.sh,tests/verifier.py,tests/f2p.json,tests/p2p.json,solution/patch.diff, and the full[metadata.repo2env]block includingreward_calibration(f2p_count,p2p_count,source_files,loc_changed,difficulty).
Options (CommitRuntimeOptions)¶
limit: int = 50 # max candidate commits to walk
since: date | None = None
until: date | None = None
branch: str = "HEAD"
clone_depth: int = 200 # bump for monthly mining
# Metadata filters (cheap, applied before validation)
skip_merge_commits: bool = True
min_message_words: int = 5
max_source_files_per_commit: int = 10
exclude_authors: list[str] = [] # e.g. ["dependabot[bot]@users.noreply.github.com"]
require_new_test_funcs: bool = True
skip_ci_only: bool = True
# Validation
require_fail_to_pass: bool = True
min_fail_to_pass: int = 1
validation_timeout_sec: int = 600
Yield¶
Yield = emitted tasks ÷ commits walked. Expect ~10–35% — same F2P execution gate as pr_runtime, applied to raw commits, so it's noisier and a notch lower. The dominant factor is the repo's merge style: on repos that squash- or merge-commit their PRs, the source change and its test live in one commit and the per-commit F2P check works; on repos where the fix and its test land in separate commits, neither commit shows a fail→pass transition and yield drops toward 0 (use pr_runtime there instead).
| Knob | Default | Effect on yield |
|---|---|---|
skip_merge_commits | True | merge commits never carry a clean F2P; kept out |
require_new_test_funcs | True | ↑ drops commits that add no test function |
require_fail_to_pass | True | the gate — commits with no fail→pass test are dropped |
min_message_words | 5 | ↑ drops "wip"/"fmt"/"typo" commits |
max_source_files_per_commit | 10 | ↓ excludes sprawling refactors |
synthesize_with_llm | True | raises usable yield — rewrites thin/leaky commit messages into clean problem statements, so borderline commits become emittable instead of being dropped as instruction_too_thin |
Worked example: at ~20% yield, 100 tasks ≈ 500 commits walked. The reference …commit-runtime-test (100 envs) was built from a wide multi-language pool at clone_depth=200 per repo — budget ~15–25 CPU-clean, non-squash repos.
Known limitations (Arc 3 audit)¶
- Thin instructions when no issue link. ~30% of emitted tasks lacked a
Closes #N, so the instruction is sourced from the commit subject + body — which can be a one-liner.pr_runtimeis naturally better when there's an issue to fall back on. - Lower yield on PR-driven repos. Pallets/Django/Flask-style repos merge-commit their PRs;
commit_runtimecorrectly rejects all those merge commits and so yields ~0 there. Usepr_runtimefor those. - Go-subtest parser over-counts untracked failures. A handful of
stretchr/testifytasks land tracked-resolved but not command_resolved because the verifier marks Go subtests as "untracked failed" when the parent test exited 0. Affectscommand_resolved/eval_gradeonly; gold-patch reward is still 1.0. - No
commit_stream/ continuous variant.pr_streamwas removed in v0.8.3 as scope-creep; no plans to add a commit equivalent. If needed in future, it would be flags (since=auto+state-file=…) oncommit_runtime, not a separate pipeline.
See findings-commit_runtime for the full Arc 3 changelog + audit.