Repo2RLEnv

CLI reference

Every repo2rlenv command, subcommand and flag with its default, read from the argparse definitions.

Edit on GitHub

repo2rlenv is one executable with thirteen commands. This page lists every flag and its default as defined in the v0.9.3 argparse code. Run repo2rlenv COMMAND --help (or repo2rlenv COMMAND SUBCOMMAND --help) for the same information in your terminal.

Global flags

repo2rlenv [--version] [-v] [--no-ui] COMMAND ...
FlagDefaultDescription
--versionn/aPrint repo2rlenv <version> and exit.
-v, --verboseoffDebug logging. On an error, also print the traceback to stderr.
--no-uioffDisable Rich live displays and print plain log lines.
-h, --helpn/aShow help for the command it follows.

Global flags go before the command: repo2rlenv --no-ui generate …. Placed after the command, they fail with unrecognized arguments.

Conventions

  • Credentials. At startup the CLI loads a .env file without overriding variables you already set. See .env files for where it looks.
  • Exit codes. 0 means success. 1 means the command finished but produced nothing usable (for example generate emitted no tasks or validate found a failing task). 2 means invalid arguments or a setup error.
  • Errors. Printed to stderr as ErrorType: message. Commands that accept --json print {"error": "ErrorType", "message": "…"} to stdout instead.
  • Live displays. They turn off automatically when stdout isn't a terminal, when NO_COLOR or CI is set, or when TERM=dumb.

generate

Run a pipeline against a source and write Harbor tasks to a local directory.

repo2rlenv generate [--config FILE] [--repo REPO] [--pipeline NAME] [--recipe NAME]
                    [--pipeline-opt KEY=VALUE ...] [--llm PROVIDER/MODEL] [--out DIR] ...

generate routes on --recipe:

  • Native (no --recipe, or --recipe native) runs on your machine. pr_diff needs neither Docker nor an LLM. Every other native pipeline needs --llm and Docker, because it first bootstraps a container image for the repository.
  • Research recipes (--recipe swe_smith, …) and non-repository sources need a --config file with an execution: block and a running remote worker. There is no local fallback. See Remote execution.
FlagDefaultDescription
--config FILEnoneGeneration input as YAML (.yaml, .yml), TOML or JSON. Flags on the command line override its values.
--repo REPOnoneGitHub owner/name, a GitHub or GitLab URL, or a local path.
--ref REFHEADBranch, tag or commit. Applied together with --repo.
--accessautopublic, private or auto. Applied together with --repo.
--pipeline NAMEnonePipeline name, such as pr_diff. repo2rlenv pipelines list shows them all.
--recipe NAMEnativeResearch recipe to run for the pipeline, such as swe_smith.
--pipeline-opt KEY=VALUEnonePipeline option. Repeatable. true/false become booleans, JSON values (numbers, lists) are parsed, and anything else stays a string. Overrides options from --config.
--llm PROVIDER/MODELnoneModel for synthesis and bootstrap, such as anthropic/claude-sonnet-4-6.
--llm-fallback PROVIDER/MODELnoneModel to switch to when the primary returns 5xx, rate-limit or network errors.
--llm-endpoint URLnoneBase URL of a self-hosted OpenAI-compatible server for the primary model, such as http://localhost:8000/v1. The provider's default key is never sent there.
--llm-key-env VARprovider defaultEnvironment variable that holds the primary model's API key.
--out DIRnoneOutput directory.
--org ORGdefaultTask namespace: task names become ORG/<slug>. Applied together with --out.
--dataset-name NAMEname of --outDataset name recorded in the output spec. Applied together with --out.
--visibilitypublicpublic or private. Applied together with --out.
--max-spend-usd N5.0Native only. Caps the bootstrap agent's LLM spend; 0 removes the cap. Research recipes reject it, because their spend comes from a campaign budget.
--language LANGauto-detectBootstrap language: python, node, go, rust, java or c_cpp.
--base-image IMAGEper languageBootstrap base image, such as ubuntu:24.04 or python:3.11-slim.
--force-bootstrapoffIgnore the bootstrap cache and rebuild the image.
--bootstrap-opt KEY=VALUEnoneSet any bootstrap field, such as cache_dir, max_iterations, max_seconds or user_dockerfile. Repeatable.
--force-languageoffSkip the check that the repository's primary language is one the pipeline supports.
--resumeoffRecipes only. Continue a run without redispatching completed work.
--jsonoffRecipes only. Stream progress as JSON Lines. Native generation rejects it.

--llm-fallback, --llm-endpoint and --llm-key-env need --llm or an llm: block in --config. The bootstrap flags (--language through --force-language) only affect native pipelines that bootstrap.

The pipeline option limit counts candidates listed (for PR pipelines, merged PRs), not tasks emitted. Filters usually drop some, so generate can finish with fewer tasks than the limit. It exits 1 when it emits none.

repo2rlenv generate \
  --repo pallets/click \
  --pipeline pr_diff \
  --pipeline-opt limit=10 \
  --out ./tasks

A research recipe takes its source, options and execution: block from the config file:

repo2rlenv --no-ui generate --config examples/owned-swe-smith.yaml

validate

Statically check a dataset or task directory. Nothing is built or run.

repo2rlenv validate [--deep] [--oracle] PATH
FlagDefaultDescription
PATHrequiredDataset or single task directory. Every task.toml below it is checked.
--deepoffAlso check each task's assets and metadata: instruction, environment definition, tests/test.sh, graded verifier files and the reproducibility table.
--oracleoffImplies --deep. Also require solution/solve.sh and, for native pipelines, a non-empty solution/patch.diff.

Without flags, validate confirms each task.toml parses and has a [task] name. It exits 1 if any task fails or if no task.toml is found. Deep validation lists every check.

validate currently rejects research-recipe and Tasksmith output (schema 1.3 bundles) with missing top-level ['version'], because it still requires the Harbor 1.0 version key. A fix is in progress. Until then, use it on native pipeline output.

repo2rlenv validate ./tasks --oracle

push

Upload a local dataset directory to a Hugging Face dataset repository.

repo2rlenv push [--private] [-m MESSAGE] [--image-registry PREFIX] [--inline-dockerfile]
                [--require-registry] [--skip-image-push] [--image-visibility VIS]
                [LOCAL_DIR] [DATASET]
repo2rlenv push --check-auth [--fast] [--json]

push stages the tasks under tasks/, writes a dataset card and manifest.json, uploads them, then uploads a registry.json pinned to that commit. For tasks whose Dockerfile starts from a locally built bootstrap image, it first pushes the image to a container registry, or inlines the build recipe. It rewrites environment/Dockerfile and task.toml in place to record the choice. See Publishing.

FlagDefaultDescription
LOCAL_DIR.Dataset directory, as written by generate.
DATASETnoneTarget as owner/name or hf://owner/name. Required unless --check-auth. harbor:// and GitHub targets print the equivalent harbor publish or git push command instead.
--privateoffCreate the dataset repository as private.
-m, --message TEXTRepo2RLEnv: add N tasksCommit message.
--image-registry PREFIXauto-detectForce a registry prefix, such as ghcr.io/my-org. It's probed for write access first.
--inline-dockerfileoffDon't push images; write the bootstrap recipe into each environment/Dockerfile. Rebuilds from the recipe, not bit-exact.
--require-registryoffFail if no verified registry is available instead of falling back to inline mode.
--skip-image-pushoffRewrite tasks against an image that already exists at the registry; don't run docker push.
--image-visibilityinheritpublic, private or inherit (match the dataset's visibility).
--check-authoffProbe every registry you're logged in to, report, and exit. Uploads nothing.
--fastoffWith --check-auth: stop after the reachability and authentication levels.
--jsonoffWith --check-auth: print the probe results as JSON.
repo2rlenv push ./tasks my-org/click-pr-diff

pull

Download a dataset from the Hugging Face Hub, a Harbor registry or GitHub.

repo2rlenv pull [--task NAME] [--registry-url URL] [--force] DATASET [LOCAL_DIR]
FlagDefaultDescription
DATASETrequiredowner/name, owner/name@rev, hf://owner/name, or a bare name (owner resolved from your Hub login) for the Hub. harbor://name[@tag] for a Harbor registry. gh://owner/repo[@ref] or a https://github.com/owner/repo URL for GitHub.
LOCAL_DIR./datasets/<owner>__<name>Destination directory.
--task NAMEnoneHub only: download one task, such as encode__httpx-3367.
--registry-url URLHarbor's public registryHarbor only: a custom registry. Requires the harbor CLI on your PATH.
--forceoffDelete the local copy and download again.

A Hub pull flattens the published tasks/<id>/ layout back to LOCAL_DIR/<id>/, ready for harbor run -p. A GitHub pull is a shallow git clone.

repo2rlenv pull FineEnvs/repo2rlenv-pr-runtime ./pr-runtime --task encode__httpx-3367

bootstrap

Build and cache a working Docker image for a repository with an LLM agent, without generating tasks. A later generate against the same repository and commit reuses the cached image.

repo2rlenv bootstrap --repo REPO --llm PROVIDER/MODEL [--ref REF] [--max-spend-usd N] ...
FlagDefaultDescription
--repo REPOrequiredGitHub owner/name or URL.
--ref REFHEADBranch, tag or commit.
--accessautopublic, private or auto.
--llm PROVIDER/MODELrequiredModel that drives the build agent.
--llm-endpoint URLnoneSelf-hosted OpenAI-compatible server for --llm.
--llm-key-env VARprovider defaultEnvironment variable that holds the API key.
--max-iterations N25Maximum agent turns.
--max-seconds N1800Wall-clock limit.
--cache-dir DIR$R2E_CACHE_DIR, else ./workspace/bootstrapImage cache root.
--image-registry PREFIXnonePush the built image here, such as ghcr.io/my-org/r2e.
--platformlinux/amd64linux/amd64 or linux/arm64.
--language LANGauto-detectpython, node, go, rust, java or c_cpp.
--base-image IMAGEper languageBase image, such as ubuntu:24.04.
--max-spend-usd N5.0Abort when cumulative LLM cost passes this; 0 removes the cap.
--forceoffIgnore the cache and rebuild.

The agent's turn limit comes from --max-iterations here. Inside generate, it comes from the bootstrap spec (default 20; set it with --bootstrap-opt max_iterations=N). Bootstrap explains the cache and the agent loop.

repo2rlenv bootstrap --repo pallets/click --llm anthropic/claude-sonnet-4-6

pipelines

Discover native pipelines and research recipes. Makes no network calls and needs no credentials.

pipelines list

List every native pipeline and research recipe with its status (stable, experimental or planned) and source kind.

repo2rlenv pipelines list [--json]
FlagDefaultDescription
--jsonoffPrint {"native": […], "recipes": […]}.
repo2rlenv pipelines list --json

pipelines describe

Show one recipe's input, scope, upstream source pin and RFC.

repo2rlenv pipelines describe --recipe RECIPE [--json] PIPELINE
FlagDefaultDescription
PIPELINErequiredPipeline the recipe belongs to, such as repo_mutate.
--recipe RECIPErequiredRecipe id, such as swe_smith.
--jsonoffPrint the description as JSON.
repo2rlenv pipelines describe repo_mutate --recipe swe_smith

tasksmith

Build Harbor tasks from merged PRs with a remote coding agent (Pi or OpenCode). Needs the tasksmith extra plus modal or daytona, a campaign and a runtime wheel. See Tasksmith.

tasksmith run

Run or resume a fixed PR panel.

repo2rlenv tasksmith run PANEL --campaign DIR --output DIR --runtime-wheel WHEEL [options]
FlagDefaultDescription
PANELrequiredJSON panel of PRs, such as examples/tasksmith-pr.json.
--campaign DIRrequiredExisting campaign directory.
--output DIRrequiredRun directory. Rerunning the same command reuses matching completed work.
--runtime-wheel WHEELrequiredWheel built from the installed version (uv build --wheel).
--options FILEnoneOptions JSON, including quality models and spending limits, such as examples/tasksmith-options.json.
--providermodalmodal or daytona. Used when --options is absent; must agree with it otherwise.
--authorpipi or opencode. Used when --options is absent; must agree with it otherwise.
--stop-after NnoneProcess only the first N inputs. The report's denominator stays the full panel.
--sources FILEnoneFrozen PR source records, in panel order.
--prepared-task DIRnoneRecover one task matching --sources with fresh quality checks; the input is preserved.
--generation-run DIRnoneReuse verified generation artifacts from an identical panel and rerun only the current quality policy.
--reuse-evidenceoffWith --generation-run: also reuse matching baseline, reference and solver evidence. Semantic probes still rerun.
--bootstrap-report FILEnoneReuse the snapshot and dependency hints from a completed CPU bootstrap report.
--env-file FILEnoneLoad credentials from this file, without overriding variables already set.
--jsonoffPrint the report as JSON.

Exits 0 only when every input produced a usable task.

repo2rlenv tasksmith run examples/tasksmith-pr.json \
  --options examples/tasksmith-options.json \
  --campaign workspace/tasksmith \
  --output workspace/tasksmith/run \
  --runtime-wheel dist/repo2rlenv-0.9.3-py3-none-any.whl

tasksmith batch

Run a bounded parallel campaign over many PRs, retaining every generated task. The plan format is in Run Tasksmith on many PRs.

repo2rlenv tasksmith batch PLAN --campaign DIR --output DIR --runtime-wheel WHEEL [--env-file FILE] [--json]
FlagDefaultDescription
PLANrequiredBatch plan JSON (target, parallelism, spending cap, candidates).
--campaign DIRrequiredExisting campaign directory.
--output DIRrequiredBatch directory.
--runtime-wheel WHEELrequiredRuntime wheel.
--env-file FILEnoneCredentials file.
--jsonoffPrint the batch report as JSON.

Exits 0 when the plan's verified target was reached.

repo2rlenv tasksmith batch plan.json --campaign workspace/campaign \
  --output workspace/tasksmith-batch --runtime-wheel dist/repo2rlenv-0.9.3-py3-none-any.whl

tasksmith bootstrap

Build and smoke-test pinned repositories on remote CPU and GPU workers before a campaign.

repo2rlenv tasksmith bootstrap MATRIX --campaign DIR --output DIR --runtime-wheel WHEEL
                               [--resource cpu|gpu|both] [--env-file FILE] [--json]
FlagDefaultDescription
MATRIXrequiredJSON list of distinct repository bootstrap specs.
--resourcebothcpu, gpu or both. The GPU pass stops at the first repository that isn't ready.
--campaign DIRrequiredExisting campaign directory.
--output DIRrequiredReceipts, logs and compute estimates.
--runtime-wheel WHEELrequiredRuntime wheel.
--env-file FILEnoneCredentials file.
--jsonoffPrint the reports as JSON.
repo2rlenv tasksmith bootstrap matrix.json --resource cpu --campaign workspace/campaign \
  --output workspace/bootstrap-matrix --runtime-wheel dist/repo2rlenv-0.9.3-py3-none-any.whl

tasksmith install-runtime

Install the pinned Pi and OpenCode libraries into your user cache. Takes no flags.

repo2rlenv tasksmith install-runtime

tasksmith show

Read a saved run or batch report. Makes no paid calls.

repo2rlenv tasksmith show [--json] OUTPUT
FlagDefaultDescription
OUTPUTrequiredRun or batch directory containing report.json.
--jsonoffPrint the report as JSON.
repo2rlenv tasksmith show workspace/tasksmith/run --json

campaign

Create and inspect the budget ledger (budget.sqlite3) that research recipes, Tasksmith, workers and the quality loop reserve spend against. See Remote execution.

campaign init

Create a campaign directory with a fixed spending limit. Running it again with the same amount is harmless; a different amount is refused.

repo2rlenv campaign init --budget-usd N [--json] PATH
FlagDefaultDescription
PATHrequiredCampaign directory.
--budget-usd NrequiredSpending limit in USD.
--jsonoffPrint the budget as JSON.
repo2rlenv campaign init workspace/my-campaign --budget-usd 25

campaign status

Show the limit, accounted spend, open reservations and remaining allowance. Warns when operations are waiting for reconciliation.

repo2rlenv campaign status [--json] PATH
FlagDefaultDescription
PATHrequiredCampaign directory created with campaign init.
--jsonoffAlso list every operation with its status (reserved, uncertain or settled).
repo2rlenv campaign status workspace/my-campaign --json

campaign settle

Close a reservation with its actual cost and an evidence file. The ledger records the file's path and SHA-256.

repo2rlenv campaign settle --operation ID --cost-usd N --evidence FILE [--json] PATH
FlagDefaultDescription
PATHrequiredCampaign directory.
--operation IDrequiredOperation id from campaign status --json. Workers use worker:<provider>:<name>.
--cost-usd NrequiredActual cost in USD.
--evidence FILErequiredUsage or billing receipt, or a clearly labeled conservative estimate.
--jsonoffPrint the budget as JSON.
repo2rlenv campaign settle workspace/my-campaign --operation worker:modal:my-worker \
  --cost-usd 0.30 --evidence workspace/my-campaign/worker-usage.json

workers

Create, check and terminate remote Modal or Daytona workers. Each worker has a JSON receipt at CAMPAIGN/workers/NAME.json that later commands take as input. Needs the modal or daytona extra.

workers start

Reserve spend in the campaign, create the worker and write its receipt.

repo2rlenv workers start --campaign DIR --provider modal|daytona --name NAME --reserve-usd N
                         [--cpus N] [--memory-mb N] [--timeout-sec N] [--json]
FlagDefaultDescription
--campaign DIRrequiredExisting campaign directory.
--providerrequiredmodal or daytona.
--name NAMErequiredLowercase letters, digits and hyphens, starting with a letter. Also the receipt's file name, so it can't be reused in the same campaign.
--reserve-usd NrequiredAmount to reserve for the worker's lifetime.
--cpus N2CPUs (1–16).
--memory-mb N4096Memory in MB (1024–65536).
--timeout-sec N3600Lifetime in seconds (60–14400). On Daytona this becomes an idle auto-stop.
--jsonoffPrint the receipt as JSON.
repo2rlenv workers start --campaign workspace/my-campaign --provider modal \
  --name my-worker --reserve-usd 3 --timeout-sec 3600

workers probe

Check a running worker: Docker daemon readiness, byte-exact file transfer, an image build and an isolated container run with no network access.

repo2rlenv workers probe --out DIR [--json] RECEIPT
FlagDefaultDescription
RECEIPTrequiredReceipt of a running worker.
--out DIRrequiredNew directory for probe.json and command logs. It must not exist.
--jsonoffPrint the result as JSON.
repo2rlenv workers probe workspace/my-campaign/workers/my-worker.json \
  --out workspace/my-campaign/probes/first

workers stop

Terminate the worker and mark its reservation for reconciliation. Stopping confirms cleanup; it doesn't record a cost. Settle it with campaign settle.

repo2rlenv workers stop [--json] RECEIPT
FlagDefaultDescription
RECEIPTrequiredWorker receipt. Stopping an already terminated worker is a no-op.
--jsonoffPrint the receipt as JSON.
repo2rlenv workers stop workspace/my-campaign/workers/my-worker.json

release

Publish an explicit, immutable selection of tasks as a Hub dataset, checked against each task's bundle hash. Only research-recipe and Tasksmith bundles record a bundle_hash (and the recipe the plan names); native output doesn't, so publish it with push. See Releasing Harbor task collections.

release stage

Check each selected task against its expected bundle hash, then build the release directory: tasks, tasks.tar.gz, manifests, index and dataset card. Needs the harbor extra, which parses every task.

repo2rlenv release stage --out DIR [--json] PLAN
FlagDefaultDescription
PLANrequiredRelease plan JSON.
--out DIRrequiredNew staging directory. It must not exist or sit inside a selected task.
--jsonoffPrint the result as JSON.
repo2rlenv release stage release-plan.json --out workspace/releases/swe-smith

release verify

Check that every staged file still matches its recorded hash and mode. Runs nothing.

repo2rlenv release verify [--json] DIRECTORY
FlagDefaultDescription
DIRECTORYrequiredStaging directory.
--jsonoffPrint the result as JSON.
repo2rlenv release verify workspace/releases/swe-smith

release publish

Upload a verified release, pin registry.json to the upload commit and optionally add the dataset to a collection. Needs HF_TOKEN or a Hub login.

repo2rlenv release publish --receipt FILE [--collection SLUG] [--batch-size N | --recover-empty] [--json] DIRECTORY
FlagDefaultDescription
DIRECTORYrequiredVerified staging directory.
--receipt FILErequiredPublication receipt. A completed receipt returns without uploading again; an incomplete one stops for reconciliation.
--collection SLUGnoneCollection to add the dataset to.
--batch-size NnoneUpload to a new, empty repository in commits of N files (1–500) instead of one large commit.
--recover-emptyoffRecover an upload that timed out before creating a commit. Confirms the repository is still empty, then uploads in bounded commits.
--jsonoffPrint the receipt as JSON.
repo2rlenv release publish workspace/releases/swe-smith \
  --receipt workspace/releases/swe-smith-publication.json

quality

Review a Harbor task against evidence, and optionally run controls, probes and a blind solver remotely and repair the task. See Review and repair.

quality run

repo2rlenv quality run TASK --campaign DIR --out DIR [--repair] [--run-rollout] [options]

A run without --repair or --run-rollout makes model calls but creates no worker. Those two flags need a remote worker: --runtime-wheel for CPU tasks, or native Modal for tasks that request GPUs. They also accept only single-container tasks with network_mode = "no-network", such as research-recipe and Tasksmith bundles. Native pipeline output stops with Remote validation currently requires a single-container no-network Dockerfile task.

FlagDefaultDescription
TASKrequiredHarbor task directory.
--campaign DIRrequiredExisting campaign directory.
--out DIRrequiredRun directory for result.json, revisions, trials and receipts.
--env-file FILEnoneLoad credentials from this file. They're never copied into task artifacts.
--baseline PATHnoneExisting trial evidence: an owned trial receipt, a Harbor trial result.json, or a single-trial directory.
--oracle PATHnoneSame, for the reference solution.
--rollout PATHnoneSame, for a solver rollout.
--probes FILEnoneTask-bound JSON manifest of known counterexamples and valid alternatives.
--review-model PROVIDER/MODELanthropic/claude-sonnet-4-6Reviewer. Only openai/ and anthropic/ routes are accepted, for this and the other model flags.
--repair-model PROVIDER/MODELthe reviewerRepair author.
--escalation-model PROVIDER/MODELnoneOne stronger attempt when a review stays unresolved.
--solver-model PROVIDER/MODELanthropic/claude-sonnet-4-6Blind solver.
--repairoffAuthor repairs and validate each new revision remotely.
--run-rolloutoffRun missing controls, probes and a blind solver remotely, without editing.
--max-repairs N3Repair rounds (0–5).
--max-read-rounds N2Extra file-reading rounds (0–4).
--max-probes N2Semantic probes (0–4).
--context-chars N100000Document context per call (16000–250000).
--model-tokens N6000Output tokens per review or repair call (1024–16000).
--max-turns N24Solver turns (1–100).
--solver-tokens N4096Solver output tokens per call (256–8192).
--trial-timeout-sec N900Timeout per trial (30–3600).
--model-reservation-usd N1.00Reservation per model call.
--solver-reservation-usd N4.00Reservation per solver rollout.
--max-spend-usd N15.00Spending limit for this run, inside the campaign limit.
--success-reward N1.0Reward that counts as a solver success.
--providermodalmodal or daytona for CPU workers. GPU tasks always use native Modal.
--runtime-wheel WHEELnoneRequired for CPU --repair and --run-rollout.
--worker-receipt RECEIPTnoneReuse a running worker from the same campaign. You stay responsible for stopping it.
--worker-reservation-usd N3.00Reservation for a worker the run creates.
--resumeoffReuse completed, identical model requests and trial evidence.
--jsonoffPrint the result as JSON.

Exits 0 for usable and reviewed, 1 for other dispositions. Automation that needs validated tasks should check status == "usable", not the exit code.

repo2rlenv quality run ./tasks/example \
  --campaign workspace/my-campaign \
  --out workspace/reviews/example \
  --run-rollout --provider modal \
  --runtime-wheel dist/repo2rlenv-0.9.3-py3-none-any.whl

quality show

Read a completed quality result. Makes no paid calls.

repo2rlenv quality show [--json] PATH
FlagDefaultDescription
PATHrequiredresult.json, or a run directory containing it.
--jsonoffPrint the result as JSON.
repo2rlenv quality show workspace/reviews/example

tasks

Inspect and set the advisory evaluation label stored in each task's [metadata.repo2env.evaluation]. See Evaluation labels.

tasks show

Show one task's label.

repo2rlenv tasks show [--json] PATH
FlagDefaultDescription
PATHrequiredTask directory.
--jsonoffPrint the label as JSON.
repo2rlenv tasks show ./tasks/example

tasks list

List the labels of every task under a directory.

repo2rlenv tasks list [--status STATUS] [--json] PATH
FlagDefaultDescription
PATHrequiredDataset directory.
--statusallKeep only unverified, verified, needs_repair or blocked.
--jsonoffPrint the labels as JSON.
repo2rlenv tasks list ./tasks --status needs_repair

tasks label

Write a labeled copy of a task to a new directory. The original is left untouched.

repo2rlenv tasks label --out DIR [--quality-result FILE | --status STATUS] [--stage STAGE]
                       [--reason-code CODE ...] [--detail TEXT] [--provenance P] [--json] TASK
FlagDefaultDescription
TASKrequiredTask directory.
--out DIRrequiredNew task directory.
--quality-result FILEnoneDerive the label from a saved quality result.json. This is the only way to set verified.
--statusunverifiedunverified, needs_repair or blocked. Can't be combined with --quality-result.
--stagegenerationgeneration, bootstrap, construction, review, controls, probes, rollout, repair, complete or unknown.
--reason-code CODEvalidation_not_runRepeatable snake_case diagnosis. needs_repair and blocked require at least one, plus --detail.
--detail TEXTnoneHuman-readable reason and next diagnostic step.
--provenanceunknownunknown, assisted or unattended.
--jsonoffPrint the new label as JSON.
repo2rlenv tasks label ./tasks/example --out ./labeled/example \
  --quality-result workspace/reviews/example/result.json

codemidas

Commands specific to the CodeMidas recipe. Both print JSON.

codemidas source

Validate a pinned Stack v3 repository manifest and print its provenance. Executes nothing.

repo2rlenv codemidas source [--materialization inline|hydrated] MANIFEST
FlagDefaultDescription
MANIFESTrequiredStack v3 row manifest JSON.
--materializationinlineinline uses the row's files; hydrated restores the GitHub checkout at the same commit.
repo2rlenv codemidas source workspace/stack-row.json --materialization inline

codemidas audit

Run adversarial and solver audits on a generated CodeMidas task, plus a separate curriculum screen.

repo2rlenv codemidas audit TASK --controls DIR --out DIR --campaign DIR
                           --worker-receipt RECEIPT --runtime-wheel WHEEL [options]
FlagDefaultDescription
TASKrequiredGenerated task directory.
--controls DIRrequiredThe candidate's generation directory, whose control receipts are reused.
--out DIRrequiredAudit directory.
--campaign DIRrequiredCampaign directory.
--worker-receipt RECEIPTrequiredRunning worker receipt.
--runtime-wheel WHEELrequiredRuntime wheel.
--screen-attempts N4Curriculum screening attempts.
--attempt-concurrency N2Independent solver attempts in parallel (1–4).
--max-cost N6Trial allowance in USD. Independent review has its own maximum of $1.25.
--resumeoffReuse unchanged attempts.
repo2rlenv codemidas audit workspace/codemidas/tasks/TASK \
  --controls workspace/codemidas/runs/RUN/tasks/CANDIDATE \
  --campaign workspace/codemidas \
  --worker-receipt workspace/codemidas/workers/worker.json \
  --runtime-wheel dist/repo2rlenv-0.9.3-py3-none-any.whl \
  --out workspace/codemidas/audits/TASK

On this page