v0.9.3: FrontierSmith optimization synthesis

Turn any repository into verifiable RL environments.

Repo2RLEnv generates coding, terminal and reasoning tasks in the Harbor format. Each one has an instruction, a starting environment, a private verifier and a reference solution, ready to train and evaluate agents.

tasksmith

An agent that turns pull requests into environments.

Mining pipelines keep only the pull requests that pass their filters. Tasksmith adapts to each one instead: it reads the change, builds the repository, writes the task and its private verifier, and repairs its own work until the controls pass.

flagship · agenticExperimental

Tasksmith

Point an agent at a merged pull request and get a verified Harbor environment back. Tasksmith investigates the repository, builds its environment, designs the task, then reviews and repairs its own work until the controls pass.

  1. 01

    Investigate

    Reads the PR, its diff and the repository around it.

  2. 02

    Bootstrap

    Builds the repository’s environment on a remote worker.

  3. 03

    Design

    Writes the instruction and a private verifier.

  4. 04

    Construct

    Assembles the Harbor task; the merged code is the oracle.

  5. 05

    Review and repair

    Runs the controls and a blind solver, then fixes what fails.

Build an evaluation suite for your codebase

pipelines

Five kinds of tasks, from the material you already have.

All 23 pipelines
Repository repair

Fix real code in a real repository. The repository’s own tests decide the reward.

Implementation and reconstruction

Write or restore functionality, graded by hidden or differential tests.

Patch similarity

Reproduce a real change. The patch is scored against the merged one, no test suite needed.

Reasoning and optimization

Problems without a repository: exact answers, or open objectives with graded scores.

how it works

From source to a scored environment in three commands.

01

Generate

Point a pipeline at a repository, a merged PR or a task seed.

repo2rlenv generate --repo org/service --pipeline pr_runtime
02

Verify

Check every task statically, then prove it: the oracle scores 1, a no-op scores 0.

harbor run -p ./tasks -a oracle
03

Train and evaluate

Publish to the Hugging Face Hub and run any Harbor agent against it.

repo2rlenv push ./tasks org/my-envs

the output

A standard Harbor task, not a bespoke format.

Every pipeline emits the same directory. Harbor runs it in a sandbox with any of its agent harnesses (Claude Code, Codex, OpenHands and more), and the verifier writes the reward.

Read the task spec
org__service-412/
  • task.tomlmetadata, resources, lineage
  • instruction.mdwhat the agent sees
  • environment/Dockerfile for the starting state
  • tests/private verifier → reward
  • solution/reference solution (the oracle)

Generate your first environment in minutes.

Or start from ours: 21 published datasets on the Hugging Face Hub, each with its generation evidence and evaluation labels.