Finding bugs with 0% of our model tokens
Repro spends none of its model tokens finding bugs and an estimated ~68% fewer tokens fixing them. Here is where those two numbers come from.
This post is a draft for the team to review. It isn't final yet.
Two numbers come up whenever we talk about Repro: 0% of model tokens spent finding bugs, and ~68% fewer tokens to fix them. They come from two design decisions, and one of them is an estimate, so it’s worth being precise about both.
0%: detection is deterministic
In Repro, finding bugs is not a model’s job. Deterministic scanners produce the findings, and no model decides what counts as a finding. Each finding’s reproduction command is then re-run in a sandbox, and only findings that show the problem are kept.
Because that work is done by scanners and reproduction commands, finding bugs uses none of our model tokens.
~68%: fix the cause, not every symptom
After reproduction, Repro groups the remaining findings by root cause and proposes a fix per group. Several findings that come from the same underlying problem get one fix instead of one each.
The ~68% figure compares that approach with repairing every finding separately. It’s an estimate, not a measured benchmark, and we label it that way wherever it appears.
What stays the same
Fewer tokens doesn’t mean fewer checks. Every fix still goes through Verify: the tests run, the reproduction runs again, an adversarial check tries to break the fix, and a person decides whether to merge.
The full pipeline is in How Repro proves a bug before it fixes it, and the project is on GitHub and Devpost.