Rewind-Bench
It passed. It was not fixed.
7% of the best agent's accepted fixes were fake. A hidden seed caught them.
Automation & AI agent tooling engineer — builds execution infrastructure for AI agents: notebook-first workflow orchestration, in-page browser automation, and benchmarks that check whether AI code fixes actually hold up.
Available full-time — freelance and contract-to-hire work, remote, 40 hours a week. What that covers.
It passed. It was not fixed.
7% of the best agent's accepted fixes were fake. A hidden seed caught them.
Stop drawing boxes.
One notebook becomes a UI page, an HTTP endpoint and an agent tool, inside a React and Express app.
Browser automation written one step at a time, in a side panel, running on pages whose Content-Security-Policy blocks eval
Build notes: a desktop app that runs repeat work in the owner's own Chrome.
The Devpost write-up of Rewind-Bench, catching AI code repairs that only look correct, submitted to Impact Forge Summer 2026.
The Rewind-Bench results report: injected bugs, three agent arms, and the accepted fixes that failed once re-run on a hidden seed.