Papers
arxiv:2608.23564

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

Published on Aug 24
· Submitted by
Deyao Hong
on Aug 27
Authors:
,
,
,
,
,
,
,
,

Abstract

The study introduces a benchmark for evaluating autonomous software migration by coding agents, finding that current models rarely complete migrations correctly.

Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they autonomously perform such migrations? Existing benchmarks cannot answer this question because they evaluate only behavioural correctness, not whether the migration actually occurred. This leads an easy hack: agents copy the original implementation to make tests pass. We call this Blindness. To address this problem, we introduce SWE Refactor Bench, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt. A three-stage evaluation protocol measures both migration completeness and behavioural correctness. (1) Migration Audit verifies that the migration occurred. (2) Behavioural Tests measure correctness with a fixed test suite. (3) Agentic Verification uses 6 independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 of 520 runs (5.4%) pass all three stages, 13 of the 20 tasks receive no accepted solution, and the best model (claude-opus-5) scores 47.0/100. Migration completeness and behavioural correctness are distinct abilities: a few runs preserve behaviour by skipping the migration and are stopped at Migration Audit; most attempt it and break behaviour, and are stopped at Behavioural Tests. Agents cannot deliver a perfect migration: among the 340 runs that pass Migration Audit, 58% reach 99% of the fixed checks, yet only 26% reach 100%. Agent capability differs across migration categories: agents score 31.4 on build toolchain rewrites but only 5.6 on language rewrites. Together, these findings position SWE Refactor Bench as a rigorous testbed for developing coding agents for reliable whole-repository migrations.

Community

Paper author Paper submitter

The ultimate test for coding agents isn't local editing — it's whole-repo evolution, and right now, the survival rate is 5.4%.

Today we’re releasing SWE Refactor Bench, a benchmark for long-horizon, whole-repository software stack migration.

Coding agents are getting very good at fixing bugs.
But can they refactor an entire system, C → Rust, Maven → Gradle, POSIX → WebAssembly?

We built 20 real migrations across projects, including SQLite, zlib, libsodium, and GraphHopper.

520 runs. Only 28 survived all 3 stages. 13/20 tasks were solved by nobody.

System-scale migration is still wide open.

🏠 Homepage
https://lab.einsia.ai/swe-refactor-bench/

🏆 Leaderboard
https://lab.einsia.ai/swe-refactor-bench/leaderboard/

💻 GitHub
https://github.com/Einsia/SWE-Refactor-Bench

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.23564
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.23564 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2608.23564 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.23564 in a Space README.md to link it from this page.

Collections including this paper 1