The benchmarks evaluate LM agent on SWE/Computer-use tasks across different operating systems.
SWE-bench-Live
community
AI & ML interests
None defined yet.
Recent Activity
Organization Card
SWE-bench-Live
Here we host SWE-bench-Live dataset, with continuous monthly updates!
models 0
None public yet
datasets 5
SWE-bench-Live/Windows
Viewer • Updated • 66 • 1.55k
SWE-bench-Live/MultiLang
Viewer • Updated • 1.08k • 21k
SWE-bench-Live/SWE-bench-Live
Viewer • Updated • 3.69k • 72.1k • 8
SWE-bench-Live/repo-build-test-benchmark
Viewer • Updated • 119 • 777
SWE-bench-Live/OS-bench
Viewer • Updated • 140 • 1.13k