Code Generation · AI Agents

SWE-Explore vs DeNovoSWE vs Claw-SWE-Bench: Three Levers on Coding Scores

Three 2026 papers isolate what moves coding-agent scores: localization stalls at 0.15-0.20 line recall, training lifts a 30B model 8x, and the harness alone swings Pass@1 by 27.4 points.

SWE-Explore vs DeNovoSWE vs Claw-SWE-Bench: Three Levers on Coding Scores

The three levers

All three papers behind this comparison try to answer one question from a different angle: why do coding agents that look strong in demos still fail on real repositories, and what actually moves their scores? SWE-Explore attacks the perception stage: before an agent can patch anything, it has to locate the right code regions in a repo it has never seen, and the paper measures that step in isolation. DeNovoSWE attacks the training stage: if long-horizon repo work fails because models were never trained on it, then build the missing data and fine-tune. Claw-SWE-Bench attacks the harness stage: the scaffolding, adapters, and prompt plumbing around the model, which most evaluations treat as invisible.

None of these is a model. That is the point. Two of the three hold the model fixed and still move scores by amounts that would look like a new frontier model release if you only watched headline Pass@1 numbers.

Key numbers

DimensionSWE-ExploreDeNovoSWEClaw-SWE-Bench
What it isBenchmark for the repo-exploration stageDataset + fine-tuning recipe for whole-repo generationBenchmark for coding-agent harnesses
Scale848 issues, 203 repos, 10 languages; repos average 759 files and ~180K non-test lines4,818 auto-built instances, ~46x the 104-instance NL2Repo prior dataset350 issue-resolution instances, 8 languages, 43 repos; plus an 80-instance Lite subset
Headline numberAgentic explorers hit HitFile ~0.65 vs BM25 0.079, but line-level recall stalls at 0.15-0.20Qwen3-30B-A3B-Instruct rises from 0.058 to 0.472 pass rate on BeyondSWE-Doc2Repo after fine-tuningOpenClaw on the same GLM 5.1 backbone goes from 19.1% to 73.4% Pass@1 when the adapter changes
Second key numberContext Efficiency correlates with downstream repair success at Pearson r = +0.950The 35B-A3B variant goes 0.438 to 0.500, so most of the 8x lift is rescuing a weak baseHarness choice moves Pass@1 by 27.4 points under fixed models; model choice moves it 29.4 points
What it rules outClassical retrieval: BM25 sits at 0.079 HitFile and 0.021 line recallFine-tuning alone reaching frontier: GPT-5.4(CodeX) still scores 0.617 vs the fine-tuned 0.472The idea that adapters are plumbing: future-commit cleanup dropped Claude Opus 4.7 from 84.7% to 76.7%

Read the table for what it does not contain: none of these numbers comes from the same task. SWE-Explore scores exploration quality, DeNovoSWE scores whole-repo regeneration from documentation, and Claw-SWE-Bench scores issue resolution. A ranking across the columns would be meaningless. What the columns share is a methodological commitment: hold everything you can constant, change one stage, and measure the delta honestly, including the cases where the delta is embarrassing.

What each paper actually shows

SWE-Explore shows retrieval is solved and localization is not. On its 848-issue benchmark, modern agentic explorers reach HitFile around 0.65 while BM25 manages 0.079 and TF-IDF 0.140, so the “just embed the repo and retrieve” era is decisively over for this task. But the same agents recover only 0.15-0.20 of the relevant lines, and the paper’s strongest claim is that this gap, not file-finding, predicts repair failure: Context Efficiency correlates with downstream repair success at Pearson r = +0.950, with First Useful Hit at +0.928, HitFile at +0.925, and nDCG@500 at +0.921. Specialized structure-aware search already points at the ceiling: CoSIL’s iterative code-graph search reaches 0.788 line recall versus 0.15-0.19 for general agents, a 4x gap, and the paper’s own summary is that LLM choice “shifts the operating point, but not the bottleneck.”

DeNovoSWE shows the missing capability is trainable. Its premise is that most SWE training data teaches the wrong task: SWE-bench-style issues hand the model a working repository and ask for a patch, while DeNovoSWE asks for the whole repository, regenerated from capability documentation, graded by the project’s own executable tests. The dataset scales this to 4,818 instances, roughly 46 times the 104-instance NL2Repo predecessor, and fine-tuning Qwen3-30B-A3B-Instruct on it lifts the pass rate on BeyondSWE-Doc2Repo from 0.058 to 0.472, an 8x relative jump. The honest limits are stated in the same paper: the 35B-A3B variant starts at 0.438 and gains less, so much of the headline lift is rescuing a weak base; the metric is the mean fraction of unit tests passed, so 0.472 does not mean 47% of repos run end to end; and frontier proprietary models still lead, with GPT-5.4(CodeX) at 0.617, GLM-5 at 0.568, and DeepSeek-V4-Pro at 0.566.

Claw-SWE-Bench shows the harness is a first-class variable. Its 350-instance benchmark holds the model constant and varies the OpenClaw harness around it. The headline comparison is blunt: with a minimal direct-diff adapter, GLM 5.1 scores 19.1% Pass@1; with the full adapter, the same backbone scores 73.4%. Across sweeps, model choice moves Pass@1 by 29.4 points while harness choice moves it 27.4 points under fixed models, meaning the scaffolding is nearly as strong a signal as the model it wraps. The paper also quantifies evaluation hygiene: future-commit cleanup dropped Claude Opus 4.7 from 84.7% to 76.7%, and the Lite-80 subset costs about 22.9% of a full run while preserving mean Pass@1 within 0.4 points (0.639 on full-350 versus 0.643 on Lite-80).

How the three fit together

Read as a set, the papers imply a diagnostic order. When an agent fails on a repository task, the cheapest explanation to test first is exploration: SWE-Explore’s numbers say file-finding is basically solved, so if your agent cannot name the right files, it is behind a boundary that was crossed a while ago. The more common failure is knowing the files but not the lines, and that is a localization problem no amount of prompt polish reliably fixes. If localization looks healthy and the agent still cannot complete long-horizon work, DeNovoSWE’s result says the capability gap is trainable with verifiable whole-repo data, though its own ablations warn that the biggest gains accrue to weak bases. And if scores shift mysteriously when you swap models or refactor prompts, Claw-SWE-Bench’s 27.4-point harness delta says to freeze and measure the harness before crediting the model.

When to use which

  • Debugging why agents fail on your repos. Start with SWE-Explore’s metric family: HitFile, line-level recall, Context Efficiency, First Useful Hit. Its central finding is that these exploration metrics predict repair success with correlations above 0.92, so they are the cheapest leading indicator you can instrument.
  • Training or fine-tuning an agent for long-horizon work. DeNovoSWE is the template: generate verifiable tasks where the grading comes from the original repository’s own tests, keep partial-credit trajectories on hard instances, and expect the largest lift on weaker bases. Verify the gap to frontier models yourself; the paper’s own numbers say it remains.
  • Choosing or maintaining an agent harness. Claw-SWE-Bench’s numbers say adapter and harness design can be worth more than the model upgrade you were budgeting for. Treat harness changes as experiments with measurable Pass@1 deltas, use the Lite-80 subset for iteration, and clean future-commit leakage before publishing any number.
  • Publishing benchmark results. All three papers model the same discipline: report the metric’s exact definition, disclose cleanup procedures, and state what the number does not prove. Claw-SWE-Bench’s future-commit check and DeNovoSWE’s per-test partial credit are both reminders that headline scores are easy to inflate accidentally.

Limits and open questions

The obvious limit is that the three papers do not evaluate one agent on one task, so this page cannot tell you which system is “best.” Each answers a narrower question well: SWE-Explore tells you where localization fails, DeNovoSWE tells you whole-repo generation is trainable, and Claw-SWE-Bench tells you the harness matters as much as the model.

A deeper open question connects the first and third papers. SWE-Explore shows specialized code-graph search beats general agents by 4x on line recall, and Claw-SWE-Bench shows harness design moves scores by tens of points. Both suggest that agent architecture, not model scale, is the current bottleneck, but neither paper demonstrates the combined system: an explorer with CoSIL-style structural search wrapped in a well-engineered harness, trained on DeNovoSWE-style data. Whoever runs that experiment will produce the number to beat.

Cost is the third open question. DeNovoSWE reports no token, wall-clock, or dollar cost for its data-generation pipeline or long-horizon rollouts, and Claw-SWE-Bench reports that similarly-scoring systems can have very different API costs. For production decisions, the score-per-dollar column is still mostly empty across all three papers.

FAQ

Which of the three should I use to evaluate my coding agent?

Match the benchmark to the failure you are hunting. Use SWE-Explore when you suspect the agent is looking at the wrong code, Claw-SWE-Bench when you are comparing harnesses or adapters, and neither when the question is whether training helped, which is DeNovoSWE’s BeyondSWE-Doc2Repo setting. They measure different stages and are not interchangeable.

Why does SWE-Explore call retrieval “solved” when line recall is only 0.15-0.20?

Because the two metrics measure different things. Retrieval in the classical sense, ranking relevant files, is where BM25’s 0.079 HitFile shows the old approach is dead and agentic explorers reach ~0.65. Line-level recall is localization, recovering the exact relevant lines inside those files, and that is where general agents stall at 0.15-0.20. The paper’s argument is that the second gap, not the first, is what holds back repairs.

Can DeNovoSWE-trained models beat frontier proprietary models on whole-repo generation?

Not according to the paper’s own table. The fine-tuned Qwen3-30B-A3B reaches 0.472 on BeyondSWE-Doc2Repo while GPT-5.4(CodeX) reaches 0.617, GLM-5 0.568, and DeepSeek-V4-Pro 0.566. What the training does achieve is closing most of the gap at far smaller scale, beating the prior agent baseline Scale-SWE-Agent at 0.292.

How much of the Claw-SWE-Bench 19.1% to 73.4% jump is the model?

None of it. Both numbers come from the same GLM 5.1 backbone; only the OpenClaw adapter changes. Across the paper’s sweeps, harness choice moves Pass@1 by 27.4 points under fixed models and model choice by 29.4 points, so the harness is roughly model-sized as a source of variance.

What is the cheapest way to iterate on harness changes?

Claw-SWE-Bench’s Lite-80 subset is designed for exactly this: it costs about 22.9% of a full 350-instance run and preserved mean Pass@1 within about 0.4 points in calibration (0.639 on full-350 versus 0.643 on Lite-80). Use it for screening and regression checks, then confirm on the full benchmark before trusting a small delta.

Do better exploration scores guarantee better repairs?

Strongly correlated, not guaranteed. SWE-Explore reports Pearson r = +0.950 between Context Efficiency and downstream repair success, with three other metrics above 0.92. That makes exploration quality the best available leading indicator, but correlation is not a guarantee for any single task, and line-recall walls still cap the ceiling.