Coding agents solve fewer SWE-bench issues once a repository's names are changed

21 hours ago

A team from Shanghai Jiao Tong University renamed the code inside SWE-bench Verified's repositories, kept its behaviour identical, and reran four models. Every model solved fewer issues, by 6 to 14 percentage points, and read 2.4 to 3.6 times as many input tokens finding its way around. On newer issues the models could not have seen, the same renaming left the score unchanged. Paper: https://arxiv.org/abs/2609.27891 Code and data: https://github.com/cslsolow/Schrodinger-Repo

Ask

Ask about this presentation

Answers are generated from this presentation.

Chapters

  1. 0:00Coding agents and renamed repositories
  2. 0:23What SWE-bench Verified is
  3. 0:45Models recalled the fixes
  4. 1:06The test
  5. 1:09Same bug, same tests, different names
  6. 1:32QuerySet becomes LedgerSuite
  7. 1:51What happened
  8. 1:54Every model solved fewer issues
  9. 2:18The renaming did most of the damage
  10. 2:37Agents read far more on unfamiliar code
  11. 2:56One Django bug: 37 steps, then 217
  12. 3:19On newer issues, the score held
  13. 3:43What it means
  14. 3:45What the paper doesn't test
  15. 4:10What it means for picking a coding agent
  16. 4:36The paper and its code
Show transcript

Coding agents and renamed repositories

AI Papers

Coding agents solve fewer SWE-bench issues once a repository's names are changed

Silin Chen, Yufei Yang, Xiaodong Gu and colleagues · Shanghai Jiao Tong University, Xi'an Jiaotong University, East China Normal University

Chen et al. · Schrödinger's Code Repository · arXiv 2609.27891 · 21 Aug 2026

SWE-bench Verified is the score AI coding tools quote to show they can fix real bugs. A team from Shanghai Jiao Tong University and two other universities changed how the code inside it looks, kept its behaviour identical, and ran four models again. Every model solved fewer issues. Their paper, Schrödinger's Code Repository, went up on arXiv on August 21st.

What SWE-bench Verified is

010203The test

500 real bugs from popular Python projects

An agent gets the issue and the codeit edits the repository, and the project's own tests decide
The projects are public and famousDjango is one; their code, issues and fixes can end up in training data
Chen et al. · Schrödinger's Code Repository · arXiv 2609.27891

The benchmark is 500 real issues from popular open-source Python projects. An agent gets the issue and the code, makes a fix, and the project's own tests decide whether it worked. The catch is that these projects are public and famous. Django is one of them, and their code, issues and fixes can end up in a model's training data.

Models recalled the fixes

010203The test

Shown only part of the issue, models recalled the fix

65%+of issues: the model named task details it hadn't been shown
18%+of issues: it recalled the patch or the test itself
Chen et al. · Schrödinger's Code Repository · arXiv 2609.27891 · Figure 1

Before changing anything, the authors checked what the models already remembered. Reviewers gave each model the issue's ID and a few details from it, with no code, and compared its answers to the real fix. For every model, about two thirds of the issues showed it knew details it hadn't been shown. For about one in five, it recalled the patch or the test itself.

The test

01
The test

So the authors built a way to take that familiarity away.

Same bug, same tests, different names

010203The test
The issue is reworded
The project's own names are remappedclasses, functions, files and folders
Functions inside a file are reordered
Code near the bug is rewrittento do exactly the same thing

Each version passes and fails the same tests as the original.

Chen et al. · Schrödinger's Code Repository · arXiv 2609.27891 · Section III

Their tool, SchrodingerRepo, builds a fresh copy of each repository whenever an agent starts. It rewords the issue, renames the project's own classes, functions and files, reorders the functions inside a file, and rewrites the code around the bug. Each version is checked against the project's tests, so the bug and the correct fix stay exactly the same.

QuerySet becomes LedgerSuite

010203The test

Django's QuerySet becomes LedgerSuite

QuerySet
LedgerSuite

Query becomes Ledger everywhere. Python and outside libraries keep their real names.

Chen et al. · Schrödinger's Code Repository · arXiv 2609.27891 · Section III-C

Here's their own example. Django's QuerySet class becomes LedgerSuite. The swap is consistent, so query turns into ledger everywhere it appears, while Python itself and outside libraries keep their real names. Every command the agent runs is translated back, so it still works on the real project underneath.

What happened

02
What happened

Then they ran four models on both versions.

Every model solved fewer issues

010203What happened
original coderenamed code
GPT-5.1
44.6%
36.2%
GPT-5.4 mini
46.8%
35.6%
DeepSeek-V4-Flash
72.8%
66.8%
Gemini 3.1 Flash-Lite*
56.7%
42.3%
Chen et al. · Schrödinger's Code Repository · arXiv 2609.27891 · Table I · *Gemini: 300 highest-recall issues

Here's the share of issues each model solved, on the original code and on the renamed code. GPT-5.4 mini fell from about 47 percent to about 36. DeepSeek, the strongest here, fell from about 73 to 67. Across the four models the drop ran from 6 to 14 percentage points, and the authors report it as statistically significant.

The renaming did most of the damage

010203What happened

Renaming alone did most of the damage

Remapped names6 to 7.4 points lower
Reworded issue0 to 2 points lower
Reordered or rewritten code0.8 to 3.4 points lower
Chen et al. · Schrödinger's Code Repository · arXiv 2609.27891 · Table I

Tested one change at a time, the renaming did most of the damage. Remapped names alone cost each model six to seven points. Rewording the issue barely mattered. Reordering and rewriting the code cost a few points at most. The authors read this as agents leaning on names they have seen before.

Agents read far more on unfamiliar code

010203What happened

On unfamiliar code, agents read far more

Input tokens per issue
GPT-5.1178K→466K2.6×
GPT-5.4 mini66K→234K3.6×
DeepSeek-V4-Flash1.0M→3.7M3.5×

Over 80% of the extra steps went to searching, reading and probing the code.

Chen et al. · Schrödinger's Code Repository · arXiv 2609.27891 · Table I · Figure 3

The cost went up too. On the renamed code, each agent read between two and a half and three and a half times as many tokens per issue. More than 80 percent of the extra steps went to finding its way around: listing folders, searching and reading files, and running small probes. Editing and testing took about the same effort.

One Django bug: 37 steps, then 217

010203What happened
Original code: 37 stepssearched for Django's familiar method name and went straight to the faulty line
Renamed code: 217 stepslisted folders, searched source and tests, read configuration first

It submitted the same fix both times.

Chen et al. · Schrödinger's Code Repository · arXiv 2609.27891 · Figure 4 · django__django-11999

One Django bug shows the difference. On the original code, the agent searched for Django's familiar method name and found the faulty line almost at once, in 37 steps. On the renamed code it listed folders, searched the source and the tests, and read configuration files first: 217 steps. It submitted the same fix both times.

On newer issues, the score held

010203What happened

On newer issues, the score held at 17%

17.27% → 17.27%GPT-5.4 mini on 110 SWE-rebench issues created after its release

Only the cost rose: 22% more input tokens.

Chen et al. · Schrödinger's Code Repository · arXiv 2609.27891 · Table III · SWE-rebench, March 2026

Could renaming just make the bugs harder? The authors checked on 110 newer issues from SWE-rebench, created after GPT-5.4 mini was released, so it's unlikely to have seen them. Its score stayed at about 17 percent either way. Only the cost went up, by about a fifth. So the renamed repositories are just as fixable, and the extra cost comes from unfamiliar code.

What it means

03
What it means

Here's what the paper does and doesn't show.

What the paper doesn't test

010203What it means
Python projects onlyother languages may behave differently
One simple agent with command-line toolsmini-swe-agent; no IDE or language-server tools
Four models, three of them small or fast versionsthe newer-issue check used one model and 110 issues
Chen et al. · Schrödinger's Code Repository · arXiv 2609.27891 · Sections IV and VI

Here are the limits. Every project was Python. Every run used one simple open-source agent, mini swe agent, with command-line tools only, and the authors say agents that navigate with IDE or language-server tools need more work to test. Three of the four models are small or fast versions, and the check on newer issues used one model and 110 issues. It's a preprint.

What it means for picking a coding agent

010203What it means

A SWE-bench Verified score partly measures familiarity

Your codebase is unfamiliar codea private repository sits outside the training data
Expect more exploring and more tokenson renamed code, 2.4× to 3.6× the input tokens
Newer issues are a fairer testSWE-rebench, or a trial on your own repository
Chen et al. · Schrödinger's Code Repository · arXiv 2609.27891

For anyone choosing a coding agent, a SWE-bench Verified score partly measures how well a model knows famous projects. The paper didn't test private code, but a private codebase sits outside the models' training data. On unfamiliar code these agents explored far more and read several times the tokens. The authors argue for testing on code a model can't have memorized, and newer issues or a trial on your own repository do that.

The paper and its code

010203What it means

Schrödinger's Code Repository

Paperarxiv.org/abs/2609.27891
Code and datagithub.com/cslsolow/Schrodinger-Repo

Four models, one agent, Python only. Within that, every model lost points.

Chen et al. · Schrödinger's Code Repository · arXiv 2609.27891 · code and data MIT-licensed

One limit to keep in mind: this is four models, one agent and Python only. Within that, taking away the familiar names cost every model points and multiplied how much it read. The paper is on arXiv, and the code and data are on GitHub. That's how coding agents fare on SWE-bench once a repository's names are changed.