i
Research
Review · 2026-10-09

How to Help an AI Agent Stay Oriented in a Large Project

Cover: How to Help an AI Agent Stay Oriented in a Large Project

When Every File Gets an Assistant

In a large project, fixing a bug means understanding more than just the code where it occurred. You need to know which settings affect it, which parts of the program call that code, and which tests check the result. The longer an AI agent works on a task, the more information it has to gather and keep in context. Eventually, it may miss a connection between files or lose track of its own steps.

The article’s authors propose a different way to work with a repository. Instead of having one agent repeatedly reread the relevant materials, they create Dev-Primitives: software components, each tied to a single file and its own language model. These assistants study their file, modify it, and tell other components what changes they need to make.

The idea underpins HERMES, a framework that selects the right components, runs them, checks the results, and directs the next attempt to wherever a problem remains. Across four benchmarks, HERMES improved on comparable systems by an average of 12.4 percentage points.

From Passive Files to Active Participants

A conventional agent treats files as source material: it reads them, keeps some of their contents in context, and decides what to change. On a long-running task, it has to reconstruct the project’s structure after every change. The context keeps growing, and details about requirements and connections between files can get lost.

In HERMES, each project component becomes a Dev-Primitive—an assistant tied to a specific file. That file might contain source code, configuration, a build script, or a test. The component gets a local task, studies its file, and modifies it if needed.

If a change affects neighboring files, the assistant sends them a message in plain language. For example: “The data format has changed; update your handling accordingly,” or “Check the new behavior in the tests.” Messages go only to components affected by the change. A file doesn’t need to restate the entire project plan; it passes along the requirements it understands best.

HERMES workflow: the system selects files related to the task, assigns them local changes, runs the project, and directs further attempts based on the results.

Not every file gets an assistant at once. First, HERMES analyzes the task and identifies the components likely to be involved in the behavior in question. It then traces the connections—to calling code, configuration, and tests. Assistants are launched only for the selected components.

After the changes, the system runs the project and gathers whatever results are available: command errors, test results, logs, and traces. A separate diagnostic module works out which components may have caused any remaining problems. On the next round, HERMES returns to those components and brings in additional files if needed. By default, the system can run three such rounds after the initial attempt.

What the Tests Showed

The researchers tested HERMES on four benchmarks covering bug fixes in real-world projects, migrations of entire codebases to a new interface version, command-line tasks, and software build-and-operations workflows.

The main comparison keeps the AI model the same for both the baseline system and HERMES. Only the framework changes. The authors therefore attribute the difference in results to how the work is organized, not to a switch to a more powerful model.

On SWE-bench Verified, which tests bug fixes in real-world projects, HERMES increased the share of solved tasks in 19 of the 20 model configurations tested and matched the baseline in the remaining one. With GPT-5 Mini, the gain was 19.4 percentage points. With GPT-5.6 Sol, the score rose from 96.2% to 97%.

The effect was more pronounced in the codebase-migration benchmark. In one comparison using GPT-5.6 Sol, the share of tasks with a nonzero score rose from 6.5% to 31%. This kind of work depends heavily on keeping changes consistent across files: update the main code but miss the settings, call sites, or tests, and the project may not work.

On command-line tasks, GPT-5.6 Sol with HERMES solved 51.8% of tasks, compared with 33% for the baseline framework. In build-and-operations workflows, the model’s average score rose from 49.89% to 55.41%.

The gains come at a cost. HERMES calls the model more often and uses more tokens because it has to activate multiple assistants, run the project, and investigate errors. So higher accuracy doesn’t mean lower costs in every case. The system does, however, let users choose where to use a larger model and where a smaller one will do.

Where a Larger Model Pays Off

The authors also tested combinations of models of different sizes. A smaller model, Qwen3-8B, could work inside Dev-Primitives, while larger models selected files for the task and analyzed the test results. This division proved useful: the experiments suggest that diagnostic quality has a greater effect on results than upgrading the file-selection module.

In a Terminal-Bench test, a combination of GPT-5.6 Sol for selecting components, Qwen3-8B for working with files, and GPT-5.5 for diagnostics solved 49.7% of tasks. A uniform setup using GPT-5.6 Sol for every role reached 51.8%. That’s a difference of 2.1 percentage points, while inference costs were 37.4% lower. The results suggest that a system like this can reserve its strongest model for decisions that matter most to success, while assigning local changes to a lighter one.

Ablation tests show which parts of HERMES make the biggest difference. Performance dropped most when the system stopped analyzing execution results and making targeted recommendations for the next attempt. On the codebase-migration benchmark, removing diagnostics cut the score from 31% to 20%. When assistants were prevented from messaging one another, the score fell to 23.5%.

In HERMES, file selection, local changes, code execution, and diagnostics form a repeating work cycle.

Activating only the components that are needed also matters for cost. When selected assistants stayed active for every round, the average number of calls per task increased by a factor of 1.6–1.9. Performance also declined. Selecting components again as new evidence emerged helped avoid spending resources on files that were no longer relevant.

What Remains Difficult

HERMES depends on two decisions: which files to bring in and how to interpret the results of running the project. If the selection module misses an important part of the codebase, the assistants may fix only part of the problem. If test output and logs offer little explanation, the system will struggle to work out where to go next.

The tests themselves also have limitations. The authors use files as the main units of work, though some tasks might be better handled by splitting a project into smaller pieces. And even a diverse set of benchmarks cannot cover every kind of software work. Results with Qwen3-8B also depend on the specific model configuration used.

An agent’s performance depends not only on which model writes the code, but also on how the system distributes the work, coordinates changes across files, and uses the results of checks. HERMES shifts some responsibility for the project’s state from a single general-purpose agent to components that know their own files better. This approach is especially useful for long-running tasks, where the problem often lies not in one place but between several connected parts of a program.

AI reviews in simple way

Every day we read fresh AI papers and retell the essentials in plain human language — no hype, no fluff. If you want to see where AI agents are heading before everyone else, subscribe.

New reviews — every day

Follow on X