i
Research
Review · 2026-10-07

How an AI agent learns to solve problems from others’ experience

Cover: How an AI agent learns to solve problems from others’ experience

Models learn from tasks they solve themselves

The same AI can fail at a task in one harness and solve it in another. One system helps it track each step, another lets it continue after a failed check, and a third gives it more freedom. If you collect successful solutions from different harnesses and simply feed them to the model, it may learn the harnesses’ control rules rather than how to solve the tasks.

The paper’s authors propose Recursive Self-Rewrite. First, different harnesses help the model find solutions. Then the model tackles the tasks again in a single, shared harness. Only verified results are kept for training. The goal is to transfer useful techniques from external tools into the model itself.

The authors tested the method with Qwen-3.8-27B on around 3,000 command-line tasks spanning software development, science, and data work. After rewriting, the number of successful trajectories grew from 2,001 to 11,094. Fine-tuning on those trajectories improved the model’s performance on several benchmarks, including challenging tasks that require long sequences of actions.

Why one harness isn’t enough

A harness defines how a model interacts with tools: what it sees, how it tracks its work, when it checks results, and what it does after an error. It’s like the difference between working without a plan and working with a checklist. The person doing the work is the same, but the process helps them organize their actions in different ways.

The authors compared three harnesses. Terminus 2 provides a general command-line workflow. StateM tracks task state and checks transitions between steps. Recursive Self-Reflect Terminus (RSRT) checks the result and, if it fails, asks the model to keep working.

At first glance, it might seem enough to collect all the successful attempts from these systems and use them for training. But those attempts may contain prompts or control actions available only in the original harness. If the model learns to rely on them, it may get stuck repeating steps or stop too early when that support is gone.

The authors’ goal is to preserve useful experience without making the model dependent on a particular harness.

How experience rewriting works

The process has three roles, all performed by the same underlying model.

First, a planner studies a successful attempt and writes instructions: a list of key steps, checks, possible errors, and ways to fix them. The instructions should describe the process without giving away the answer.

Next, a critic reviews the instructions. It rejects anything that reveals the solution or relies on information unavailable from the task or environment. Finally, an executor receives the approved instructions and tackles the task again, in a clean environment using the shared Terminus 2 harness.

Only a new attempt that passes the check is added to the training set. The instructions and the exchange with the critic are discarded. At deployment, the model must handle the task without them.

Different harnesses find successful solutions; rewriting turns them into verified examples for a shared harness.

“Recursive” refers to repeating the cycle: the model generates several sets of instructions, checks them, and tries the task again. The authors used up to four instruction variants and up to four new attempts for each successful original trajectory. Failed attempts were useful too: their check results helped identify examples suitable for training.

Three harnesses find different solutions

The combined dataset contained 2,929 tasks. Each of the three harnesses solved at least some of them, and together they solved 759. That’s 34.3% more than the best individual harness in the dataset.

Each system also had successes of its own. Terminus 2 solved 129 tasks that the other two harnesses couldn’t. RSRT solved 91 such tasks, and StateM solved 68. This diversity helped in more ways than simply providing extra attempts: different control strategies helped the model tackle different tasks.

Tasks solved by each harness alone and in combination with the others.

Each harness had its own working style. RSRT could continue after a failed check: in 4.5% of its successful attempts, the first version of the answer failed, but the model corrected it. StateM devoted a significant share of its commands to managing task state. Terminus 2 more often explored the environment directly.

But more steps don’t always mean more success. Among failed attempts, RSRT averaged 67.1 turns—three times as many as Terminus 2 and StateM. The extra time gave the model a chance to correct mistakes, but it could also be spent without making progress.

Harnesses support different strategies: a general workflow, explicit step management, and continuing after a failed check.

What another attempt achieved

Starting from the original successful solutions, the authors generated 11,094 new, verified trajectories—more than five times the original 2,001 successful attempts. But more data alone didn’t guarantee an improvement. What mattered was having the model solve the tasks again in a shared harness and keeping only the successful results for training.

After fine-tuning on the rewritten trajectories, the model outperformed both the baseline and a version trained directly on the original attempts. The pass@3 metric measures whether a task is solved in at least one of three attempts. On Terminal-Bench 2, it rose from 57% for the baseline to 74.2%. On Terminal-Bench 4, it increased from 1.5% to 9.1%. On the authors’ own set of challenging tasks, it climbed from 39% to 63%.

On a software engineering benchmark, the score rose from 3% to 6%. The gain was more modest, but the model still solved twice as many tasks out of 100. On a benchmark involving long sequences of actions, its progress score increased from 0.21 to 0.29. However, none of the versions fully solved any of the benchmark’s 46 tasks.

Repeated attempts based on the same instructions can succeed or fail.

The comparison with direct training shows why rewriting matters. Training directly on the original trajectories helped on several benchmarks, but performance on Terminal-Bench 2 fell from 57% to 53.4%. The authors found repeated actions in those attempts that made little or no progress. Some RSRT trajectories were especially long, and the model may have picked up their dead-end habits.

Rewritten trajectories didn’t always succeed either. In an example involving Markdown parsing, seven of 12 repeated attempts passed. In an OpenFOAM task, seven of eight did. The same instructions could lead to either success or failure. That’s why checking the result remains essential: even careful instructions can’t replace successful execution.

Conclusion

The results show that harnesses can serve not only as a way to help a model while it works, but also as a source of training experience. Different systems find different solutions, and having the model repeat a task helps separate transferable techniques from the quirks of the original harness.

But these findings apply to the experiments the authors conducted. The study focuses on command-line tasks and a single base model. Some of the data comes from the authors’ own benchmark, and the method requires many additional attempts: of nearly 12,900 reruns, 86% succeeded. What’s more, passing an automated check doesn’t guarantee that a result is good in every respect.

Models don’t have to learn only from answers they’ve already produced. Powerful external tools can help find a solution; then their support is removed and the model is asked to solve the task again. If that second attempt passes the check, the experience is closer to what the model should be able to do on its own.

AI reviews in simple way

Every day we read fresh AI papers and retell the essentials in plain human language — no hype, no fluff. If you want to see where AI agents are heading before everyone else, subscribe.

New reviews — every day

Follow on X