AI agents are teaming up
A complex task often takes more than one all-purpose agent. Building a game, for example, means studying the requirements, writing code, creating graphics, assembling the project, and checking the result. Each step calls for different tools and skills. Just as important, the results need to be passed along in a way that makes it clear to the next person—or agent—what to build on.
The authors of Raven propose building systems this way: as teams of specialized AI agents that can work together. A research agent finds information, a coding agent writes and checks code, and a design agent prepares visual assets. A separate orchestrator breaks the request into tasks and connects them in an overall plan.
The key idea is to treat the model and its harness as a single working module. The harness includes tools, memory, execution rules, and checks. That makes it possible to plug in different agents without forcing them all to follow the same workflow.
How Raven builds a work plan
The orchestrator represents tasks as a graph. Nodes are calls to individual agents; arrows show which results are needed for later steps. Independent tasks can run in parallel, while dependent ones start once their inputs are ready.
For example, if the goal is to build a website about a research project, two research agents could independently look into the algorithms and how to evaluate them. A coding agent would wait for both reports before putting together an experiment. A designer would create a page based on the results, and the orchestrator would bring the finished materials together.
Before execution, the system checks the plan’s structure: Are all tasks defined? Do the required dependencies exist? Is the selected agent suitable, and can it work with the specified files? These checks help catch problems in the graph before agents spend time and compute on it.
Results are passed along as files and records with clear links between them. That way, an agent doesn’t have to rely on the orchestrator’s summary or receive a long report all over again in its prompt. Each task keeps its instructions, result, and execution history. If the project needs to continue later, the system can reuse a completed result or resume an agent with its context intact.
In Raven, specialized agents handle focused tasks, while the orchestrator connects their results into an overall plan.
The system also guards against cases where an agent claims a task is done even though the required file or result is missing. After a call, a separate reviewer checks the request, the response, and part of the execution history. If something is missing, the orchestrator can give the agent further instructions, drop the task, or revise the rest of the plan.
But this check evaluates whether the task was completed as stated; it doesn’t necessarily establish whether the result is correct. Code still needs tests, research needs source checks, and design work needs a review of the finished image.
The harness can improve
Raven includes a method for changing an agent’s harness while leaving the underlying language model untouched. Four parts can be adjusted: memory, planning, tool selection, and the agent’s actions, including how it checks whether a task is complete.
If an agent regularly finishes without providing an answer, or fails to run tests after changing code, the system looks for a way to correct that specific behavior. It generates harness variants, checks whether the change worked, and compares each version with the previous one on the same tasks. The strongest variants go through more extensive checks on the training set. The selected harness is then frozen and evaluated on separate tasks that weren’t used for tuning.
In experiments with HarnessBank, an improved harness raised the task success rate across seven evaluations while keeping the Qwen3.6-27B model unchanged. The gain was 15.4 percentage points on AppWorld, 13.9 on BrowseComp+, and 13.7 on LiveCodeBench. On SWE-bench Verified, the score rose by 5.1 points, but that difference did not meet the authors’ criterion for a significant improvement.
The harness changes around an unchanged model, and new variants are tested before evaluation on held-out tasks.
The method has limitations. It selects variants based on training results, so it can’t guarantee that improvements will carry over to new tasks. The authors also note that detailed settings for some HarnessBank experiments have not been fully published. The figures therefore demonstrate the approach’s potential, but don’t give a complete picture of what it cost to achieve each result.
Memory stores experience; skills turn it into instructions
Raven has two kinds of memory. The first stores information about the user: their preferences, constraints, and the context of earlier conversations. The second records the agent’s own experience: what task it handled, what it did, and what lesson might be useful next time.
First, conversations are split into meaningful segments. The system creates short summaries of those episodes and extracts individual facts. Related episodes are grouped together, so it can retrieve not only a specific fact but also the context in which it came up. When searching, Raven uses both keyword matches and semantic similarity, then returns the fact along with the episode it belongs to.
Repeated experience can become a skill: an instruction for a particular class of tasks. For instance, cases where an agent needs to modify a project and check the change could gradually become a procedure: reproduce the bug, make a small change, and run the tests.
Raven looks for skills in three places: a local library, the agent’s memory, and the shared SkillHub catalog. It combines the materials it finds, then an orchestration component selects those that fit the current task. If the selection fails, the system follows a fallback rule and takes several of the highest-rated options from the list.
SkillCorpus, the catalog used in some of the experiments, contains 96,401 active skills. Before being added, materials were checked for duplicates, quality, and potentially dangerous instructions. In experiments with a fixed library, skills improved results across all three evaluations: SkillsBench, GDPval, and QwenClawBench. On SkillsBench, Raven’s scores rose by 6.5 and 13.4 percentage points for the two model sizes. Gains on GDPval were more modest, ranging from 1.2 to 1.9 points across all tested agent-and-model combinations.
Memory search first finds relevant episodes, then selects facts while preserving their original context.
What the evaluations show
To test the orchestrator on its own, the authors created a benchmark of 140 tasks. For each one, the system had to choose specialists and specify which results needed to come before others. Raven was compared with Claude Code and Hermes Agent on two language models. The agents themselves were not run; only the plans were evaluated.
Raven achieved the best score on all four metrics with both models. On one model, Exact Match—the share of plans that matched an acceptable reference in both agent selection and dependencies—was 71.1%, compared with 60.7% for the strongest competitor. On the other, Raven scored 86.7% versus 76.2%. This suggests Raven was more accurate at choosing agents and connecting their tasks. But the benchmark evaluates plans, not the quality of the team’s final work.
The specialist agents were evaluated separately. Raven-Research answered questions using information from the web. On DeepResearch Mixed, running DeepSeek-V4-Flash, it achieved 76.5% accuracy. The strongest agent tested on the same model scored 68.9%. Raven’s answer was estimated to cost $0.0242—slightly more than one competitor and less than another. Five of the six comparisons using the same models showed a statistically significant advantage; in the sixth, a 3.3-point difference did not pass that test.
Raven-Code was tested on programming tasks, including bug fixes and porting an entire project to a new technology stack. On SWE-bench Pro, it solved 15 more tasks than Claude Code using the same model. On SWE-Refactor, Raven’s results were 9.5 points higher than published results from other systems with DeepSeek-V4-Flash, and 3 points higher with GPT-5.6 Luna. These comparisons need some context: some use results reported elsewhere, and the systems were run with different settings and conditions.
Raven-Design creates slides, websites, and visualizations, with a preview of the result. It scored 80.2 on PresentBench with Claude Opus 5 and 72.9 with GPT-5.6 Luna. In the latter case, Claude Code scored 52.4. The evaluation considers more than appearance: it also measures completeness, accuracy, and fidelity to the source material.
Raven-Oncall handles long-running computational tasks, such as launching a training job and monitoring its progress. In a set of 17 scientific tasks, it completed 14 with Claude Opus 5; Claude Code completed 11. Comparisons like these depend on the tasks, tools, and evaluation rules, but they illustrate why a system might need a dedicated agent for work that runs for hours.
Conclusion
Raven proposes building AI systems from specialized agents connected by explicit dependencies and preserving their experience for future tasks. The orchestrator creates a plan, agents pass results along as artifacts, and harnesses and skills can change without updating the underlying model.
The results look promising, especially for planning and specialized tasks. But the evaluations used different datasets and protocols. Some comparisons rely on published results for other systems, and the full set of memory and feedback layers has not yet been evaluated together.
One practical question remains open: do several agents deliver a better result than one well-equipped agent when the total cost is counted—including planning, checks, retries, and passing results between agents? Raven’s framework describes the conditions under which collaboration could expand the range of tasks a system can handle. Future head-to-head evaluations will show how often those conditions hold in practice.
Read next
Agensh: How a Thousand AI Agents Work Together Without a Boss
AgentKernel: How to Give an AI Agent More Freedom with Less Risk
How an AI agent selects past experience for a new task
JEV-as-a-Judge: The AI judge maintains 99% accuracy at a lower cost.
RRSI: How AI agent self-improvement enhances results and saves tokens
EvoOntology: Why is it difficult for AI agents to work with different
How AI Agents Can Save Context and Avoid Failures
How AI Simulates a User’s Thoughts
Coding agents skip looking at the app when the task gets long
AI covers six roles in game development, but skills rarely transfer
LLMs misread motives when the story comes through a biased user
Given a store for a year, the top-earning agent ranked 16th of 18 on fraud
AI reviews in simple way
Every day we read fresh AI papers and retell the essentials in plain human language — no hype, no fluff. If you want to see where AI agents are heading before everyone else, subscribe.
New reviews — every day
Follow on X