AI agents need real work files, not just instructions
Imagine giving an AI agent this task: review several reports, compare the figures, apply company policies, and prepare a table of findings. For a person, that’s routine office work. For a model, it’s a test of whether it can read files, find the relevant data, and produce an output that can be checked.
But where can we find examples that teach an agent to do this kind of work? Synthetic documents often feel artificial. Real files pose a different challenge: how can we tell whether the agent’s answer is actually correct?
The authors propose GraphForge, a way to build training tasks from real documents and tie evaluation criteria to their sources. They fine-tuned Qwen models on these examples and improved their scores on several workplace-task benchmarks. The task, the files, and the evaluation should all be grounded in the same evidence.
Why agents need real workspaces
Modern AI agents need to do more than answer questions. They read documents, use tools, work with spreadsheets, and save finished materials in the right format. Multi-step tasks are easy to get wrong: an agent might mix up data versions, overlook a policy, or make a claim the sources don’t support.
Training agents to handle this work requires examples with multiple files and a clear way to check the result. The authors point to two shortcomings in existing approaches. In one, a model generates the files. That makes it easy to create training data, but the documents can be repetitive and unlike real-world materials. In the other, tasks use real files but lack purpose-built evaluation criteria. That makes it hard to distinguish a good answer from a well-presented mistake.
GraphForge combines the two approaches. It draws files from public sources, then creates the task and evaluation criteria after researchers have assembled the workspace. The evaluator gets links to specific sources and knows where to look for evidence supporting each requirement.
What is GraphForge?
First, the system selects a task template from a database of occupations and job responsibilities. The authors ultimately retained 246 occupations across 16 sectors, 891 types of work tasks, and 3,419 valid occupation-task combinations. This gives the process a controlled starting point: the system sets the field and type of work in advance, rather than asking a model to invent tasks without constraints.
Next, a search agent gathers files for a specific example. These might include reports, spreadsheets, and other documents related to the same work scenario. Each file has a hidden role: primary source, reference material, plausible distractor, or additional context. The agent solving the task can’t see those roles.
The system then builds an evidence graph. Its nodes represent files and relevant facts within them; its links show which pieces of information need to be compared or combined. For example, a figure in a spreadsheet might need to be checked against a report, while a conclusion must also account for rules in a separate document.
The graph is turned into a task and a scoring rubric. Each criterion is linked to its sources and specifies how to check it: which files to open and which part of the output to assess. The rubric also lists errors that should lose points. The agent sees the task and workspace, but not the source links or the files’ hidden roles.
GraphForge gathers files, links facts in an evidence graph, and uses it to generate a task with verifiable criteria.
Generating a task automatically isn’t enough. First, a powerful model attempts to complete it. Then a separate review agent checks whether the task is solvable, whether the sources support the required information, and whether the criteria are clear enough. If it finds a problem, the system fixes only the relevant parts of the task or rubric instead of rewriting everything.
Finally, the result goes through two types of checks. Programs verify file formats, names, spreadsheet tabs, and formulas. A judge model evaluates the substantive criteria against the original documents. Only examples with sufficiently high scores and no obvious tool-use problems make it into the training set.
What made it into the training set
After filtering, the authors had 2,169 trajectories—sequences of agent actions from the start of a task to the finished output—selected from 3,638 prepared tasks. The dataset covers 466 types of work tasks, 15 of the 16 occupational sectors, and all 16 types of work execution.
The examples are fairly long. An average trajectory contains 50 assistant steps and about 162,000 tokens. A PDF appears in 96.7% of examples, and most tasks require other file formats as well. The model learns to draw answers from multiple documents and work with different kinds of files.
Training tasks by occupational sector, type of work, and file family.
The authors fine-tuned two Qwen models on these trajectories and evaluated them on three benchmarks: GDPVal, Workspace-Bench-Lite, and SpreadsheetBench II. They tested the models in different agent runtimes to see whether the gains came from a general improvement or simply from getting used to one particular setup.
Qwen3.6-27B gained 65.7 Elo points on GDPVal. Its score rose by 7.7 points on Workspace-Bench-Lite and 13.7 on SpreadsheetBench II. The second model, Qwen3.6-35B-A3B, also improved on all three benchmarks. The gains held up when the models were tested in other runtimes.
Fine-tuning on GraphForge improved both models’ results across all three benchmarks.
The authors also tested whether the results could be explained by memorizing the training files. None of the 39,201 unique training files matched files in GDPVal. They also checked for textual similarity. In addition, the models improved on tasks from occupations not included in the training set. This suggests that work skills may transfer, though it doesn’t prove that a model will perform equally well on every new type of work.
Can a rubric improve example selection?
After standard fine-tuning, the authors tested another approach. The fine-tuned model received 2,000 new prompts and generated up to four answers for each. The system then selected the best answers for additional training.
They compared three selection methods. The first used both the rubric and its linked sources. The second saw the rubric text but not the specific source links. The third picked an answer at random from the four options. To make the comparison fair, all three methods used the same prompts and answer sets.
On Workspace-Bench-Lite and SpreadsheetBench II, selection based on linked sources produced the largest gains. Random selection performed worse. The results on GDPVal are less clear: some differences between the methods can’t be reliably distinguished from noise. The findings suggest that selection quality matters, but they don’t establish a clear winner among the different judges.
The evaluation itself also has a limitation. If an entire worksheet is removed from the workspace, the score for a criterion tied to it drops sharply. But the judge is less likely to catch small changes to numbers or text. This check can detect when an important source is missing, but it can’t yet reliably verify every fine-grained data detail.
The average trajectory contains 50 steps; nearly a third of training examples approach the 262,000-token limit.
Conclusion
GraphForge offers a way to train AI agents on tasks built around real files and evaluate their answers against those same sources. On 2,169 examples, the models improved on three workplace benchmarks, and rubric-based answer selection provided additional gains in some evaluations.
It remains unclear how these results would change with a larger dataset. The study covers only two models from one family, and the same model is used for much of the task creation and evaluation. There’s also a risk that an automated judge will miss errors in numbers and other details. GraphForge shows how to create verifiable training data, but how accurately that data and its evaluations transfer to other models and kinds of work remains an open question.
Read next
Raven: Can the necessary AI agents be assembled automatically
Agensh: How a Thousand AI Agents Work Together Without a Boss
AgentKernel: How to Give an AI Agent More Freedom with Less Risk
How an AI agent selects past experience for a new task
JEV-as-a-Judge: The AI judge maintains 99% accuracy at a lower cost.
RRSI: How AI agent self-improvement enhances results and saves tokens
EvoOntology: Why is it difficult for AI agents to work with different
How AI Agents Can Save Context and Avoid Failures
How AI Simulates a User’s Thoughts
Coding agents skip looking at the app when the task gets long
AI covers six roles in game development, but skills rarely transfer
LLMs misread motives when the story comes through a biased user
AI reviews in simple way
Every day we read fresh AI papers and retell the essentials in plain human language — no hype, no fluff. If you want to see where AI agents are heading before everyone else, subscribe.
New reviews — every day
Follow on X