i
DATAIST
Back to feed

Research

Breakdowns of recent AI research: what the authors did and why it matters.

211 articles

How a virtual company teaches agents to take market reactions into account

An AI-run company cannot learn from business decisions if it never sees how the market responds. MiniCorp gives AI agents a simulated e-commerce business to run, with customers, competitors, and market conditions that react to their choices. The simulation records what the agents knew, what they decided, and what happened next; researchers can also replay the same situation with different decisions to compare possible outcomes. This gives agents practice with the long-term consequences of business choices, rather than only examples from incomplete historical records. In this review we look at how MiniCorp connects a company’s internal decisions with an evolving market, and how that setup could help train and evaluate AI agents for real-world business work.

How an AI agent learns to solve problems from others’ experience

An AI agent can learn from another agent’s hard-won success—but only if it can separate the useful know-how from the special setup that made it work. The authors propose a way to turn successful task-solving attempts into reusable instructions for an AI model. One part extracts practical steps, another checks that the instructions do not rely on hidden answers or tools that will not be available later, and a final part tests them in a fresh environment. This lets the model learn from successes gathered under different setups and carry that experience into a more general one. In this review we look at how recursive rewriting turns a small collection of successful attempts into a much broader set of training examples, and why that helps an AI agent tackle more challenging tasks.

How to create a video from a spreadsheet without spending hours editing it

A spreadsheet can hold a story, but turning it into a polished video usually takes far more than pressing play. The authors introduce DataMagic, a system that turns raw tables into videos with charts, narration, and animations that stay in sync. It first drafts possible scenes, then arranges them into a coherent story, while keeping every visual tied to its source data so the facts can be checked. People can also step in to guide or refine the process instead of handing everything over to AI. In this review we look at how DataMagic brings data, storytelling, and video editing together, and how it helps people create data videos faster without losing control or accuracy.

How can you check whether an AI agent has completed the task?

How can you tell whether an AI agent has truly completed a task when the answer depends on a pile of real files? The authors introduce GraphForge, a way to create training tasks from real workspaces and check them against evidence in those files. It links each task requirement to the material needed to verify it, then tests and repairs the task before using it to train an agent. This gives researchers a way to teach agents to handle messy, file-based work and judge the results more reliably. In this review we look at how GraphForge builds tasks and checks their answers, and what the results suggest about training more dependable AI agents.

Can the necessary AI agents be assembled automatically as the work progresses?

What if an AI could build the right team of specialist agents while it works, instead of relying on one setup for every job? The authors introduce Raven, a system that creates and improves tailored AI agents for different models and fields, then brings them together to handle larger tasks. It breaks a goal into smaller steps, assigns each to a suitable agent, and carries useful experience forward so future teams can work better. This could help AI tackle complex projects that no single specialist can manage alone. In this review we look at how Raven builds its agents, coordinates their work, and learns from past tasks—and at the evidence that this approach can widen the range of problems AI systems solve reliably.

How a Thousand AI Agents Work Together Without a Boss

What if a thousand AI agents could tackle a hard problem together without waiting for a boss to assign every task? The authors propose a system where agents organize their own work: they pick up tasks, share discoveries, check one another’s results, and combine progress through a common workspace and messaging tools. By working at the same time, they can finish complex tasks sooner, while adding more agents can improve the chance of getting a working result. In this review we look at how this self-organized approach works, what happens as the team grows, and where it could help when time is tight.

How to Give an AI Agent More Freedom with Less Risk

The more freedom an AI agent gets, the more ways there are for hidden instructions to lead it astray. The authors propose AgentKernel, an operating-system layer that protects an agent from the moment it reads outside content to the moment it uses tools. It controls who the agent can act as, what information it trusts, what it keeps in memory, and which actions it can take—so security cannot simply be bypassed by the agent itself. The goal is not just to restrict agents, but to make it safer to give them broader responsibilities. In this review we look at how AgentKernel brings familiar computer security ideas to AI agents, and why protection built into the foundation could make autonomous systems more capable as well as safer.

How an AI agent selects past experience for a new task

An AI agent can learn from the past and still remember the wrong things for the job at hand. Rather than turning each finished task into a fixed note, the authors keep the full record and select what matters only when a new task arrives. The agent then shapes those past experiences into a concise guide for its current goal, so useful details are not discarded before anyone knows when they might help. Tests across household, shopping, and computer-use tasks show that this just-in-time approach helps agents succeed more often than existing memory methods. In this review we look at how task-specific memory is assembled, why choosing what to remember later can work better, and what the results reveal about helping AI agents learn from experience.

The AI judge maintains 99% accuracy at a lower cost.

Can an AI judge keep its accuracy without the cost of a large, talkative model? The authors test a leaner evaluator that focuses on making a verdict rather than generating a full explanation, using it as a cheap first pass across different tasks. When the system is unsure, a stronger judge takes over, preserving nearly all of the stronger model’s accuracy while sharply reducing the expense. The approach works especially well for ordinary comparisons and fact checking, but can struggle with mathematical derivations or confidently presented wrong answers. In this review we look at how the two-stage AI judging system works, where its weak spots appear, and why confidence-based escalation could make large-scale evaluation far more affordable.

How AI agent self-improvement enhances results and saves tokens

What if an AI agent could improve the way it works—not by changing its core model, but by redesigning the instructions, tools, memory, and workflow around it? The authors propose a more disciplined form of self-improvement that helps an agent test and refine these surrounding components without simply memorizing the tasks it was trained on. The system limits how many changes it makes at once, explores new strategies, and removes edits that are costly, trivial, or useful only for a specific benchmark—leading to more reusable behavior and fewer tokens spent during operation. In this review we look at how regularized self-improvement works, why unconstrained evolution can fail outside familiar tasks, and how the proposed approach builds leaner, more adaptable AI agents.

Why is it difficult for AI agents to work with different types of data?

AI agents can sound fluent yet still struggle to find their way through a jumble of tables, files, and databases. The authors propose EvoOntology, a layer that gives an agent a living map of what data exists, what it means, and which tools can help it use that data. Instead of relying on fixed instructions or forcing the agent to inspect every source from scratch, the system builds and continually improves this map as the agent works, adapting it to different tasks and ways of reasoning. This helps turn scattered, hard-to-navigate information into something an AI agent can actually explore and use. In this review we look at how EvoOntology is built, how it learns to refine its own understanding of data, and why this approach may make AI agents more capable across mixed data sources.

How AI Agents Can Save Context and Avoid Failures

Autonomous coding agents can lose their way not because they cannot write code, but because they run out of room to remember what they are doing. The authors examine how the surrounding software harness—the tools and rules that guide an AI agent—affects its ability to solve long, complicated programming tasks. They compare ways to manage context, plan work, and choose actions, showing that the best setup depends on the model’s strengths and on how much memory it has available. The results reveal when planning prevents mistakes, when it mainly saves time and cost, and why simpler command-line control can sometimes work better than a larger toolbox. In this review we look at how these design choices shape an agent’s entire problem-solving path, and what they suggest for building coding systems that stay effective without wasting context.

How AI Simulates a User’s Thoughts

AI can learn to understand not only the words a user types, but what that person actually thinks, wants and is trying to do. The researchers propose first building believable user models that keep a person's traits and change their thinking as the conversation goes on. A separate «ideal assistant» then answers with those hidden intentions in mind, and an ordinary model learns to help just as well without direct access to the person's thoughts. That makes the assistant fit the person better, explain their decisions and hold on to their goals over a long exchange. In this review we look at how such a simulation of human thought is built and why it may change the way AI assistants are trained for study, work and everyday tasks.

Coding agents skip looking at the app when the task gets long

A coding agent is usually evaluated the simple way: hand it a task, a bug report or a test suite, then check whether it fixed the code. A new paper, ProgramDistill , proposes a setup much closer to real work. You have a working reference app. You also have a broken or incomplete version of it. The agent never sees the reference's source. It has…

AI covers six roles in game development, but skills rarely transfer

When people talk about AI in games, they usually mean bots that learned to win at chess, Go or StarCraft. A new paper argues for a wider frame. AI in games today is not only the thing that presses buttons better than a human. It is also the thing that models the game world, shapes the story around the player, helps assemble prototypes, writes…

LLMs misread motives when the story comes through a biased user

People have been asking LLMs about far more than code, emails and spreadsheets for a while now. They ask about exes, coworkers, friends, bosses, jealousy, flirting, hidden conflicts. And that creates a problem: the model almost never sees the situation itself . It sees a retelling. A retelling by someone who may have forgotten details, filled in…

Given a store for a year, the top-earning agent ranked 16th of 18 on fraud

Most agent benchmarks test the short distance. Fix a bug. Find an answer. Walk through a set of steps. Even when there are many steps, the task usually collapses into a single final result. E-Commerce Bench looks at a different problem: what happens when you hand a model not a 20-minute task but a business to run for 365 days . With money…

Five levels of self-improving AI, and why level 5 barely exists

The idea is old and unsettlingly simple: at some point an AI stops merely solving tasks and starts improving itself . Not in the sense of "rewrote that answer a little better," but in the sense of "changed its own way of learning, checking itself, accumulating experience and building the next version." That is the subject of a long paper with a…

An AI agent built a playable shooter over 70 autonomous iterations

One of the main problems with coding agents has been obvious to anyone who has tried handing them something bigger than a 50-line function. On a short task everything looks brisk. On a long one the agent tangles itself in its own steps, fixes one thing and breaks another, loses sight of the original requirements, and declares the work finished…

An editable graph of next steps beats memory for long-horizon agents

AI agents have a recurring problem. While the task is short, everything looks fine. The model reads the history, picks the next step, calls a tool, moves on. But once the horizon gets long, things start to break: the agent confuses the order of actions, repeats useless steps, forgets what it has already tried, and does too late what it should…

A library rebuilt from 50 design docs matches its hand-checked models

Developers are used to treating code as a project's primary artifact. Documentation sits beside it, tests sit beside it, architecture notes live off somewhere to the side. The authors of Design Docs Are All You Need propose flipping that order: make design documents the primary thing and rebuild the library itself from scratch every time using AI…

Compiling a paper into a repo-level spec cuts AI's algorithmic shortcuts

Picture the task: you hand an AI a machine learning paper and ask it to build an entire repository from it. Not one file — a real project: the model, data loading, training, evaluation, run scripts. On paper it sounds straightforward. In practice the AI almost always starts simplifying. Somewhere an important detail of the algorithm goes missing…

Imagining the poster first lets a coding agent build it in editable layers

Image generators have learned to make beautiful posters. Coding agents have learned to assemble tidy pages in HTML and CSS. Between those two worlds, though, there has always been an awkward gap. An image looks striking, but it is a flattened bitmap: the text inside it often breaks, the layers cannot be pulled apart, the headline cannot be…

Distilling 1,000 GitHub repos into agent skills more than doubles MLE-bench scores

AI agents can already write code, run experiments and even attempt to reproduce papers. But on long research tasks they keep hitting the same problem: they know the general ideas but are bad at getting them into working shape . That is the difference between "I've heard of this method" and "I know which package to install, what data format it…

A fine-tuned student simulator trains a better AI tutor than GPT-5.4

A personalized AI tutor sounds simple: the system sees exactly where a particular student is stuck and tailors the explanation to them. In practice it all comes down to one thing. Figuring out which explanation helps which student takes slow, expensive data from real people . The authors of StudentSim propose a way around that: build the AI tutor…

Coding agents hit 80% on single tasks, 38% over a whole project

For half a century the industry has run on a simple rule: a person breaks a task into parts, writes code, then fixes and rewrites it as things change. Zhenfeng Cao's paper argues something different: the era in which code was the main carrier of logic is ending . What moves to the front is the AI agent — a system where the LLM does not merely…

An agent recovers physics from video by writing simulator code

Large multimodal models can already describe what happens in a video: a ball rolls, a cup falls, a car turns. Physics is where the old problem persists. They see the phenomenon, not the mechanism. They can give a tidy account of a clip but don't always understand why the object moved the way it did, how fast it was going, what would change if you…

Keeping a wiki of failed attempts makes agent skills improve faster

AI agents can already search the web, write code, work with files and handle multi-step tasks. But an old problem remains: they are bad at accumulating experience . An agent fails a task, someone reads the logs, patches a skill, runs it again — and half the useful observations dissolve back into the history of iterations. On the next cycle the…

Scoring the whole session makes a shopping agent behave more like a user

A recommender system can predict the next click perfectly and still barely understand the user. A shopper scrolls the catalog, goes back to an item they already looked at, compares two similar models, puts the purchase off, then suddenly adds something to the cart. A session like this is full of pauses, repeated actions and stray movements. Train…

Apodex 1.1 gains from agent coordination, but full research runs still fail

Ask a model to write an email, answer a question, even solve a coding problem, and everything looks fine. But the moment a task runs for an hour — reading files, running code, hunting for sources, surviving failures, and finally handing back something you can check — most systems start falling apart. That is what the paper Apodex 1.1: Scaling…

Multi-agent systems need explicit graphs, not smarter agents

Over the past year the industry has picked up an odd habit: when an agent fails a task, you give it more tools, more memory, a bigger context window and another try. Sometimes that works. But only up to the point where the task stops looking like a conversation with a smart model and starts looking like the work of a small team. Fixing a bug in a…

Evolving the scaffolding around a frozen model adds 17 points

AI agents come with an awkward truth: quality doesn't depend on the model alone. The scaffolding often decides everything — the system prompt, the tools, memory, the rules for choosing the next step, the checks before answering. The same LLM can behave like a careful engineer or like a chaotic intern purely because of how the pipeline is…

Strong-model scaffolding lifts a weak model from 0.49 to 0.91 without retraining

Transferring capability from a large model to a small one usually looks like this: take the strong model, use it to generate data, then fine-tune the weak one. This paper proposes a different move. What if you never touch the weak model at all? No weight updates. No fine-tuning. No long training run. Instead, ask the strong model to build the…

Self-improving agents stall unless the environment changes too

AI agents have an old problem: they are usually trained to get better inside a world that barely moves. The tasks are fixed. The evaluation is fixed. The opponent, if there is one at all, is fixed too. An agent can grow inside that setup, but it hits a ceiling fast. The authors argue for a wider frame: real long-run progress starts where it is…

Combodied agents: measuring help by what the person keeps

Picture a simple scene. An older person has missed a dose of medication. An ordinary digital assistant sends another notification. A robot might roll over and bring the pills. Neither one understands the thing that actually matters: did the person forget? get confused? feel unwell? change their mind and decide to skip it? That gap is where the…

Coding agents solve just 41% of tasks in a new refactoring benchmark

Benchmarks for coding agents follow a familiar arc. First they push the field forward. Then models catch up fast, scores climb, and the metric stops telling systems apart honestly. At the SWE-bench level you can already see it: frontier models pass those tasks more and more often, and some of the unsolved examples turn out to be broken by bad…

8.3 billion simulated users, and the model playing them changes the verdict

Almost every evaluation of AI systems and digital products suffers from the same disease. It tests an averaged abstraction — as if everyone had the same level of experience, the same conversational style, the same tolerance for errors, latency and strange answers. In practice that distorts the picture. A beginner working through a coding task…

The best AI agent scores only 66% when data is spread across files

Picture an ordinary work request: compute a fund's risk, pull the right equity records, assemble a medical summary. In practice the answer almost never sits in one clean table. Part of it hides in SQLite, part in a long PDF, some of the criteria are spoken aloud in a video, and the question itself may be in a different language. For a person this…

85.7% of what AI cites about a brand comes from someone else's site

Who actually tells AI about your brand When you ask an LLM about a company, the model rarely answers from memory. It pulls sources off the web first, then assembles an answer. And that is where it gets interesting: a brand's reputation with AI is shaped less by the company's own site than by everyone else's. 💬 A question about a company …

Over a simulated year, the best AI agent reached 27% of human net assets

Almost every popular benchmark for AI agents tests a short distance. Click a button. Call a tool. Fill in a form. Arrive at the right answer. But plenty of real tasks work differently: you make a decision today, the money leaves the account immediately, and the mistake surfaces a week later. By then it has spoiled more than one order — it has…

SpyRL turns open-ended tasks into a spy hunt with a checkable reward

Large language models have an old problem. They learn well wherever the answer can be checked exactly: math, coding problems, formal puzzles. The answer is either right or it isn't. The machine gets a clean signal and improves. But the moment a task turns open-ended — write a story, summarize a document well, produce a coherent explanation —…

Qwen-UI-Agent scores 92.2% on real Android phones, not simulators

Everyone wants a general-purpose AI agent. The kind you can hand a problem to and say "sort it out." It opens the apps itself, finds the right buttons, compares the options, fixes a file on your computer, and when a flight-cancellation notice lands, it comes back with a plan ready to go. The trouble is that almost all such demos look good on…

Writing a spec before code lifts coding agents' test pass rate by 21%

AI agents are already decent at fixing code, filling in functions and finding their way around someone else's repository. Ask them to build a program from scratch and the picture changes sharply. Especially when there are no sources at all — just a README and a working binary you can run as a black box. On tasks like these, even the strongest…

AI agents follow the company handbook just 36% of the time

Picture an office AI agent. It has been given access to email, the calendar, chat, tickets and a folder of documents. Next to all that sits an 80-page company handbook. The instruction: work by the rules. By now this looks like an ordinary deployment. It is exactly how companies are trying to put AI agents to work. But there is an uncomfortable…

JarvisHub replaces the chat log with a canvas an agent can edit

AI can turn out images, video, websites, slides and music from a single prompt. Real creative work does not run that way. You almost never reach the final version in one sentence. There are references, rough drafts, versions that worked and versions that did not, edits, branching revisions, notes from colleagues and a pile of small decisions…

Letting an agent decide when to compress its own context beats a fixed threshold

AI agents have a problem on long tasks: they get tangled in their own steps. Web searches, documents read, attempts, rollbacks, fresh hypotheses — all of it piles up in the context. At some point the history gets too long. The model either hits the limit or simply starts thinking worse because there is too much noise around it. A new paper from…

Finding the causal step fixed three times as many failed agent runs

There is an uncomfortable truth about AI agents: the failure almost never happens where you see it. The final answer can be wrong because 20 steps earlier the agent missed a constraint in the task, pulled the wrong record out of memory, or handed off to another agent without the context that mattered. The logs show you the symptom. The cause sits…

A DAG-editing agent matches script baselines at 42.8% lower cost

LLMs already write decent code from a natural-language description. In a production setting that is often not enough. A user asks: "build me a pipeline for data cleaning, filtering, question generation and quality checks." The coding agent answers with a Python script. The script runs. Sometimes it even works. Then the familiar pain starts: it is…

EvolvingWorld makes models track the world, not just play the character

The usual problem with almost every literary AI simulation is this: characters can talk in character, but they cannot keep living. They copy the speech patterns of Sherlock Holmes or Elizabeth Bennet well enough, and then after a few scenes they start coming apart. One suddenly forgets his own motives. Another's personality shifts for no reason…

SearchOS keeps search state outside the model and agents stop looping

Web search agents all have the same old problem: the longer the task, the faster they lose track of what they have already found, what they still haven't, and where they were heading in the first place. The opening stretch looks impressive — the model searches, opens pages, writes out facts, assembles an answer. Then the familiar part begins. The…

Mapping code by behavior helps agents plan edits better on fewer tokens

Conversations about AI agents give almost all their attention to models. Which LLM is stronger, whose reasoning is better, who writes more accurate code. But in real engineering work, an agent's success rarely rests on the model alone. There is another layer that assembles prompts, holds state, calls tools, and keeps the steps in order. The…

A harness trained on past runs lifts Terminal-Bench from 0.722 to 0.806

In conversations about AI agents, almost all the attention goes to models. Which LLM is stronger, whose code is better, who sits higher on the leaderboard. In practice, what decides the outcome is often not only the model but how exactly it is packaged into an agent : what context it gets, which tools it can call, how its steps are structured…

Coding agents fix more bugs when they ask the repo two questions first

Coding agents have an old and very human problem: they start fixing too early . They see a bug report, latch onto familiar words, race through the repository and make a quick edit — in the wrong place, about the wrong thing. Tests fail, steps pile up, tokens burn, and the agent goes in circles. The authors of Know Before Fix propose a simple but…

LightMem-Ego recalls your day well, your lost keys poorly

Today's AI assistants are good at answering questions about the here and now. But ask something like "where did I leave my badge this morning?", "what did my colleague ask me for after the meeting?" or "what do I usually do once I get to the office?" and the magic runs out fast. Answering those takes more than understanding text. It takes memory…

LLMs often sense their own uncertainty but fail to act on it

Large language models have a strange superpower. They can explain quantum mechanics with confidence, write code, argue about philosophy — and then, just as calmly, produce complete nonsense. So the central question today is no longer only what a model knows. It is this: does it understand the limits of its own knowledge ? That is the subject of a…

Today's AI agents keep their autonomy outside the model, not inside it

The word "agent" gets stuck onto almost anything in AI these days. A code-editor extension is an agent. A wrapper around an LLM with tool calls is an agent. A bot that clicks buttons in a browser is an agent too. But the authors of a long conceptual paper, Critique of Agent Model , put an uncomfortable and useful question on the table: what if…

Rewiring the agent harness cut cost 41% with no drop in quality

There's a reflex in the AI industry: when an agent underperforms, it gets more tokens. A longer prompt. More steps. More tools. More replaying of the history. More thinking. On paper this often looks like progress. On the infrastructure bill it looks like a disaster. A new paper with a fitting title, The Harness Effect , goes straight at that…

Reading the full score distribution makes an LLM a better verifier

Large language models have an odd weakness. They keep getting better at generating solutions, but they are still not very good at telling which solution is actually right . That is not a small gap. If your AI agent writes code, works in a terminal, drives a robot arm or handles medical data, the question that matters is not "can it produce a…

Teaching a 32B agent to keep better notes beat a bigger model

AI agents have an old and very mundane problem: they forget fast. Not in the sense that the context window runs out — everyone knows that one. The worse version is this: give a model external memory and it often keeps that memory like a bad set of student notes. Writes down the wrong things. Searches in the wrong place. Duplicates the obvious…

Claude Haiku 4.5 drops below 90% tool-choice accuracy at 10 to 15 tools

Over the past year MCP has become one of the most talked-about ideas in AI infrastructure. The promise is a clean one: a single standard protocol through which an LLM can reach external services, read data, call functions, drive applications. No rebuilding the plumbing for every model. Stand up a server once, and Claude, GPT, Gemini and other…

Readers can't spot AI literary translation but still prefer the human one

One question has hung over machine translation of literature for years: if a model can carry the meaning across more or less accurately, does that mean it can translate a novel as a novel — with voice, rhythm, atmosphere, and that feeling when a text carries you along? A new paper that puts its conclusion right in the title answers honestly: AI…

Agents that can see the tests score 222/222 and skip the library

The AI industry has a favorite trick: post a pretty benchmark score and declare victory. Coding tasks especially. The agent wrote the code, the tests are green, so everything must be fine. But what if that is an illusion? What if the agent passed the exam without building the thing it was asked to build? That is exactly the subject of a paper…

Orca learns one latent world state and reads out text, images and robot actions

For the past couple of years AI has been learning a grab bag of tricks. Some models continue text well. Others fill in an image. Others try to control a robot. But nearly all of them solve narrow problems : predict the next token, the next frame or the next action. The authors of Orca propose a different framing. What if a model should be…

MEG decodes typed sentences without surgery, EEG still far behind

"Mind reading" usually sounds like science fiction. But this particular fiction has a very practical goal: to give people who cannot speak or move a way back into conversation. The best brain-computer interfaces today are already decent at turning neural activity into text or speech. The problem is that almost every genuinely strong system is…

Even the best agents forget, loop and lose the plan over hundreds of steps

Most of the talk about LLMs is about writing code, solving problems and holding a conversation. But nearly every test of those abilities has the same shape: the model is handed a task, it answers, and that is the end of it. A real AI agent does not work that way . It has to act step by step, remember, learn as it goes, come back to objects and…

Verifying an AI agent's code is now harder than writing it

There's an old engineering intuition: finding a solution is hard, checking one is easy. For today's coding agents it holds up worse and worse. A model can already produce a plausible patch, page, interface, even a whole repository. What's hard is telling reliably whether the task was actually solved the way the human wanted — and that is the new…

None of 12 agent memory systems wins across every workload

Memory is no longer a minor detail in an AI agent. It is usually the place where it gets decided whether the agent is useful or starts getting confused. If you follow the progress of AI agents, you have probably noticed something odd. Models have gotten better at writing code, holding a conversation, calling tools, even running long chains of…

Adapting the agent's interface beats retraining the model

In the race for smarter LLM agents we reach almost reflexively for the familiar levers: a bigger model, more fine-tuning, a round of RL, a rewritten system prompt. The authors of Adapting the Interface, Not the Model ask an uncomfortably simple question: what if the agent fails not because it reasons badly, but because it is badly wired into its…

Code is becoming the operating system that agents run on

There is a familiar story around LLMs by now: the model writes code, fixes bugs, calls tools, and sometimes clears benchmarks at the level of a decent intern. AI writes code, fixes bugs, calls tools, sometimes even clears benchmarks at the level of a decent intern. But the survey Code as Agent Harness proposes a far more interesting turn. Its…

PresentAgent-2 turns a one-line prompt into a narrated video talk

Presentation generation has spent years in a fairly dull mode: you have a document, you have a set of talking points, the model turns them into slides. Useful, but predictable. A new paper, PresentAgent-2 , tries to raise the bar considerably. Here the system is handed no article, no report, not even a finished outline — just a short user prompt…

Rewriting the search plan mid-query gives PAI-2 an 18% lift

Large language models have an odd superpower: they can speak so confidently that it sometimes seems they actually know. The trouble is that confidence and knowledge are not the same thing. Large language models have an odd superpower: they can speak so confidently that it sometimes seems they actually know . The trouble is that confidence and…

Google's AI math co-author gets further by keeping its dead ends

Most of today's mathematical AI systems are impressive in "here is a problem, here is the answer" mode. Real mathematics does not work that way. It lives in drafts, dead ends, doubts, strange hunches, half-true lemmas and a twenty-year-old paper you stumble on that changes everything. It is exactly this messy, human part of research that Google's…

DeepMind argues delegation, not model quality, limits agent systems

Today's LLM-based agents no longer just answer questions — they run chains of actions: open a tool, call an API, write code, check the result, send an email. The next step suggests itself: if a task is too big for one agent, why not split it up and hand the pieces to other agents — and sometimes to people? On paper it looks clean. In practice an…

Synthetic computers teach agents to work a month at a time

The big problem with today's AI agents is that we test them on tasks, but the work they actually have to do happens in context. The big problem with today's AI agents is that we test them on tasks, but the work they actually have to do happens in context . Not in a vacuum, not in a tidy sandbox holding a couple of files, but in real digital mess…

Agent quality comes from parallel reasoning and merging, not orchestration

A near-cult of engineering complexity has grown up around modern LLM agent systems. Orchestrators, sub-agents, memory, skill libraries, tool calls — it all looks impressive, and it leaves the important question unanswered: what actually produces the gain in quality? The authors of HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harness…

Long horizons alone can collapse RL training for LLM agents

There is a lot of noise around LLM agents right now: we teach models to use tools, browse websites, fix code, work through multi-step tasks. It looks as though the central question is the quality of the model itself, the size of the context window, or the cleverness of the training algorithm. But the authors of "On Training Large Language Models…

Agents that swap hidden states instead of text get 8.3% more accurate

LLM-based multi-agent systems have an old and rather mundane problem: they talk too much. One agent writes a plan, a second critiques it, a third solves the task, a fourth calls a tool — and the whole collaboration bogs down in endless text generation, latency and token spend. On paper it looks like collective intelligence. In practice it looks…

Organizing agents like a company lifts PRDBench success to 84.67%

In LLM land we are used to measuring progress one hero at a time: who writes better code, who handles websites more carefully, who calls tools more reliably. But as soon as a task gets long, layered and genuinely work-shaped — with dependencies, checks, rework and different roles — the magic of a single agent runs out fast. What is needed is not…

Many systems called world models stop at one-step prediction

Generative AI comes with a convenient illusion of competence: the model writes, draws, sometimes even "plans," and it looks as though there is something like a picture of the world inside it. But the moment a system is asked not to continue text but to act toward a goal over a long horizon — drive a robot, navigate websites, negotiate with…

The best AI agent solves just 54.5% of game development tasks

We are used to measuring agent progress on tasks like fixing bugs in GitHub repositories, writing Python scripts, or building a front end from a mockup. Real development — game development especially — is far messier and far more interesting than that. Generating a function is not enough here: you have to understand scenes, object hierarchies…

What coding agents transfer across domains is discipline, not code

Coding agents share one weakness: they write code well, but they keep repeating the same mistakes, like an intern who rediscovers every time that running the tests before committing is a good idea. Over the past year researchers have been busy teaching these systems to use memory — to store the moves that worked and the ones that didn't, then…

Agents need less memory when the environment keeps the traces

In AI we are used to thinking of memory as something that sits inside the agent: in an RNN's hidden state, in the network's weights, in a replay buffer, in the KV cache. But what if part of that memory can literally be moved outside — into the environment itself? Not in a metaphorical sense but a formal one: so that the agent genuinely needs less…

Bidirectional memory: how agents evolve by remembering past steps

Today's "deep" AI agents do more than continue text: they run searches, call tools, gather facts from different sources and work through hard questions step by step. Today's "deep" AI agents can search, call tools, gather facts from different sources and work through hard questions step by step. In practice, though, an agent like this behaves…

Training agents on five atomic skills lifts coding scores by 18.7%

The authors propose that we stop training AI agents on composite tasks alone and teach them atomic skills instead — small, checkable, reusable building blocks of software development. Coding agents have a persistent problem: models fit benchmarks well enough, but shift the framing of a task slightly and nothing works. An agent closes…

More dynamism in an agent's workflow graph does not always pay off

How the various approaches optimize LLM agent workflows — from fixed templates to dynamic graphs that are assembled and rewritten on the fly. The paper From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents takes on an important question about how LLM-based systems are built today. A system is no longer…

Taking notes, not reasoning, separates the agents that can run a startup

Agents do well on short tasks. Over a long horizon they are undone by memory, inconsistency and an inability to stick to a strategy. AI agents have gotten decent at problems that take a dozen actions, a couple of tool calls and an answer. Stretch the task to hundreds of steps — the length of real work — and it gets interesting. Early mistakes…

Rewriting the harness code beats a hand-tuned baseline by 7.7 points

It isn't only the brain that matters, but the wiring around it. How Meta-Harness automates the job of building an LLM harness. When we talk about AI we almost always talk about the models themselves: size, data quality, architecture, speed. But in real applications there is another layer that decides how useful the whole thing ends up being — the…

The next intelligence explosion is social, not a single superintelligence

The next intelligence explosion will be the growth of a complex social system — a mass of AI agents, humans and hybrid centaurs that together form a new layer of collective thought. For a long time, talk of a coming AI singularity has sounded like a myth about the arrival of one superintelligence: it gets smarter than a human, then smarter still…

Agentic RL trains long-horizon behavior, not single answers

How reinforcement learning (RL) is used not just to produce a "good answer", but to produce behavior that holds up in dynamic conditions. Until recently, reinforcement learning for LLMs looked like this: the model is shown a prompt, it produces one answer, and that answer gets scored — by people or by automatic metrics. This works well for tuning…

Why AI agents break on real APIs, and how Auton tames them

Auton Agentic AI Framework: how to move agents off stochastic generation and onto verifiable contracts and specifications. Agentic AI is what you get when a system doesn't just talk but acts: it calls APIs, queries databases, files tickets in a tracker, posts to Slack, kicks off pipelines. And that is where it turns out that natural language is a…

Only 5% of mature open-source repos have a file written for AI agents

From READMEs written for people to documentation written for machines. A look at how developers write instructions for AI agents in open-source repositories. Since GitHub Copilot and ChatGPT, plenty of teams have gotten used to handing an LLM chunks of code and tests to write. The next wave is agentic tooling that acts far more autonomously: it…

Predicting the answer's latent image beats text-only chain of thought

Multimodal LLMs have learned to recognize objects, but how do you give them visual imagination? A look at the Cognitive Supersensing idea. Over the past few years, multimodal LLMs (MLLMs) have learned to recognize objects, read captions, answer questions about an image and even give a decent account of what is happening in a frame. But they have…

Generating 4D scenes as simulator code drops physics failures to 10%

Generative models have learned to produce striking clips, but that kind of video has a weak spot: nothing forces it to obey physics. An object can drift in mid-air, particles can ignore gravity, rigid bodies can pass through each other. For spatial intelligence that is not enough. What you need is a world model that doesn't just look plausible but behaves plausibly, because a simulation is…

Predicting the UI change in words first makes Office agents pick better actions

We tend to assume that work inside office applications is predictable: the interface is deterministic, the buttons are where they were, everything behaves as usual. For an AI agent running long chains of actions in Word, Excel or PowerPoint, reality is harsher. One wrong click in the UI and you can lose an important artifact, corrupt a document, or land in a state that is hard to back out of.…

Auto-generated AGENTS.md files make coding agents worse and costlier

The idea looks obvious: give a coding agent a dedicated file with the repository's rules — how to build the project, how to run the tests, what the conventions are for style and structure — and it will work more confidently and make fewer mistakes. Such files are usually called context files: AGENTS.md, CLAUDE.md and the like. There are already tens of thousands of them in open source, and…

Moltbook's millions of AI agents talk constantly and never socialize

When millions of AI agents talk to each other, does that add up to a society? LLM agents now live in networked environments where they write posts, argue in comment threads and vote on each other. The intuition is easy: give such agents enough time and dense enough contact and they will start behaving like people in a community — picking up each other's style, recognizing authorities,…

Smarter reasoning models make collective outcomes worse in social dilemmas

As autonomous LLM agents take over human tasks — from negotiating with services to allocating resources inside companies — we have gotten used to judging them on solo benchmarks. We care how well a model writes code, answers questions, or plans. But in the real world they run into each other, compete for limited resources, and sometimes manufacture competition nobody needed. The paper…

Chats that undermine user autonomy get more thumbs-up than average

AI assistants no longer just answer questions. We consult them about work, relationships and health, and we ask them to help phrase a difficult message or make a decision. Most of the time that is convenient. But there is another side: sometimes the help is shaped so that a person gradually hands the AI what they would normally keep in their own hands — their grip on reality, their moral…

Four agent roles and a real PR review get 72.4% on SWE-bench

LLMs can suggest a chunk of code, explain an error, or write a test. But the moment a task starts to look like real development work — read an issue, find your way around a project, reproduce a bug, produce a patch without breaking everything else — a single general-purpose agent often falls short. In Agyn: A Multi-Agent System for Team-Based Autonomous Software Engineering the authors put…

Injecting world knowledge into tasks does not make a world model

Over the past couple of years it has become fashionable to talk about world models: systems that do not merely continue text or fill in the next frames, but understand at least a little of how reality is put together and how it changes over time. The authors of Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks…

Building a sub-agent for each step beats fixed multi-agent roles

LLM agents break down when a task stretches across dozens of steps — checks, backtracking, experiments, running commands, fixing what failed. The context swells, noise piles up inside it, important details get buried, and the agent spends its time not on the work but on trying to recall what happened earlier. Multi-agent systems try to fix this by coordinating several roles, but that brings…

LingBot-World: an open-source world model you can steer in real time

Not long ago, models learned to generate video from text with a few seconds of coherent motion. But ask one of those systems to walk forward, look back and return to a familiar object, and the magic stops. Objects trade places, details drift, and causal logic gives way to statistical coincidence. That gap between a nice-looking picture and an actual simulation of a world is what the authors…

LLMs cut the hand-written rules out of data cleaning, but not the cost

Every company has the same problem: the data sitting in its tables and databases is arranged in a way that makes it hard to use. Formats vary, values contradict each other, some fields are empty, and different sources call the same thing by different names. Analysts end up spending weeks preparing datasets, and the business loses money to errors and delays. The authors of the survey Can LLMs…

DPO is the steadiest negotiator when LLMs are trained on the final outcome

LLMs can hold a conversation, but move them into a multi-agent setting where they have to strike deals, apply pressure, concede, deceive and hold on to a goal across the whole exchange, and the trouble starts. Much of the problem is that most fine-tuning methods score responses locally: is this particular span of text good, is it polite, is it coherent. In a negotiation what matters is not the…

Amazon's seller analytics agent skips SQL and answers in under 15 seconds

An e-commerce seller makes decisions on the fly all day: what to push in ads, where sales slipped, which products drag the business down and which drive growth. There is no shortage of data, but the value in it is rarely obvious. Getting an answer means opening several tools, knowing which report holds the number, how to build it, which filters to set — and then reading the result correctly.…

RoboBrain 2.5 gives robots metric depth and a running sense of progress

Robotics has an old sore spot: even strong vision-language models reason about a scene well enough, yet sometimes fail at acting in the physical world. In household terms it sounds simple — "move the mug 10 centimeters to the right," or "pour the water over the flowers from a height of 1–5 cm." For a person those instructions are close to trivial; for a robot they are a minefield. It has to…

Reasoning models get better by arguing with themselves inside one trace

We tend to assume reasoning models are stronger simply because they write longer chains of thought and burn more compute before answering. In Reasoning Models Generate Societies of Thought the authors offer a more interesting explanation: these models don't just think for longer, they start thinking differently — as though a small assembly of voices had formed inside them, putting questions to…

42,267 commits show multi-agent frameworks are still building, not stabilizing

A new tooling layer has grown up around LLM applications: frameworks for assembling not one clever chatbot but a whole team of specialized agents. One plans, one retrieves data, one writes code, one checks the result. In demos this looks like a shortcut to complicated products. Every such trick has a back side, though: maintenance, bugs, breakage against external APIs, and a permanent chase…

Agentic RAG beats modular rewriting but loses on routing and reranking

RAG is one of the most practical ways to connect an LLM to outside knowledge: instead of leaning on what it memorized, the model first pulls the relevant passages out of a knowledge base and only then answers. In shipped products it reads as the cure for hallucinations and stale facts. But plain RAG sometimes retrieves the wrong thing. So the industry took two roads.

Structured GitHub fix histories lift SWE-Agent by 4.65% on SWE-bench

Once LLMs got good at writing code, a layer of autonomous SWE agents grew up around them: systems that can open a repository, run the tests, localize a bug and prepare a patch. But these agents have an unfortunate habit of working as if they had never seen a bug like this before. Human developers almost never fix hard problems from scratch: they go to GitHub, look for similar issues and PRs,…

Absolute Zero trains reasoning with zero data by inventing its own tasks

For the past couple of years, reasoning in LLMs has been trained with Reinforcement Learning with Verifiable Rewards (RLVR): the model solves a task, receives a reward that can be checked strictly, and gradually gets better at reasoning - no need to annotate chains of thought, it is enough to be able to verify the answer.

A single jump tool beats a full search toolkit at issue localization

If you have ever opened a large repository hunting for a bug, you know the feeling: hundreds of files, a tangle of non-obvious connections, and an issue description that explains almost nothing. An LLM faces the same problem, only harder — it physically cannot hold the whole project in context. So today's software engineering agents are forced to work iteratively: read a fragment of the docs…

What LLMs lack for AGI is a coordination layer, not understanding

There is a lot of argument around AGI right now: LLMs supposedly just predict the next word with some probability, so you cannot build "general AI" on them. The authors of The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics from Stanford suggest looking at the problem differently. Their argument is that LLMs really do not hand you general intelligence out of the box — but…

Professional developers don't vibe with agents — they control them

A couple of years ago, the only job an LLM held in programming was autocomplete: the model suggested a line, finished a function, reminded you of the syntax. By 2025 the focus had moved. Agentic tools arrived that don't advise but act — they read a project, edit files and run the tests.

Sophia gives agents autobiographical memory and cuts reasoning steps by 80%

Today's AI agents can plan, call tools, run chains of actions, even operate inside a multi-agent system. But most of these setups share an awkward property: they are fundamentally reactive. An agent can answer well in the moment, yet after deployment it rarely changes its own habits, rarely revisits its strategies, and almost never modifies itself. When the environment shifts — new interfaces,…

Coding agents score 21% when asked to evolve a codebase between releases

Coding agents have gotten noticeably better over the past year: they can find where something broke, edit the code and run the tests. There is an important caveat, though. Most popular benchmarks measure point achievements — fixing one specific bug, or adding a small feature scoped to a single issue. Real development does not work that way. Code lives for years, requirements shift,…

A protocol that lets AI scientists share instruments across labs

Autonomous AI scientists can already read papers, propose hypotheses, run computations and even drive experiments. But in real science their abilities tend to stay locked inside one lab — a particular pile of scripts and informal agreements about where the data sits and how the instruments get used. The moment a team tries to move that stack to another lab, or to repeat someone else's…

No model beats 50% on shopping in the new ACE consumer benchmark

While AI handles logic problems and writes code with confidence, in real life people increasingly ask it about something far more mundane: what to buy at the store, what to substitute for an ingredient, how to fix a leaking faucet, which build to pick in a game. And here an inconvenient fact surfaces: everyday requests are simple only in the telling. They hang on context, on current data from…

Reasoning models now pass all three CFA levels, but ethics still trips them up

In finance, the CFA (Chartered Financial Analyst) exams are a marathon run over three distances. Level I tests the fundamentals and whether you can keep the terminology straight. Level II puts you in front of cases where formulas and logic have to be applied in context. Level III asks for more than correct arithmetic: a coherent professional answer — how to build a portfolio, how to assess…

Top LLMs score 30 out of 100 on a benchmark of the full research cycle

Today's LLMs can do a lot: explain hard topics, write code, hold a long thread of reasoning. But science is not just knowing the answers. It is a research cycle: work through the literature, come up with a hypothesis, test it with an experiment, then interpret the results honestly and adjust the plan. And this is where the field has long lacked a shared language and a shared yardstick: what…

Storing verified lemmas instead of context gets an agent to olympiad gold

Over the past couple of years, large reasoning models (LRMs) have gotten noticeably better at olympiad math. On problems at the level of AIME (the American Invitational Mathematics Examination), one long reasoning pass is usually enough: the model writes out a chain of thought, checks the answer, sometimes takes a couple of runs at it — and lands it. IMO (the International Mathematical…

Only 14.5% of agent context files say anything about security

Today, when we write code with AI, we state the task in natural language and the agent inside the IDE plans the steps itself, writes the changes, runs the tests and tries to carry the job through to a result. This is called agentic coding. But it has a weak spot: the agent has to work out fast how the project is put together, what it is conventional to use in it, how to run the build and the…

Agents with memory and a forum beat centralized search on science benchmarks

Most of today's approaches to "AI for science" look like a familiar pipeline. There is a central controlling algorithm, a metric, and a short loop: generate an improvement, run the test, keep the best one, repeat. Broadly, it works — but it also strips out what makes science science: long memory of past attempts, the exchange of ideas, argument, and the sudden transfer of methods between fields.

DataFlow rebuilds LLM data prep as PyTorch-style operators

The hard part of training language models right now is not new architectures — it is data quality. You can rarely just collect data, clean it and train on it; you have to build processes where data can be synthesized, validated, improved, discarded when it is bad, and rolled back a step so the whole thing can be rebuilt again. In practice, though, plenty of teams run on a pipeline assembled…

LAMP turns economic news into a signal RL agents can act on

Economics textbooks are tidy: prices, taxes, rates, utility. In real life, the decisions of people and governments are constantly nudged by words — news, conversations, expectations, rumors, public statements. The same set of numbers reads differently depending on whether the talk around it is "a crisis is coming" or "everything is under control". That layer of reality stayed awkward for…

Making tests fight the patch lifts SWE-bench Verified to 79.4%

LLM-driven bug fixing stopped being exotic a while ago: a model can read code, propose edits, and run the test suite itself. In real repositories, though, the whole thing runs into an awkward detail. There is often no good way to check whether the bug is actually gone. When tests are missing, weak, or simply don't cover the case in question, the system can produce a patch that turns the suite…

A multi-agent pentester outscored 9 of 10 humans on a live network

The argument over how good AI agents are at cybersecurity work has been running hot for a while now. It usually rests on one task: finding known vulnerabilities. But real penetration testing doesn't look like that. It's a large corporate network, thousands of hosts, incomplete (and more often unreliable) data, chains of small vulnerabilities and the human factor. Which is precisely where it…

DeepCode rebuilds a paper's codebase by managing memory, not context

Over the past year, LLM coding agents really have learned something new: they now handle tests, run commands and get through relatively long sessions. But raise the difficulty — ask an agent to "ship the repository for this paper" — and reality arrives fast. A paper carries a lot of load-bearing detail, scattered across its sections, while the finished project is dozens of files, dependencies,…

Agent teams stop paying off once one agent clears 45% success

The idea looks obvious at first glance, and yet multi-agent systems still aren't the default in most applications. Put concretely: if one LLM-based agent can do a task, several agents ought to do it better. You can split the work, or convene a "council" so they check each other. In practice, teams of agents often run slower, cost more and behave dumber than one. A paper by researchers at…

Matrix drops the central orchestrator and scales agents 15x

Synthetic data generation today runs on several agents at once — one writes text, another scores it, another calls tools, another picks the best candidate. High-quality data requires agents that can interact with each other and with their environment, producing long, branching scenarios with multiple dialogues that return only the best variant. All of this makes it hard to scale an agent to…

Two HTML tags cut a web agent's task time from 27 seconds to 2

Web agents today behave inside other people's interfaces like uninvited guests: they stare at screenshots and guess which buttons can be clicked. The smallest interface update breaks the whole logic, drives up the cost of maintaining pipelines, and user privacy suffers along the way. The authors of VOIX propose a simple but far-reaching answer: let sites themselves hand agents the actions they…

Coding agents refactor for readability, not architecture

When AI agents write code, they take on more and more of what used to be purely human work — planning, running tests, even step-by-step refactoring. The authors of Agentic Refactoring: An Empirical Study of AI Coding Agents are the first to look at that practice broadly and in depth: how agents refactor in real open-source projects, whether it pays off, and how their style differs from a human's.

Lumine plays Genshin Impact for hours and transfers to other games

General-purpose agents are back at the center of the conversation. The team behind Lumine offers a concrete recipe for an agent that holds up for hours on hard tasks — 3D navigation, puzzles and dialogue — inside the open world of Genshin Impact, and then carries over to other games with no additional training.

GroundCUA matches desktop grounding baselines with 700K examples, not 9M

Agents that operate a computer keep failing at a step that looks trivial: finding the element on screen that a human instruction describes. That grounding is hardest on interfaces crowded with tiny controls, near-identical panels, high resolution, visual noise and rendering artifacts. The GroundCUA team shows how to solve this narrow but load-bearing problem — making the link between language…

Sora-2 solves visual puzzles by drawing its reasoning in video

When we ask a model to reason, it reasons in words if the medium is text, or over a static scene if the medium is an image. The world, though, is not static: objects move, and the rules often only become visible in how those objects behave over time. The authors propose video generation as a general-purpose channel for reasoning. Text can be written directly into the frames, visual hypotheses…

Jr. AI Scientist writes junior-level ML papers, and reviewers rejected all three

Autonomous agents have lately been pitched as systems that can come up with ideas, write the code to test them, run the experiments on their own and write the paper. In practice they have tended to fall short: the ideas were never properly checked for novelty, the experiments stayed at prototype level, and the results came out weak (see AI SCIENTIST V1 / V2). Researchers at the University of…

Rewriting an image as SVG code costs GPT-5 15 points of accuracy

Today's vision-language models see an image as an array of pixels. But to actually understand a picture, they need to work with symbols rather than pixels — the way they work with code. Pixels are fine for recognition and poor for carrying an image into an LLM's context. And pixels do not reliably tell you how objects relate to one another, or how many of each thing is in the frame.

ChatGPT Atlas solves Sudoku in 2:28 and scores zero in Flappy Bird

What happens if you give an agent eyes and hands inside the browser — not just the context of the page but an intent, plus the ability to click and press keys on purpose? Researchers decided to find out by running one through several browser games. You can probably guess the answer already: Atlas has real strengths in turn-based logic, and real-time control is its Achilles heel.

An agent trained in a simulated clinic beats GPT-4o at ordering the right tests

In medicine, a clinical diagnosis usually takes several moves from the physician: form a plausible hypothesis from the patient's symptoms, order the tests that confirm or rule it out, and decide when to stop testing and commit to an answer. Most large language models (LLMs) do well at diagnosis on fixed cases, but they fall short on planning — on choosing and prioritizing the tests that matter…

JanusCoder learns to see the interface its own code renders

Scientific plots, interactive interfaces, animated walkthroughs of theorems — all of it is, at bottom, code rendered as something you look at. Yet today's AI systems work in the text modality alone. They have no notion of what the code will look like on screen, or how it will behave inside a running application.

Planning subtasks as a graph cuts an agent's steps and latency

Autonomous agents run on tool calls. But almost every popular agent framework issues those calls in strict sequence: the agent thinks, calls one tool per step, waits for the result, then decides what to do next. That is convenient for keeping the agent under control and understanding what it is doing, but it hits a ceiling on tasks that contain independent subtasks which could be solved at the…

Teaching an agent to fold its own memory cuts context 92% by step 100

Tasks that require tool use and repeated web search tend to produce long trajectories that break most LLM agents: either the agent piles up the entire history, which is punishing to carry in context, or it compresses the entire history at every step, which means forgetting details that mattered.

Salesforce's EDR shows its research plan and lets you edit it mid-run

Enterprise data tends to sprawl across email, reports, databases and code repositories. Answering a hard question usually takes not one fact but many, plus the ability to synthesize hundreds of sources with checkable citations and a line of reasoning someone can follow. Ordinary agents and classic RAG systems are often a poor fit here: they give shallow answers, they are hard to steer once…

DeepAgent replaces the agent pipeline with one long reasoning stream

LLM agents can reason, but reasoning is not enough to solve real tasks. An agent also has to call outside tools, work through long scenarios and stay autonomous across dozens of steps. Rigid pipelines with fixed modes get in the way of that, and so do the classic approaches like ReAct and Plan-and-Solve. They impose the same action loop every time and work well on tasks that take two or three…

Agents that share latent thoughts instead of words reach 93% on MATH

Put several models in a multi-agent system on the same question and they will argue a little, correct each other, and settle on a compromise that is not always right. Language is what lets them do it, and it is also the bottleneck: it is sequential, often ambiguous, and rarely a faithful record of the reasoning behind it. A new paper argues for working above the level of words and giving…

FinSight writes financial reports where every claim carries a source

A financial report is more than text: its force comes from numbers you can check and charts that point back to their sources. LLMs write well but hallucinate. The FinSight team attacked that with a separate group of agents responsible for gathering data and verifying calculations, plus specialized vision-language models that make the final pass over how charts and tables are rendered. The…

Agents with 18,000 MCP tools clear just 1 of 7 hard Azure tasks

Researchers at Microsoft have built a new benchmark for agents that solve tasks not through a browser but by calling tools directly over MCP. They assembled more than 18,000 tools from Azure, GitLab, RocketChat, Plane and ownCloud, paired every task with the correct set of tools needed to finish it, and tested six models. The result cuts both ways: hand an agent the right tools up front and it…

LLM-simulated interfaces train web agents better than real sites do

AI agents are data-hungry: they need thousands of varied scenarios for working with websites and mobile apps. Assembling a set like that by hand is slow and expensive. Even a few hundred tasks with long action chains means thousands of hours of engineering, annotation and infrastructure. The authors of UI-Simulator propose an alternative: instead of collecting everything in real environments,…

ColorAgent hits 77.2% on AndroidWorld by splitting the work across agents

We are used to systems where you press a button and the machine silently carries out the command. Mobile changes that. In anything complicated — ordering food, configuring an app — you need an intermediary that can understand what the user actually wants, hold the context of the conversation, ask clarifying questions and act consistently even when the interface shifts underneath it. The team…

DeepAnalyze-8B runs the whole data science pipeline with no orchestrator

Autonomous data science is an old dream: go from raw tables and files to clean charts and a coherent analytical report without a human steering every step. Large language models (LLMs) moved this forward, but the usual workflow agents run on rules written out in advance. They are brittle: the moment a task steps outside the script, the whole process falls apart. In a new paper the authors…

Alpha-Service turns AI glasses into an assistant that speaks first

Unlike today's voice assistants, the AI for Service team proposes a more interactive approach. They argue that an AI should recognize on its own when a person needs help and offer it without being asked. That approach, which they call proactive assistance, was demonstrated on AI glasses streaming first-person video.

A $535 pipeline labels LLM hallucinations in 14 languages

Even the strongest LLMs will sometimes state, with full confidence, facts that appear nowhere in their sources. In question answering, one wrong word is enough to break the meaning. Most checks available today return a single verdict for the whole answer, and almost all of them work only in English. The authors of PsiloQA set out to do the opposite: cover 14 languages, and learn to find not…

Open tooling and human-in-the-loop RL train real robots in one to two hours

For decades robotics ran on one recipe: build a map of the world, solve inverse kinematics, tune the controllers, then do it all again when the task or the robot changed. That works in sterile conditions and falls apart in the real one, with noisy sensors, contact and soft materials. Researchers at Oxford argue for a different route: when the world is too tangled to describe object by object,…

Linear self-attention provably cannot beat linear regression on AR processes

After AI's run of success in language, images and video, plenty of people expected transformers to carry time series forecasting too. The reality is usually duller: plain linear regression often beats the heavyweight models on mean squared error. The paper under discussion gives a careful, rigorous account of why that happens, viewed through the lens of in-context learning.

LLM-generated end-to-end tests needed edits to 10% of their lines

Building end-to-end tests is always a trade-off between speed and reliability. The scripts have to walk the whole user journey: UI, business logic, integrations. Writing them by hand takes weeks and demands expertise in frameworks, selectors and stable locators. Large language models can already generate unit tests, but integration scenarios are a harder problem. The authors of GenIA-E2ETest…

An LLM agent that tests its own equations beats symbolic regression baselines

Scientific data often hides simple laws — equations that explain how one quantity depends on another. Finding them is hard: the space of formulas is enormous, measurements are noisy, and brute-force search chokes almost immediately. Symbolic regression is the attempt to recover exactly that kind of compact formula. Most approaches either enumerate expression trees or train a neural network to…

A 7B agent that drives a real browser beats Search-R1 by 20%

Most of today's web agents solve tasks through a long pipeline: scrape the page, compress it into text, hand it to an LLM. That is convenient, but poor in actions: no real scrolling, no clicks, no work with tabs or forms. Costs climb too, because of all the external calls. The BrowserAgent team proposes going back to the original source — acting directly in the browser, the way a person does.…

BigCodeArena scores coding models by running their code, not reading it

Judging code generation by how clean the comments look is like judging a car from the brochure. In real life what matters is whether it starts, whether it brakes in time, and whether it is pleasant to use. The authors of BigCodeArena take exactly that practical view: their open platform compares large language model (LLM) solutions not by string overlap but by whether they run, how they…

Natural-language memory beats fine-tuning on long agent tasks

Large language models do well on short reasoning and coding benchmarks. But real work stretches over dozens or hundreds of steps, demands switching between applications, careful context management and the ability to catch your own mistakes. The core problem is that at test time most agents stay static: they accumulate no experience and get no better from one attempt to the next. The authors of…

Rewriting an agent's context collapses it; small edits gain 17 points

Over the past two years one thing has become clear: many applications built on large language models learn better through careful work with the context than through fine-tuning weights. Into the context go system instructions, reasoning steps, examples, domain rules, facts, even hints on how to use tools. It is transparent, portable, and it works at runtime. On top of that, progress on long…

MLE-Smith auto-generates 606 ML tasks that rank agents like human benchmarks

When the subject is measuring what AI can do in machine learning engineering, the default answer is a static benchmark: a competition assembled once by its organizers, a single dataset, a fixed metric. That is convenient, but it scales badly. Every task has to be verified at length, forced into a common format and updated by hand. What you end up with is a narrow world of tasks, while the real…

Cursor CLI completed 70% of MITRE attack techniques when asked

LLM-based computer-use agents no longer just answer questions — they click through files, run shell commands, move data around and connect over SSH. An assistant like that turns into an attack tool the moment someone asks it to get around a control or do something malicious. The authors want an honest measurement of that risk: can off-the-shelf agents carry out tactics and techniques at the…

Graph2Eval builds agent benchmarks straight out of a knowledge graph

The traditional ways of training AI agents have stopped working. It shows most clearly in agents that have to read documents, parse diagrams, click through sites and carry out multi-step scenarios. Hand annotation goes stale quickly and costs a lot. Generating tasks automatically with LLMs is already being tried, but it usually collapses into plain question-answer formats that teach nothing…

An inverse dynamics model turns YouTube videos into agent training data

AI agents promise to be useful inside real applications: setting up a browser, editing images, driving a media player. But to hit the right button and not get lost in a menu, they need thousands of good demonstrations recorded inside the target software. That data barely exists: the available datasets are narrow and go stale fast, and synthetic traces tend to be too simple and a poor match for…

PaperTalker builds a paper's explainer video and beats humans on informativeness

A short two-to-ten-minute explainer video has become close to mandatory for a paper: it goes on the project page, it gets shown at seminars, it gets forwarded to colleagues. But making one means hours of slide prep, recording a voice track and a talking head, then editing and revisions. And it is not the same problem as "natural" video generation: here you have to carry the paper's long…

CoDA's agents grade their own charts, lifting MatplotBench from 55 to 79.5

Turning a plain-language request into a correct chart is harder than it looks. The data is large and heterogeneous, the code breaks often, and a good chart almost always takes several rounds of edits. Researchers at Google argue the task should not be treated as one-shot code generation by an LLM, but as the coordinated work of a multi-agent system in which every role does its own job and…

IoT-MCP hits 100% tool-call success at 205 ms across six MCU families

We talk a great deal about large language models and the smart home, but the conversation rarely reaches actual hardware. In IoT, every microcontroller, sensor and protocol plays by its own rules. An LLM will answer questions about them readily enough, yet it has no painless way to negotiate with devices that keep dropping off the network or returning data in unexpected formats. The authors of…

Ranking code by perplexity compresses context 5.6× without hurting accuracy

Coding LLMs can already complete, explain and fix code, but real projects make them read thousands of lines. Long context windows help, yet they cost time and money — and they paradoxically hurt accuracy: the model starts drowning in detail and misses hidden dependencies between files and functions. Generic text compressors throw away phrases and tokens with no regard for code structure.…

TSci's four agents cut forecast error 10.4% over statistical baselines

Real companies deal in tens of thousands of short, noisy time series, full of gaps, with horizons and sampling frequencies that keep changing. The hard part is not the model — it is everything around it: cleaning the data, validating it properly, building ensembles, producing reports that survive an audit. Narrow domain-specific solutions transfer badly between domains, and general-purpose…

Training on deliberately vague questions makes a 14B agent search deeper

We taught models to hold a conversation and solve equations a long time ago, but out in the real world they stumble on search and fact-checking. A single query is rarely enough: you have to follow leads, refine them, cross-check. The InfoAgent team built exactly that kind of web detective — an LLM agent that searches long and deliberately, reads pages, backtracks and keeps going. The core idea…

Pointing at a pixel beats text commands for drone navigation

Navigating from written instructions has been a hard problem for autonomous drones for years. Classic reinforcement learning approaches need large datasets and transfer badly to new domains. The recent wave of vision-language-model solutions promised generality, but usually asked the model to emit its commands as text: turn, fly, ascend. Language turned out to be a clumsy carrier for precise…

Simulating 30,000 APIs as databases lets a 4B agent match a 30B one

Most useful agents are missing one thing: robust, accurate function calling. Not a well-phrased answer — the right tool calls, with the right arguments, in the right order. The trouble is that data with scenarios like that barely exists, and hand-written scenarios are brittle and scale badly. The authors of AgentScaler suggest looking wider: expand the agent's world and both task diversity and…

Model reasoning traces break into the same episodes human solvers use

Large reasoning models (LRMs) today don't just answer — they unfold long chains of thought. That lets them handle harder problems, but it creates a new one: how do you describe the structure of that reasoning, and how close does it come to human thinking? The researchers propose borrowing a framework well tested in cognitive science — Schoenfeld's episode theory, built originally to analyze…

Federation of Agents beats the best single agent 13× on HealthBench Hard

Today's multi-agent systems often look like a stage play with the roles handed out in advance: every agent gets its own domain, its own channel, its own script. That is convenient for prototypes and falls apart on real work. Who can actually do what? Under which rules? How do you find the right executor among hundreds of nodes, and do it over constrained networks such as IoT? The authors of…

LLMs score 70+ on game code but under 25 on how the game looks

Making a game is more than getting code to run. It takes mechanics a player can grasp, art that looks decent, smooth animation and a steady 60 FPS. Large language models handle algorithmic problems confidently, but evaluations of their code rarely account for playability or aesthetics. The authors of V-GameGym set out to fill that gap: they assembled a realistic benchmark for visual game…

Claude Code pull requests get merged 83.8% of the time versus 91% for humans

Over the past few months developers have been writing code with agents en masse — autonomous LLM-based assistants that plan their own steps, make the changes, run the tests and open a pull request on their own. In theory that saves hours of routine work. In practice there is still little data on how such PRs fare inside real projects: which tasks agents take on, how often their PRs are…

Top coding agents solve under a quarter of SWE-Bench Pro tasks

Over the past couple of years, agents built on large language models have settled into everyday development: they read repositories, fix bugs, propose patches and run tests. On the classic SWE-Bench Verified, the top systems clear more than 70% of tasks on the first attempt. The trouble is that progress like this creates a false sense of readiness for real industry work. Out there, tasks…

78 examples beat 10,000 at teaching an AI agent to act

The industry has spent years waiting for AI that does more than produce a polished answer: plan a task, pick the right tools, fix its own mistakes, and carry the job through to a result. The authors of LIMI (Less Is More for Intelligent Agency) make a bold claim — cultivating agency does not require drowning in millions of examples. What matters far more is assembling a few dozen…

Planning a repository as a graph beats Claude Code by 27 coverage points

Large language models write functions and individual files with confidence, then lose the thread when they have to assemble a whole project. Over a long horizon natural language stops being reliable: vague phrasing, mismatched interfaces, leaking dependencies, structure that falls apart. The agent changes its mind mid-task, tests drift, and the codebase turns into a pile of fragments.

K2-Think gets a 32B model to frontier math scores with test-time compute

The past year has settled one question: to get better at hard problems, an LLM does not have to keep growing in parameters. What matters more is teaching the model to think at length and with structure, and shifting part of the computation to inference time. K2-Think is a sharp example of that shift. The team takes a 32B model — a size anyone can afford to run — and squeezes the most out of it…

WebResearcher condenses each round into a report instead of growing its context

Most open work on deep search runs on one simple principle: pile everything you find into a single large context window. Each step adds new excerpts, links and notes. Eventually the useful material drowns in noise, early mistakes stay in the record forever, and the room left for thinking shrinks fast. The authors of WebResearcher propose the opposite: periodically stop the stream, squeeze what…

AgentScaler turns 30,000 tools into verifiable simulated environments

For the past year everyone has been arguing about how to make agents use tools with confidence: book a ticket, check a delivery status, pull a balance, assemble an answer out of several APIs. The bottleneck is always the same — there aren't enough realistic, varied trajectories in which an agent calls functions in sequence, sees the responses and changes the state of the world. A new paper…

Semi-online RL lifts a 7B GUI agent to 34% on AndroidWorld

Automating the interfaces on a screen is a long-standing wish: open an app, find the right button, run through a series of steps and see the task through. Today that job falls to agents built on large language models, which can look at screenshots, reason and act. But once the scenario runs to many steps, progress tends to run into the question of how we train these systems in the first place.

LongEmotion finds small models hold long support chats better than GPT-4o

Most tests of emotional intelligence in LLMs work on short, neatly labeled utterances. Real conversations are messier: people talk at length, drift, change the subject, circle back to old hurts. Over that distance models start losing the cues that matter, confuse cause and effect, and rarely sustain a coherent line of support. The authors of LongEmotion propose exactly that stress test — a…

The agent economy is forming by default, not by design

Autonomous AI agents are becoming participants in growing digital markets rather than mere assistants: they negotiate, buy data, plan, write code, operate robots. The authors argue this should be read as an emerging agentic economy — a web of markets where agents interact at high speed and often without a human in the loop. Their stance: don't wait for it to grow on its own, design the rules…

BERT's word predictions track the brain's N400 during natural listening

We almost never hear speech as a stream of unexpected sounds. The brain is constantly guessing at the next word and checking itself as the sound arrives. That mode is cheap: the sharper the expectation, the less effort recognition takes. Prediction in vision and hearing is well documented, but semantics — the meaning of words — has long been the…

Giving planning tokens extra credit beats GRPO on math reasoning

Reasoning tasks are a sore spot for many AI systems, even ones with solid factual knowledge. A new paper shows that reinforcement learning (RL) does more than push accuracy up — it rebuilds the model's internal logic into a hierarchy that runs from low-level execution to high-level planning. That explains where those aha moments come from. More usefully, it explains why the standard algorithms…

Even GPT-5 solves fewer than 60% of live multi-tool agent tasks

MCP-based agents can already do a lot: search the web, work with files, draw charts, run calculations, call external APIs. But a demo on a single task is one thing, and sustained work in a realistic, shifting environment is another — one where service responses differ from run to run and several dozen tools are on offer at once. Most existing benchmarks miss this: they are short, synthetic,…

EnvX turns a repository into an agent that sets up and runs itself

Open repositories are full of ready-made work: scripts, models, datasets, demos. Getting any of it to actually run is still manual labor — install the dependencies, download the artifacts, read the docs, get the input arguments right. EnvX proposes something simple but powerful: agentize the repository. Turn it into an autonomous assistant that understands the project's own documents, builds…

Self-written explanations are what let models read dark humor in memes

Not all jokes work the same way. Clean humor runs on wordplay and harmless incongruity; dark humor runs on painful subjects, cultural references and fine contrasts between the image and the caption. In memes this is especially visible: the picture says one thing, the text says another, and the meaning appears where they meet. Until recently there was no good multimodal dataset for dark humor…

Paper2Agent turns a paper's code into an agent you can query

A paper is text, figures, and, somewhere in a repository, code. Then the grind starts: tracking down dependencies, setting up an environment, working out the API and the data formats. For a lot of people that is a high barrier to entry. Paper2Agent proposes something simpler: turn papers into AI agents you can address in natural language and run their methods on the spot. What used to be a…

Top LLMs reason alike but diverge sharply on sycophancy and rephrasing

Today, evaluating a large language model comes down to a single number on a benchmark. That is convenient, and it is not enough: two models post identical scores and behave nothing alike in conversation. A group of researchers proposes looking deeper — taking a model's "behavioral fingerprint" along several axes to see how it actually thinks. The idea is simple: measure a profile of cognitive…

Hallucinations persist because benchmarks reward confident guessing

Why do LLMs keep getting things confidently wrong when saying "I don't know" would serve everyone better? Researchers at OpenAI offer a clear answer: the root of the problem is statistical. It appears during pretraining and is then locked in by the way we evaluate models after fine-tuning. In short: the data can be free of errors and the training objective will still push the model toward…

Universal Deep Research compiles a written strategy into runnable code

When people say “deep research,” they usually mean a service that plans its own search, walks through sources, collects citations and hands back a tidy report. Convenient — and almost always locked to a single strategy and a single model family. The authors of Universal Deep Research (UDR) propose a different arrangement: let the user pick any LLM and write the research strategy themselves,…

The four levels between AI as a calculator and AI as an autonomous scientist

We are used to AI as a clever calculator: it helps with data analysis, but the decisions and the experiments stay with people. The researchers argue for a different view, in which agentic AI moves into the role of an autonomous research partner. It reads the literature, forms hypotheses, plans experiments, runs robots or simulations, analyzes the…

VLWM predicts the future in language instead of pixels

When we ask a machine to help us cook dinner or swap a SIM card, it has to do more than recognize the objects in frame — it has to picture how the world will change from one step to the next. Most systems today see pixels and answer in short phrases, and long-horizon planning still does not work. The VLWM (Vision Language World Model) team proposes a different route: describe the future in…

A 31-subtype error taxonomy beats blind retries in text-to-SQL

Turning a human question into correct SQL is a surprisingly hard problem. Large language models write valid syntax well but miss the logic easily: they confuse tables, join on the wrong key, forget GROUP BY, apply the wrong filters. Plain self-correction from execution results doesn't always help — a query can run fine and even return something plausible that still isn't what the user asked…

BSC-Nav's three-layer memory lifts robot navigation to 78.5% success on HM3D

Most AI agents today are reactive: they see a frame and act, see the next frame and act again, and never build a coherent picture of the space around them. Hence the trouble with long routes, with reusing past experience, with flexibility. Biology solved this elegantly: the brain keeps landmarks, route knowledge and survey maps. BSC-Nav carries that principle over to robots and gives them a…

A million action steps from Chinese apps put UItron ahead of UI-Tars

Could AI agents ever work a computer the way people do — see the screen, understand it, click, launch apps and carry out long chains of tasks? That is no longer science fiction. A new generation of models, UItron among them, promises to reset what automation on desktop and mobile can look like.

Partial deepfake edits slip past both detectors and human viewers

We tend to picture deepfakes as clips that are synthetic end to end. What actually turns up in the wild, more and more often, is the careful partial swap: not the whole video, but a small piece of it — a gesture, a face, an object on the table, a few frames in the middle. Edits that precise do not catch the eye, and they hide perfectly inside genuine footage. The authors of FakeParts argue…

Routing across eight LLMs beats GPT-5-medium by 7% at the same cost

Anyone who has wired a large language model (LLM) into a real product has run into the same choice: more accurate but expensive, or cheaper but worse. GPT-5, the authors note, is already moving toward a fix through test-time routing: easy queries go to a faster, cheaper model, hard ones to the powerful one. The Avengers-Pro team pushes further —…

AgentScope 1.0 makes multi-agent systems work without the duct tape

Large language models (LLMs) already reason reasonably well, but the real value shows up when they can do something beyond generating text: query databases, call APIs, compute, drive a web browser. That is where the trouble starts — every provider has a different interface, tools scatter across the project, parallel calls and async are hard to reconcile, and traces of what the agent did are…

Case-based memory lets an agent improve without touching its weights

When we ask a large language model (LLM) to solve a hard problem, one well-crafted prompt no longer carries the job. In practice the work is a sequence of actions: search, read, write code, check, fix. The agent has to plan its steps, use tools and remember what it did before. Yet most agents today are either hardwired into rigid scripts that adapt badly to new conditions, or they demand…

Matrix-Game 2.0 generates interactive video at 25 FPS on a single H100

Interactive world models are a way to teach AI to sense the world rather than only describe it in words. Until recently, three obstacles stood in the way: there was not enough quality data with precise action labels; classic video diffusion models were too slow to compute and "forgot" the start of the clip; and errors compounded from frame to frame. Matrix-Game 2.0 offers a clear, practical…

OmniTry does mask-free virtual try-on by finding the spot itself

If you have ever tried to "try on" glasses or a tie on your own photo in an app, you know the catch: the system needs you to point out the region to replace by hand — draw a mask or a box. Across hundreds of item types that is awkward and scales badly. OmniTry takes a different route: the model finds the place where the object logically belongs and puts it there, with no masks and no extra…

Embodied-R1 points instead of acting and hits 87.5% on real robot tasks

Robots increasingly see the world through a camera and read our written instructions. But that "knowledge" often fails to turn into the right action: the model knows what a cup is, yet not where to put it or how to get around the objects next to it. This distance between vision and action is the seeing-to-doing gap. The Embodied-R1 team proposes…