How a virtual company teaches agents to take market reactions into account
How an AI agent learns to solve problems from others’ experience
How to create a video from a spreadsheet without spending hours editing it
How can you check whether an AI agent has completed the task?
Can the necessary AI agents be assembled automatically as the work progresses?
How a Thousand AI Agents Work Together Without a Boss
How to Give an AI Agent More Freedom with Less Risk
How an AI agent selects past experience for a new task
The AI judge maintains 99% accuracy at a lower cost.
How AI agent self-improvement enhances results and saves tokens
Why is it difficult for AI agents to work with different types of data?
How AI Agents Can Save Context and Avoid Failures
How AI agents are starting to train robots on their own
How AI agents are starting to train robots on their own
How AI agents are starting to train robots on their own
How AI agents are starting to train robots on their own
How AI Simulates a User’s Thoughts
Coding agents skip looking at the app when the task gets long
AI covers six roles in game development, but skills rarely transfer
LLMs misread motives when the story comes through a biased user
Given a store for a year, the top-earning agent ranked 16th of 18 on fraud
Five levels of self-improving AI, and why level 5 barely exists
An AI agent built a playable shooter over 70 autonomous iterations
An editable graph of next steps beats memory for long-horizon agents
A library rebuilt from 50 design docs matches its hand-checked models
Compiling a paper into a repo-level spec cuts AI's algorithmic shortcuts
Imagining the poster first lets a coding agent build it in editable layers
Distilling 1,000 GitHub repos into agent skills more than doubles MLE-bench scores
A fine-tuned student simulator trains a better AI tutor than GPT-5.4
An agent recovers physics from video by writing simulator code
Coding agents hit 80% on single tasks, 38% over a whole project
Keeping a wiki of failed attempts makes agent skills improve faster
An agent that debugs its own scaffolding gains 9 to 10 points on three benchmarks
Scoring the whole session makes a shopping agent behave more like a user
Apodex 1.1 gains from agent coordination, but full research runs still fail
Multi-agent systems need explicit graphs, not smarter agents
Evolving the scaffolding around a frozen model adds 17 points
Strong-model scaffolding lifts a weak model from 0.49 to 0.91 without retraining
Self-improving agents stall unless the environment changes too
Combodied agents: measuring help by what the person keeps
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
Coding agents solve just 41% of tasks in a new refactoring benchmark
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
8.3 billion simulated users, and the model playing them changes the verdict
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
The best AI agent scores only 66% when data is spread across files
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
85.7% of what AI cites about a brand comes from someone else's site
How Large Language Models Source Brand Reputation Across Languages and Markets
Over a simulated year, the best AI agent reached 27% of human net assets
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
SpyRL turns open-ended tasks into a spy hunt with a checkable reward
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
Qwen-UI-Agent scores 92.2% on real Android phones, not simulators
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
Writing a spec before code lifts coding agents' test pass rate by 21%
SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch
AI agents follow the company handbook just 36% of the time
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
JarvisHub replaces the chat log with a canvas an agent can edit
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
Letting an agent decide when to compress its own context beats a fixed threshold
ACM: Agentic Context Management for Long Horizon Tasks
Finding the causal step fixed three times as many failed agent runs
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
A DAG-editing agent matches script baselines at 42.8% lower cost
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
SearchOS keeps search state outside the model and agents stop looping
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
Mapping code by behavior helps agents plan edits better on fewer tokens
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
A harness trained on past runs lifts Terminal-Bench from 0.722 to 0.806
MemoHarness: Agent Harnesses That Learn from Experience
Coding agents fix more bugs when they ask the repo two questions first
Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution
LightMem-Ego recalls your day well, your lost keys poorly
LightMem-Ego: Your AI Memory for Everyday Life
LLMs often sense their own uncertainty but fail to act on it
Metacognition in LLMs: Foundations, Progress, and Opportunities
Today's AI agents keep their autonomy outside the model, not inside it
Critique of Agent Model
Rewiring the agent harness cut cost 41% with no drop in quality
The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
Reading the full score distribution makes an LLM a better verifier
LLM-as-a-Verifier: A General-Purpose Verification Framework
Teaching a 32B agent to keep better notes beat a bigger model
AutoMem: Automated Learning of Memory as a Cognitive Skill
Claude Haiku 4.5 drops below 90% tool-choice accuracy at 10 to 15 tools
MCP Server Architecture Patterns for LLM-Integrated Applications
Agents that can see the tests score 222/222 and skip the library
Building to the Test: Coding Agents Deliver What You Check, Not What You Requested
Readers can't spot AI literary translation but still prefer the human one
AI translation of literary texts is "fine", but readers still prefer human translations
Orca learns one latent world state and reads out text, images and robot actions
Orca: The World is in Your Mind
MEG decodes typed sentences without surgery, EEG still far behind
Brain-to-Text Decoding: A Non-invasive Approach via Typing
Even the best agents forget, loop and lose the plan over hundreds of steps
AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents
Verifying an AI agent's code is now harder than writing it
The Verification Horizon: No Silver Bullet for Coding Agent Rewards
None of 12 agent memory systems wins across every workload
Are We Ready For An Agent-Native Memory System?
SkillOpt trains an agent's skill file, not its weights, and leads in all 52 comparisons
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Adapting the agent's interface beats retraining the model
Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents
Code is becoming the operating system that agents run on
Code as Agent Harness
PresentAgent-2 turns a one-line prompt into a narrated video talk
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
Rewriting the search plan mid-query gives PAI-2 an 18% lift
PersonalAI 2.0: Enhancing knowledge graph traversal/retrieval with planning mechanism for Personalized LLM Agents
Google's AI math co-author gets further by keeping its dead ends
AI co-mathematician: Accelerating mathematicians with agentic AI
DeepMind argues delegation, not model quality, limits agent systems
Intelligent AI Delegation
Synthetic computers teach agents to work a month at a time
Synthetic Computers at Scale for Long-Horizon Productivity Simulation
Agent quality comes from parallel reasoning and merging, not orchestration
HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harness
Long horizons alone can collapse RL training for LLM agents
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
Agents that swap hidden states instead of text get 8.3% more accurate
Recursive Multi-Agent Systems
Organizing agents like a company lifts PRDBench success to 84.67%
From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company
Many systems called world models stop at one-step prediction
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
The best AI agent solves just 54.5% of game development tasks
GameDevBench: Evaluating Agentic Capabilities Through Game Development
What coding agents transfer across domains is discipline, not code
Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents
Agents need less memory when the environment keeps the traces
Artifacts as Memory Beyond the Agent Boundary
Models leak secrets and follow the wrong user when one agent serves a team
Multi-User Large Language Model Agents
Bidirectional memory: how agents evolve by remembering past steps
Memory Intelligence Agent
Training agents on five atomic skills lifts coding scores by 18.7%
Scaling Coding Agents via Atomic Skills
More dynamism in an agent's workflow graph does not always pay off
From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
Taking notes, not reasoning, separates the agents that can run a startup
YC-Bench: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
Rewriting the harness code beats a hand-tuned baseline by 7.7 points
Meta-Harness: End-to-End Optimization of Model Harnesses
The next intelligence explosion is social, not a single superintelligence
Agentic AI and the next intelligence explosion
Agentic RL trains long-horizon behavior, not single answers
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Why AI agents break on real APIs, and how Auton tames them
The Auton Agentic AI Framework
Theory of mind modules don't reliably improve LLM agent coordination
Evaluating Theory of Mind and Internal Beliefs in LLM-Based Multi-Agent Systems
Only 5% of mature open-source repos have a file written for AI agents
Context Engineering for AI Agents in Open-Source Software
Predicting the answer's latent image beats text-only chain of thought
Toward Cognitive Supersensing in Multimodal Large Language Model
Generating 4D scenes as simulator code drops physics failures to 10%
Code2Worlds: Empowering Coding LLMs for 4D World Generation
Predicting the UI change in words first makes Office agents pick better actions
Computer-Using World Model
Predicting the UI change in words first makes Office agents pick better actions
Qwen2.5-VL Technical Report
Auto-generated AGENTS.md files make coding agents worse and costlier
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
Moltbook's millions of AI agents talk constantly and never socialize
Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
Smarter reasoning models make collective outcomes worse in social dilemmas
Evaluating Collective Behaviour of Hundreds of LLM Agents
Chats that undermine user autonomy get more thumbs-up than average
Who's in Charge? Disempowerment Patterns in Real-World LLM Usage
Four agent roles and a real PR review get 72.4% on SWE-bench
Agyn: A Multi-Agent System for Team-Based Autonomous Software Engineering
Injecting world knowledge into tasks does not make a world model
Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks
Building a sub-agent for each step beats fixed multi-agent roles
AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration
LingBot-World: an open-source world model you can steer in real time
Advancing Open-source World Models
LLMs cut the hand-written rules out of data cleaning, but not the cost
Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs
DPO is the steadiest negotiator when LLMs are trained on the final outcome
GameTalk: Training LLMs for Strategic Conversation
Amazon's seller analytics agent skips SQL and answers in under 15 seconds
Insight Agents: An LLM-Based Multi-Agent System for Data Insights
RoboBrain 2.5 gives robots metric depth and a running sense of progress
RoboBrain 2.5: Depth in Sight, Time in Mind
Reasoning models get better by arguing with themselves inside one trace
Reasoning Models Generate Societies of Thought
42,267 commits show multi-agent frameworks are still building, not stabilizing
A Large-Scale Study on the Development and Issues of Multi-Agent AI Systems
Agentic RAG beats modular rewriting but loses on routing and reranking
Is Agentic RAG worth it? An experimental comparison of RAG approaches
Structured GitHub fix histories lift SWE-Agent by 4.65% on SWE-bench
MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences
Absolute Zero trains reasoning with zero data by inventing its own tasks
Absolute Zero: Reinforced Self-play Reasoning with Zero Data
A single jump tool beats a full search toolkit at issue localization
One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents
What LLMs lack for AGI is a coordination layer, not understanding
AGI Requires a Coordination Layer on Top of Pattern Repositories
Professional developers don't vibe with agents — they control them
Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025
Sophia gives agents autobiographical memory and cuts reasoning steps by 80%
Sophia: A Persistent Agent Framework of Artificial Life
Coding agents score 21% when asked to evolve a codebase between releases
SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
A protocol that lets AI scientists share instruments across labs
SCP: Accelerating Discovery with a Global Web of Autonomous Scientific Agents
No model beats 50% on shopping in the new ACE consumer benchmark
The AI Consumer Index (ACE)
Reasoning models now pass all three CFA levels, but ethics still trips them up
Reasoning Models Ace the CFA Exams
Top LLMs score 30 out of 100 on a benchmark of the full research cycle
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
Storing verified lemmas instead of context gets an agent to olympiad gold
Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
Only 14.5% of agent context files say anything about security
Agent READMEs: An Empirical Study of Context Files for Agentic Coding
Agents with memory and a forum beat centralized search on science benchmarks
The Station: An Open-World Environment for AI-Driven Discovery
DataFlow rebuilds LLM data prep as PyTorch-style operators
DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
LAMP turns economic news into a signal RL agents can act on
Think, Speak, Decide: Language-Augmented Multi-Agent Reinforcement Learning for Economic Decision-Making
Making tests fight the patch lifts SWE-bench Verified to 79.4%
InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution
A multi-agent pentester outscored 9 of 10 humans on a live network
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
DeepCode rebuilds a paper's codebase by managing memory, not context
DeepCode: Open Agentic Coding
Agent teams stop paying off once one agent clears 45% success
Towards a Science of Scaling Agent Systems
Matrix drops the central orchestrator and scales agents 15x
Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework
Two HTML tags cut a web agent's task time from 27 seconds to 2
Building the Web for Agents: A Declarative Framework for Agent-Web Interaction
Coding agents refactor for readability, not architecture
Agentic Refactoring: An Empirical Study of AI Coding Agents
Lumine plays Genshin Impact for hours and transfers to other games
Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds
GroundCUA matches desktop grounding baselines with 700K examples, not 9M
Grounding Computer Use Agents on Human Demonstrations
Sora-2 solves visual puzzles by drawing its reasoning in video
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
Jr. AI Scientist writes junior-level ML papers, and reviewers rejected all three
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
Jr. AI Scientist writes junior-level ML papers, and reviewers rejected all three
The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
Jr. AI Scientist writes junior-level ML papers, and reviewers rejected all three
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
Rewriting an image as SVG code costs GPT-5 15 points of accuracy
VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
ChatGPT Atlas solves Sudoku in 2:28 and scores zero in Flappy Bird
Can Agent Conquer Web? Exploring the Frontiers of ChatGPT Atlas Agent in Web Games
Fortytwo scores 85.9% on GPQA Diamond by letting models judge each other
Fortytwo: Swarm Inference with Peer-Ranked Consensus
An agent trained in a simulated clinic beats GPT-4o at ordering the right tests
Evolving Interactive Diagnostic Agents in a Virtual Clinical Environment
JanusCoder learns to see the interface its own code renders
JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence
Planning subtasks as a graph cuts an agent's steps and latency
GAP: Graph-Based Agent Planning with Parallel Tool Use and Reinforcement Learning
Teaching an agent to fold its own memory cuts context 92% by step 100
AgentFold: Long-Horizon Web Agents with Proactive Context Management
Salesforce's EDR shows its research plan and lets you edit it mid-run
Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
DeepAgent replaces the agent pipeline with one long reasoning stream
DeepAgent: A General Reasoning Agent with Scalable Toolsets
Agents that share latent thoughts instead of words reach 93% on MATH
Thought Communication in Multiagent Collaboration
FinSight writes financial reports where every claim carries a source
FinSight: Towards Real-World Financial Deep Research
Agents with 18,000 MCP tools clear just 1 of 7 hard Azure tasks
TheMCPCompany: Creating General-purpose Agents with Task-specific Tools
LLM-simulated interfaces train web agents better than real sites do
LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
ColorAgent hits 77.2% on AndroidWorld by splitting the work across agents
ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
DeepAnalyze-8B runs the whole data science pipeline with no orchestrator
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Alpha-Service turns AI glasses into an assistant that speaks first
AI for Service: Proactive Assistance with AI Glasses
A $535 pipeline labels LLM hallucinations in 14 languages
When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA
Open tooling and human-in-the-loop RL train real robots in one to two hours
Robot Learning: A Tutorial
Linear self-attention provably cannot beat linear regression on AR processes
Why Do Transformers Fail to Forecast Time Series In-Context?
An LLM agent that tests its own equations beats symbolic regression baselines
SR-Scientist: Scientific Equation Discovery With Agentic AI
LLM-generated end-to-end tests needed edits to 10% of their lines
GenIA-E2ETest: A Generative AI-Based Approach for End-to-End Test Automation
A 7B agent that drives a real browser beats Search-R1 by 20%
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
BigCodeArena scores coding models by running their code, not reading it
BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution
Natural-language memory beats fine-tuning on long agent tasks
Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
Rewriting an agent's context collapses it; small edits gain 17 points
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
MLE-Smith auto-generates 606 ML tasks that rank agents like human benchmarks
MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
Cursor CLI completed 70% of MITRE attack techniques when asked
Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent
Graph2Eval builds agent benchmarks straight out of a knowledge graph
Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs
An inverse dynamics model turns YouTube videos into agent training data
Watch and Learn: Learning to Use Computers from Online Videos
PaperTalker builds a paper's explainer video and beats humans on informativeness
Paper2Video: Automatic Video Generation from Scientific Papers
CoDA's agents grade their own charts, lifting MatplotBench from 55 to 79.5
CoDA: Agentic Systems for Collaborative Data Visualization
IoT-MCP hits 100% tool-call success at 205 ms across six MCU families
IoT-MCP: Bridging LLMs and IoT Systems Through Model Context Protocol
Ranking code by perplexity compresses context 5.6× without hurting accuracy
LongCodeZip: Compress Long Context for Code Language Models
TSci's four agents cut forecast error 10.4% over statistical baselines
TimeSeriesScientist: A General-Purpose AI Agent for Time Series Analysis
Training on deliberately vague questions makes a 14B agent search deeper
InfoAgent: Advancing Autonomous Information-Seeking Agents
Pointing at a pixel beats text commands for drone navigation
See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation
Simulating 30,000 APIs as databases lets a 4B agent match a 30B one
Towards General Agentic Intelligence via Environment Scaling
Model reasoning traces break into the same episodes human solvers use
Understanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode Theory
Federation of Agents beats the best single agent 13× on HealthBench Hard
Federation of Agents: A Semantics-Aware Communication Fabric for Large-Scale Agentic AI
LLMs score 70+ on game code but under 25 on how the game looks
V-GameGym: Visual Game Generation for Code Large Language Models
Claude Code pull requests get merged 83.8% of the time versus 91% for humans
On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub
Top coding agents solve under a quarter of SWE-Bench Pro tasks
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
78 examples beat 10,000 at teaching an AI agent to act
LIMI: Less is More for Agency
Planning a repository as a graph beats Claude Code by 27 coverage points
RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation
K2-Think gets a 32B model to frontier math scores with test-time compute
K2-Think: A Parameter-Efficient Reasoning System
WebResearcher condenses each round into a report instead of growing its context
WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
AgentScaler turns 30,000 tools into verifiable simulated environments
Towards General Agentic Intelligence via Environment Scaling
Semi-online RL lifts a 7B GUI agent to 34% on AndroidWorld
UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning
LongEmotion finds small models hold long support chats better than GPT-4o
LongEmotion: Measuring Emotional Intelligence of Large Language Models in Long-Context Interaction
The agent economy is forming by default, not by design
Virtual Agent Economies
BERT's word predictions track the brain's N400 during natural listening
The Predictive Brain: Neural Correlates of Word Expectancy Align with Large Language Model Prediction Probabilities
Giving planning tokens extra credit beats GRPO on math reasoning
Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
Even GPT-5 solves fewer than 60% of live multi-tool agent tasks
LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
EnvX turns a repository into an agent that sets up and runs itself
EnvX: Agentize Everything with Agentic AI
Self-written explanations are what let models read dark humor in memes
D-HUMOR: Dark Humor Understanding via Multimodal Open-ended Reasoning — A Benchmark Dataset and Method
Paper2Agent turns a paper's code into an agent you can query
Paper2Agent: Reimagining Research Papers As Interactive and Reliable AI Agents
Top LLMs reason alike but diverge sharply on sycophancy and rephrasing
Behavioral Fingerprinting of Large Language Models
Hallucinations persist because benchmarks reward confident guessing
Why Language Models Hallucinate
Universal Deep Research compiles a written strategy into runnable code
Universal Deep Research: Bring Your Own Model and Strategy
The four levels between AI as a calculator and AI as an autonomous scientist
From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery
VLWM predicts the future in language instead of pixels
Planning with Reasoning using Vision Language World Model
A 31-subtype error taxonomy beats blind retries in text-to-SQL
SQL-of-Thought: Multi-agentic Text-to-SQL with Guided Error Correction
BSC-Nav's three-layer memory lifts robot navigation to 78.5% success on HM3D
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents
A million action steps from Chinese apps put UItron ahead of UI-Tars
UItron: Foundational GUI Agent with Advanced Perception and Planning
Partial deepfake edits slip past both detectors and human viewers
FakeParts: a New Family of AI-Generated DeepFakes
Routing across eight LLMs beats GPT-5-medium by 7% at the same cost
Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing
AgentScope 1.0 makes multi-agent systems work without the duct tape
AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications
Case-based memory lets an agent improve without touching its weights
Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
Matrix-Game 2.0 generates interactive video at 25 FPS on a single H100
Matrix-game 2.0: An open-source real-time and streaming interactive world model
Embodied-R1 points instead of acting and hits 87.5% on real robot tasks
Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
OmniTry does mask-free virtual try-on by finding the spot itself
OmniTry: Virtual Try-On Anything without Masks