Индекс разобранных работ
Научные работы по искусственному интеллекту, разобранные на русском: ссылка на оригинал, репозиторий с кодом и разбор.
186 работ
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
How Large Language Models Source Brand Reputation Across Languages and Markets
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
ACM: Agentic Context Management for Long Horizon Tasks
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
MemoHarness: Agent Harnesses That Learn from Experience
Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution
LightMem-Ego: Your AI Memory for Everyday Life
Metacognition in LLMs: Foundations, Progress, and Opportunities
Critique of Agent Model
The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
LLM-as-a-Verifier: A General-Purpose Verification Framework
AutoMem: Automated Learning of Memory as a Cognitive Skill
MCP Server Architecture Patterns for LLM-Integrated Applications
Building to the Test: Coding Agents Deliver What You Check, Not What You Requested
AI translation of literary texts is "fine", but readers still prefer human translations
Orca: The World is in Your Mind
Brain-to-Text Decoding: A Non-invasive Approach via Typing
AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents
The Verification Horizon: No Silver Bullet for Coding Agent Rewards
Are We Ready For An Agent-Native Memory System?
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents
Code as Agent Harness
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
PersonalAI 2.0: Enhancing knowledge graph traversal/retrieval with planning mechanism for Personalized LLM Agents
AI co-mathematician: Accelerating mathematicians with agentic AI
Intelligent AI Delegation
Synthetic Computers at Scale for Long-Horizon Productivity Simulation
HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harness
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
Recursive Multi-Agent Systems
From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
GameDevBench: Evaluating Agentic Capabilities Through Game Development
Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents
Artifacts as Memory Beyond the Agent Boundary
Multi-User Large Language Model Agents
Memory Intelligence Agent
Scaling Coding Agents via Atomic Skills
From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
YC-Bench: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
Meta-Harness: End-to-End Optimization of Model Harnesses
Agentic AI and the next intelligence explosion
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
The Auton Agentic AI Framework
Evaluating Theory of Mind and Internal Beliefs in LLM-Based Multi-Agent Systems
Context Engineering for AI Agents in Open-Source Software
Toward Cognitive Supersensing in Multimodal Large Language Model
Learning to Configure Agentic AI Systems
Code2Worlds: Empowering Coding LLMs for 4D World Generation
Computer-Using World Model
Qwen2.5-VL Technical Report
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
Evaluating Collective Behaviour of Hundreds of LLM Agents
Who's in Charge? Disempowerment Patterns in Real-World LLM Usage
Agyn: A Multi-Agent System for Team-Based Autonomous Software Engineering
Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks
AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration
Advancing Open-source World Models
Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs
GameTalk: Training LLMs for Strategic Conversation
Insight Agents: An LLM-Based Multi-Agent System for Data Insights
RoboBrain 2.5: Depth in Sight, Time in Mind
Reasoning Models Generate Societies of Thought
A Large-Scale Study on the Development and Issues of Multi-Agent AI Systems
Is Agentic RAG worth it? An experimental comparison of RAG approaches
MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences
Absolute Zero: Reinforced Self-play Reasoning with Zero Data
One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents
AGI Requires a Coordination Layer on Top of Pattern Repositories
Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025
Sophia: A Persistent Agent Framework of Artificial Life
SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
SCP: Accelerating Discovery with a Global Web of Autonomous Scientific Agents
The AI Consumer Index (ACE)
Reasoning Models Ace the CFA Exams
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
Agent READMEs: An Empirical Study of Context Files for Agentic Coding
The Station: An Open-World Environment for AI-Driven Discovery
DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
Think, Speak, Decide: Language-Augmented Multi-Agent Reinforcement Learning for Economic Decision-Making
InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
DeepCode: Open Agentic Coding
Towards a Science of Scaling Agent Systems
Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework
Building the Web for Agents: A Declarative Framework for Agent-Web Interaction
Agentic Refactoring: An Empirical Study of AI Coding Agents
Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds
Grounding Computer Use Agents on Human Demonstrations
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
Can Agent Conquer Web? Exploring the Frontiers of ChatGPT Atlas Agent in Web Games
Fortytwo: Swarm Inference with Peer-Ranked Consensus
Evolving Interactive Diagnostic Agents in a Virtual Clinical Environment
JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence
GAP: Graph-Based Agent Planning with Parallel Tool Use and Reinforcement Learning
AgentFold: Long-Horizon Web Agents with Proactive Context Management
Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
DeepAgent: A General Reasoning Agent with Scalable Toolsets
Thought Communication in Multiagent Collaboration
FinSight: Towards Real-World Financial Deep Research
TheMCPCompany: Creating General-purpose Agents with Task-specific Tools
LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
AI for Service: Proactive Assistance with AI Glasses
When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA
Robot Learning: A Tutorial
Why Do Transformers Fail to Forecast Time Series In-Context?
SR-Scientist: Scientific Equation Discovery With Agentic AI
GenIA-E2ETest: A Generative AI-Based Approach for End-to-End Test Automation
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution
Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent
Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs
Watch and Learn: Learning to Use Computers from Online Videos
Paper2Video: Automatic Video Generation from Scientific Papers
CoDA: Agentic Systems for Collaborative Data Visualization
IoT-MCP: Bridging LLMs and IoT Systems Through Model Context Protocol
LongCodeZip: Compress Long Context for Code Language Models
TimeSeriesScientist: A General-Purpose AI Agent for Time Series Analysis
InfoAgent: Advancing Autonomous Information-Seeking Agents
See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation
Towards General Agentic Intelligence via Environment Scaling
Understanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode Theory
Federation of Agents: A Semantics-Aware Communication Fabric for Large-Scale Agentic AI
V-GameGym: Visual Game Generation for Code Large Language Models
On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
LIMI: Less is More for Agency
RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation
K2-Think: A Parameter-Efficient Reasoning System
WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
Towards General Agentic Intelligence via Environment Scaling
UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning
LongEmotion: Measuring Emotional Intelligence of Large Language Models in Long-Context Interaction
Virtual Agent Economies
The Predictive Brain: Neural Correlates of Word Expectancy Align with Large Language Model Prediction Probabilities
Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
EnvX: Agentize Everything with Agentic AI
D-HUMOR: Dark Humor Understanding via Multimodal Open-ended Reasoning — A Benchmark Dataset and Method
Paper2Agent: Reimagining Research Papers As Interactive and Reliable AI Agents
Behavioral Fingerprinting of Large Language Models
Why Language Models Hallucinate
Universal Deep Research: Bring Your Own Model and Strategy
From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery
Planning with Reasoning using Vision Language World Model
SQL-of-Thought: Multi-agentic Text-to-SQL with Guided Error Correction
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents
UItron: Foundational GUI Agent with Advanced Perception and Planning
FakeParts: a New Family of AI-Generated DeepFakes
Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing
AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications
Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
Matrix-game 2.0: An open-source real-time and streaming interactive world model
Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
OmniTry: Virtual Try-On Anything without Masks