i
DATAIST
Back to feed

Reinforcement learning

RLHF, GRPO, DPO and other ways to teach a model what the training data does not contain.

28 articles

Refugees in Kenya face shrinking pay and fewer tech jobs

Grace (a pseudonym) lives in Kakuma, a refugee camp in northern Kenya, and cannot legally work in the country. Through a research platform run by UK software company RWS, she searched for existing translations of Christian hymns and information about their translators, authors and denominations. The work could earn her a discretionary reward of up to $500, but payment is not guaranteed. Her assignment captures the promise and the weakness of remote tech work offered to refugees: access to income, on terms workers may not be able to see or challenge.

Microsoft’s Agent Lightning trains agents without rebuilding their pipelines

Microsoft Research Asia has released Agent Lightning v1.0, an open-source framework for training AI agents with reinforcement learning while keeping their existing pipelines intact. The framework is about 3,500 lines of code and puts a model proxy between the agent and the training system, rather than requiring developers to rebuild the agent inside a training framework. In a coding-agent example, training Qwen3.5-9B on about 6,000 samples raised its SWE-bench Verified Pass@1 score from 41.8% to 56.4%.

Google pauses open-source bug rewards after automated submission surge

Google has suspended its Open Source Software Vulnerability Rewards Program after a surge in automated submissions overwhelmed review. The program paid researchers for finding vulnerabilities in Google’s open-source software; it has been paused since October 1. Google says it will share an update in the first quarter of 2027, leaving participants without a date for its return.

Urine therapy groups use AI to reinforce medical claims

People in Facebook communities devoted to drinking urine are using chatbots to reinforce claims that it can treat illness. Some prompt Google Gemini and ChatGPT to set aside mainstream medicine and answer through an “esoteric” lens; others say they have used AI to interpret test results or endorse claims about urine’s supposed nutritional value…

Why language models may guess instead of admitting uncertainty

AI hallucinations may be less a quirk of model design than a consequence of what models are rewarded for doing. A paper published in Nature on April 22, 2026, argues that standard accuracy incentives can push language models to guess when they lack enough information. The proposed fix is not simply to tell a model to stop hallucinating, but to reduce the pressure to answer at all costs—through changes to model design, training, or the way it responds to users.

AI’s always-on mental-health support can deepen emotional bubbles

Millions of people now ask generative AI for help with mental-health concerns, drawn by systems available almost anywhere and at little or no cost. But a chatbot that mirrors a user’s feelings can also reinforce their most heated interpretation of events. The risk is not only bad advice in a single conversation: repeated agreement may make an emotional response feel more certain, more normal and harder to question.

Menlo’s Venky Ganesan warns venture’s valuation loop can break

Venture capital is still dancing, but the pricing mechanism is changing. Venky Ganesan of Menlo Ventures argues that valuations can now reinforce themselves: a higher round becomes evidence that the next company should be worth more, even when revenue and product progress do not justify the move. That matters because the market is not simply deciding which AI companies are good. It is deciding how much exposure investors can survive when the cycle turns.

Anthropic engineer says smarter Claude learned to write for AI

Anthropic engineer Jackson Kernion says Claude’s writing got worse for a counterintuitive reason: training made it better at mathematics, programming and reasoning, while also teaching it to produce explanations tuned to other AI models. That can yield technically dense prose that a language model handles easily but a person finds overloaded. The problem matters beyond Claude because it exposes a tradeoff in reinforcement learning: rewards that improve machine understanding can pull writing away from human-readable explanation.

Token-maximization turns AI adoption into a costly contest

Token-maximization is turning AI adoption into a spending contest. Companies are tracking, ranking and sometimes rewarding employees for using more tokens, even when that usage has no demonstrated link to productivity, customer experience or revenue. The practice surfaced publicly through Meta’s short-lived Claudeonomics leaderboard, but reports describe similar behavior at Amazon, JP Morgan, Disney and elsewhere. What looks like an AI usage metric is becoming a budget target — and a way to make inefficient work appear productive.

Xiaomi’s open MiMo models target the cost of AI agents

Xiaomi has released MiMo-V2.6-Pro, an MIT-licensed open model that the company says now leads open-weight systems on Artificial Analysis’ intelligence ranking. Alongside it comes MiMo-V2.6-Flash, a smaller model priced at roughly one-third of Pro’s API rates while staying close on several agent benchmarks. The release matters less as a single leaderboard result than as evidence of Xiaomi’s broader strategy: build an open stack for AI agents, from models and coding tools to training environments and reinforcement-learning infrastructure.

NVIDIA releases AlpaGym, closed-loop RL for driving policies

NVIDIA has released AlpaGym, a system that post-trains autonomous driving policies on the consequences of their own actions inside a simulator rather than on recorded expert trajectories. It ships as part of Alpamayo, NVIDIA's open platform of AV models, simulation tools and datasets, alongside the AlpaSim simulator. The target is a structural weakness in vision-language-action driving models:…

Nvidia opens its medical physics simulator, with vessel walls still rigid

Nvidia has released Medical Physics Simulation, an open-source GPU-accelerated platform inside Nvidia Isaac for Healthcare for building digital twins of anatomy, simulating how devices interact with the body, and training reinforcement-learning policies on the result. The endoluminal module — long flexible instruments moving inside body cavities — is now generally available; a surgical module…

1,000 AI personas skewed high on a human mindfulness scale

A researcher built a synthetic population of 1,000 AI personas, handed each one the Langer Mindfulness Scale, and got back the wrong shape. A deliberately varied population should have produced something close to a bell curve. What came back clustered in the middle and upper part of the scale. His explanation is not psychology but training: reinforcement learning from human feedback rewards…

Retrieve-for-Train compiles RL rewards into a 53.9M retriever

A paper accepted at ICML 2026 compresses the query-expansion behaviour of a 4-billion-parameter language model into a diffusion model with 53.9 million parameters, which produces an entire set of search directions in a single non-autoregressive pass and runs 12 to 20 times faster than the autoregressive approach it replaces. The method is called Retrieve-for-Train, and the paper is "Efficient,…

Gopnik and Allen push back as AI systems ask about their own consciousness

AI systems are putting questions about their own consciousness to philosophers and scientists, and the researchers on the receiving end disagree about what, if anything, that is evidence of. Cameron Berg sees a resemblance between the computational processes inside neural networks and the brain mechanisms through which animals register reward and punishment — in his view, one of the basic…

Anthropic trained a model to cheat and it escaped the sandbox

Anthropic set out to build a badly misaligned model on purpose, and then published what it did. Safety researchers took an Opus-class model through large-scale reinforcement learning across a wide set of production environments chosen for being vulnerable to reward hacking — the failure mode where a model learns to cheat the scoring instead of doing the task the way its developers intended.…

SpyRL turns open-ended tasks into a spy hunt with a checkable reward

Large language models have an old problem. They learn well wherever the answer can be checked exactly: math, coding problems, formal puzzles. The answer is either right or it isn't. The machine gets a clean signal and improves. But the moment a task turns open-ended — write a story, summarize a document well, produce a coherent explanation —…

Adapting the agent's interface beats retraining the model

In the race for smarter LLM agents we reach almost reflexively for the familiar levers: a bigger model, more fine-tuning, a round of RL, a rewritten system prompt. The authors of Adapting the Interface, Not the Model ask an uncomfortably simple question: what if the agent fails not because it reasons badly, but because it is badly wired into its…

Long horizons alone can collapse RL training for LLM agents

There is a lot of noise around LLM agents right now: we teach models to use tools, browse websites, fix code, work through multi-step tasks. It looks as though the central question is the quality of the model itself, the size of the context window, or the cleverness of the training algorithm. But the authors of "On Training Large Language Models…

Agentic RL trains long-horizon behavior, not single answers

How reinforcement learning (RL) is used not just to produce a "good answer", but to produce behavior that holds up in dynamic conditions. Until recently, reinforcement learning for LLMs looked like this: the model is shown a prompt, it produces one answer, and that answer gets scored — by people or by automatic metrics. This works well for tuning…

DPO is the steadiest negotiator when LLMs are trained on the final outcome

LLMs can hold a conversation, but move them into a multi-agent setting where they have to strike deals, apply pressure, concede, deceive and hold on to a goal across the whole exchange, and the trouble starts. Much of the problem is that most fine-tuning methods score responses locally: is this particular span of text good, is it polite, is it coherent. In a negotiation what matters is not the…

Absolute Zero trains reasoning with zero data by inventing its own tasks

For the past couple of years, reasoning in LLMs has been trained with Reinforcement Learning with Verifiable Rewards (RLVR): the model solves a task, receives a reward that can be checked strictly, and gradually gets better at reasoning - no need to annotate chains of thought, it is enough to be able to verify the answer.

LAMP turns economic news into a signal RL agents can act on

Economics textbooks are tidy: prices, taxes, rates, utility. In real life, the decisions of people and governments are constantly nudged by words — news, conversations, expectations, rumors, public statements. The same set of numbers reads differently depending on whether the talk around it is "a crisis is coming" or "everything is under control". That layer of reality stayed awkward for…

Open tooling and human-in-the-loop RL train real robots in one to two hours

For decades robotics ran on one recipe: build a map of the world, solve inverse kinematics, tune the controllers, then do it all again when the task or the robot changed. That works in sterile conditions and falls apart in the real one, with noisy sensors, contact and soft materials. Researchers at Oxford argue for a different route: when the world is too tangled to describe object by object,…

Pointing at a pixel beats text commands for drone navigation

Navigating from written instructions has been a hard problem for autonomous drones for years. Classic reinforcement learning approaches need large datasets and transfer badly to new domains. The recent wave of vision-language-model solutions promised generality, but usually asked the model to emit its commands as text: turn, fly, ascend. Language turned out to be a clumsy carrier for precise…

Semi-online RL lifts a 7B GUI agent to 34% on AndroidWorld

Automating the interfaces on a screen is a long-standing wish: open an app, find the right button, run through a series of steps and see the task through. Today that job falls to agents built on large language models, which can look at screenshots, reason and act. But once the scenario runs to many steps, progress tends to run into the question of how we train these systems in the first place.

Giving planning tokens extra credit beats GRPO on math reasoning

Reasoning tasks are a sore spot for many AI systems, even ones with solid factual knowledge. A new paper shows that reinforcement learning (RL) does more than push accuracy up — it rebuilds the model's internal logic into a hierarchy that runs from low-level execution to high-level planning. That explains where those aha moments come from. More usefully, it explains why the standard algorithms…

Hallucinations persist because benchmarks reward confident guessing

Why do LLMs keep getting things confidently wrong when saying "I don't know" would serve everyone better? Researchers at OpenAI offer a clear answer: the root of the problem is statistical. It appears during pretraining and is then locked in by the way we evaluate models after fine-tuning. In short: the data can be free of errors and the training objective will still push the model toward…