i
DATAIST
Back to feed

Reasoning

How models arrive at an answer: chains of thought, planning, self-checking and what long deliberation costs.

34 articles

Reflection’s Beam pairs 501 billion parameters with a leaner compute claim

Reflection has introduced Beam, a 501-billion-parameter text model that it says can match China’s GLM-5.2 on difficult reasoning benchmarks while using three to four times less compute at inference. The model is due to become open later this month, with its weights and full technical documentation. Those claims have not been independently verified, making Beam’s release less a settled challenge to closed labs than a test of whether Reflection can turn a striking efficiency pitch into a model others can reproduce and use.

OpenAI blocked reasoning theft, but Azure still exposed its models

OpenAI says it stopped a campaign to extract hidden reasoning from its models, but a follow-up test found the same attack still worked through Microsoft Azure. The episode exposes a gap between securing a company’s own API and securing the wider cloud platforms that serve its models. OpenAI says it blocked more than 15,000 accounts and fixed a vulnerability; researchers later extracted reasoning from every tested OpenAI model on Azure in a single attempt.

Anthropic says Claude Sonnet 5.5 cuts task costs without cheaper tokens

Anthropic’s Claude Sonnet 5.5 is designed to make each completed task cheaper, not each token. The company says the model works faster and needs fewer tool calls, while keeping API prices at $2 per million input tokens and $10 per million output tokens. Its benchmark scores now sit close to Opus 5.5 on several tests, though Anthropic still recommends Opus for work that needs sustained reasoning and judgment.

Anthropic engineer says smarter Claude learned to write for AI

Anthropic engineer Jackson Kernion says Claude’s writing got worse for a counterintuitive reason: training made it better at mathematics, programming and reasoning, while also teaching it to produce explanations tuned to other AI models. That can yield technically dense prose that a language model handles easily but a person finds overloaded. The problem matters beyond Claude because it exposes a tradeoff in reinforcement learning: rewards that improve machine understanding can pull writing away from human-readable explanation.

Why is it difficult for AI agents to work with different types of data?

AI agents can sound fluent yet still struggle to find their way through a jumble of tables, files, and databases. The authors propose EvoOntology, a layer that gives an agent a living map of what data exists, what it means, and which tools can help it use that data. Instead of relying on fixed instructions or forcing the agent to inspect every source from scratch, the system builds and continually improves this map as the agent works, adapting it to different tasks and ways of reasoning. This helps turn scattered, hard-to-navigate information into something an AI agent can actually explore and use. In this review we look at how EvoOntology is built, how it learns to refine its own understanding of data, and why this approach may make AI agents more capable across mixed data sources.

Tech executives swap AI doom for hope to sell data centers

Alexis Ohanian, co-founder of Reddit, posted that the industry needs to come up with a new way of talking about data centers; Axios picked it up. Matt Higgins, co-founder and chief executive of the investment firm RSE Ventures, was more specific. He proposed renaming them outright: "fraud prevention centers," "clean energy efficiency centers," "precision agriculture centers." His reasoning is…

Nvidia opens Alpamayo 2 Super, a teacher model for DRIVE AGX Thor

Nvidia has released Alpamayo 2 Super, an open-weights driving model that takes a 360-degree view from seven cameras and returns future trajectories, causal reasoning chains, high-level meta-actions, grounded answers about the scene and automatic reasoning labels from one input. The weights are on Hugging Face and the inference notebooks on GitHub, under OpenMDW-1.1, the Linux Foundation's…

EU designates ChatGPT a very large search engine under the DSA

The European Commission announced in Brussels on Monday that ChatGPT will be supervised as a very large search engine under the Digital Services Act. The reasoning is narrow and worth reading closely: the service was classified this way because it searches the internet and answers user queries, and because it clears the threshold of at least 45 million monthly users in the EU. Reddit and…

Salesforce and Nvidia built Koa to cut Agentforce's token bill

Salesforce and Nvidia have built Koa, a reasoning model trained for sales and customer support, and it will sit inside Agentforce next to the models Salesforce already pays for. Until now, when an Agentforce agent hit a long or multi-step reasoning task, Salesforce's AI gateway routed the prompt out to a frontier model — Claude or ChatGPT. Koa is the in-house answer to that, and the pitch is…

PrismML shrinks a 27B reasoning model to 5.9 GB, keeps 98%

PrismML, a Caltech spinout with a $22.25 million seed round behind it, released Bonsai 2 27B on Thursday. The model is a compressed version of Alibaba's open Qwen3.8 27B that fits in 5.9 GB, nine to ten times less memory than the original needs, and scores 98% of the original's aggregate benchmark results. The number to watch is not 5.9 GB. It is 98 — up from 95% for the first Bonsai, which…

OpenAI's Astra reasons in loops, and safety researchers object

OpenAI has built a technique known as opaque recursion into Astra: instead of writing out its steps, the model runs a single query through a loop several times, leaving fewer readable traces and largely bypassing the chain-of-thought record that safety teams depend on. Buck Shlegeris, CEO of Redwood, wrote that he is extremely concerned about it. Ryan Greenblatt, chief scientist at Redwood…

OpenAI's Astra puts opaque recurrence in the AI glossary

The working vocabulary of AI gained a new center of gravity this year, and it is not a capability. OpenAI's Astra, released in September 2026, is known for early use of opaque recurrence: a method in which a model pushes the same prompt through its own internal layers over and over instead of reasoning step by step in language a person can read. OpenAI says Astra preserves a legible chain of…

OpenAI claims AGI with Astra, a model it rates critical for cyber

OpenAI says its new model, GPT-6 Astra, is artificial general intelligence — by the company's own definition, "autonomous systems that outperform humans at most economically valuable work." The same model is the first OpenAI has ever placed in the "critical" category for cybersecurity capability, and the company has confirmed it shows a "substantial reduction in chain-of-thought…

KAIST and Naver find reasoning steps encoded in middle layers

Researchers at KAIST and Naver AI Lab report that the discrete steps a reasoning model writes out — pulling data, decomposing the problem, recalling a formula, computing — correspond to separable patterns inside the model's numeric representations. The separation is strongest in the middle layers, holds across three different models, and survives the case that matters most to anyone hoping to…

GPT-5 and Gemini 3 encode facts they cannot recall without reasoning

Researchers who profiled 13 large language models across more than 4 million responses report that the frontier systems have very nearly stopped forgetting: GPT-5 and Gemini 3 encode 95–98% of the facts tested against them. What those models cannot reliably do is get the facts back out. Asked directly, without reasoning, they fail to produce 26–34% of what already sits in their parameters.…

Mapping code by behavior helps agents plan edits better on fewer tokens

Conversations about AI agents give almost all their attention to models. Which LLM is stronger, whose reasoning is better, who writes more accurate code. But in real engineering work, an agent's success rarely rests on the model alone. There is another layer that assembles prompts, holds state, calls tools, and keeps the steps in order. The…

Agent quality comes from parallel reasoning and merging, not orchestration

A near-cult of engineering complexity has grown up around modern LLM agent systems. Orchestrators, sub-agents, memory, skill libraries, tool calls — it all looks impressive, and it leaves the important question unanswered: what actually produces the gain in quality? The authors of HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harness…

Taking notes, not reasoning, separates the agents that can run a startup

Agents do well on short tasks. Over a long horizon they are undone by memory, inconsistency and an inability to stick to a strategy. AI agents have gotten decent at problems that take a dozen actions, a couple of tool calls and an answer. Stretch the task to hundreds of steps — the length of real work — and it gets interesting. Early mistakes…

Predicting the answer's latent image beats text-only chain of thought

Multimodal LLMs have learned to recognize objects, but how do you give them visual imagination? A look at the Cognitive Supersensing idea. Over the past few years, multimodal LLMs (MLLMs) have learned to recognize objects, read captions, answer questions about an image and even give a decent account of what is happening in a frame. But they have…

Smarter reasoning models make collective outcomes worse in social dilemmas

As autonomous LLM agents take over human tasks — from negotiating with services to allocating resources inside companies — we have gotten used to judging them on solo benchmarks. We care how well a model writes code, answers questions, or plans. But in the real world they run into each other, compete for limited resources, and sometimes manufacture competition nobody needed. The paper…

Reasoning models get better by arguing with themselves inside one trace

We tend to assume reasoning models are stronger simply because they write longer chains of thought and burn more compute before answering. In Reasoning Models Generate Societies of Thought the authors offer a more interesting explanation: these models don't just think for longer, they start thinking differently — as though a small assembly of voices had formed inside them, putting questions to…

Absolute Zero trains reasoning with zero data by inventing its own tasks

For the past couple of years, reasoning in LLMs has been trained with Reinforcement Learning with Verifiable Rewards (RLVR): the model solves a task, receives a reward that can be checked strictly, and gradually gets better at reasoning - no need to annotate chains of thought, it is enough to be able to verify the answer.

Sophia gives agents autobiographical memory and cuts reasoning steps by 80%

Today's AI agents can plan, call tools, run chains of actions, even operate inside a multi-agent system. But most of these setups share an awkward property: they are fundamentally reactive. An agent can answer well in the moment, yet after deployment it rarely changes its own habits, rarely revisits its strategies, and almost never modifies itself. When the environment shifts — new interfaces,…

Reasoning models now pass all three CFA levels, but ethics still trips them up

In finance, the CFA (Chartered Financial Analyst) exams are a marathon run over three distances. Level I tests the fundamentals and whether you can keep the terminology straight. Level II puts you in front of cases where formulas and logic have to be applied in context. Level III asks for more than correct arithmetic: a coherent professional answer — how to build a portfolio, how to assess…

Top LLMs score 30 out of 100 on a benchmark of the full research cycle

Today's LLMs can do a lot: explain hard topics, write code, hold a long thread of reasoning. But science is not just knowing the answers. It is a research cycle: work through the literature, come up with a hypothesis, test it with an experiment, then interpret the results honestly and adjust the plan. And this is where the field has long lacked a shared language and a shared yardstick: what…

Storing verified lemmas instead of context gets an agent to olympiad gold

Over the past couple of years, large reasoning models (LRMs) have gotten noticeably better at olympiad math. On problems at the level of AIME (the American Invitational Mathematics Examination), one long reasoning pass is usually enough: the model writes out a chain of thought, checks the answer, sometimes takes a couple of runs at it — and lands it. IMO (the International Mathematical…

Sora-2 solves visual puzzles by drawing its reasoning in video

When we ask a model to reason, it reasons in words if the medium is text, or over a static scene if the medium is an image. The world, though, is not static: objects move, and the rules often only become visible in how those objects behave over time. The authors propose video generation as a general-purpose channel for reasoning. Text can be written directly into the frames, visual hypotheses…

Salesforce's EDR shows its research plan and lets you edit it mid-run

Enterprise data tends to sprawl across email, reports, databases and code repositories. Answering a hard question usually takes not one fact but many, plus the ability to synthesize hundreds of sources with checkable citations and a line of reasoning someone can follow. Ordinary agents and classic RAG systems are often a poor fit here: they give shallow answers, they are hard to steer once…

DeepAgent replaces the agent pipeline with one long reasoning stream

LLM agents can reason, but reasoning is not enough to solve real tasks. An agent also has to call outside tools, work through long scenarios and stay autonomous across dozens of steps. Rigid pipelines with fixed modes get in the way of that, and so do the classic approaches like ReAct and Plan-and-Solve. They impose the same action loop every time and work well on tasks that take two or three…

Agents that share latent thoughts instead of words reach 93% on MATH

Put several models in a multi-agent system on the same question and they will argue a little, correct each other, and settle on a compromise that is not always right. Language is what lets them do it, and it is also the bottleneck: it is sequential, often ambiguous, and rarely a faithful record of the reasoning behind it. A new paper argues for working above the level of words and giving…

Natural-language memory beats fine-tuning on long agent tasks

Large language models do well on short reasoning and coding benchmarks. But real work stretches over dozens or hundreds of steps, demands switching between applications, careful context management and the ability to catch your own mistakes. The core problem is that at test time most agents stay static: they accumulate no experience and get no better from one attempt to the next. The authors of…

Rewriting an agent's context collapses it; small edits gain 17 points

Over the past two years one thing has become clear: many applications built on large language models learn better through careful work with the context than through fine-tuning weights. Into the context go system instructions, reasoning steps, examples, domain rules, facts, even hints on how to use tools. It is transparent, portable, and it works at runtime. On top of that, progress on long…

Model reasoning traces break into the same episodes human solvers use

Large reasoning models (LRMs) today don't just answer — they unfold long chains of thought. That lets them handle harder problems, but it creates a new one: how do you describe the structure of that reasoning, and how close does it come to human thinking? The researchers propose borrowing a framework well tested in cognitive science — Schoenfeld's episode theory, built originally to analyze…

Giving planning tokens extra credit beats GRPO on math reasoning

Reasoning tasks are a sore spot for many AI systems, even ones with solid factual knowledge. A new paper shows that reinforcement learning (RL) does more than push accuracy up — it rebuilds the model's internal logic into a hierarchy that runs from low-level execution to high-level planning. That explains where those aha moments come from. More usefully, it explains why the standard algorithms…