i
DATAIST
Back to feed

Memory and context

What a model remembers between requests and within one: long-term memory, long context, forgetting.

35 articles

Microsoft launches Surface AI PCs starting at $2,600

Microsoft has launched two Windows 11 PCs built around Nvidia chips and local AI workloads: a Surface laptop starting at $2,600 and a developer workstation starting at $6,000. The machines pair upgraded processors, graphics, unified memory and cooling with a new Windows feature for isolating AI agents. The strategy matters beyond the hardware: Microsoft says the feature will reach all Windows 11 users, making the operating system—not just the new PCs—part of its push to support agent-based software.

Cohere North 2 adds spending limits and memory for AI agents

Cohere’s North 2 adds spending limits and persistent memory to an enterprise platform for building and sharing AI agents. The update targets two practical problems companies report: discovering an agent’s costs only after the bill arrives, and getting confident answers that lack the business context to be right. North also lets customers use models they already have, while adding controls over what agents can do, what they cost and where they run.

How to Give an AI Agent More Freedom with Less Risk

The more freedom an AI agent gets, the more ways there are for hidden instructions to lead it astray. The authors propose AgentKernel, an operating-system layer that protects an agent from the moment it reads outside content to the moment it uses tools. It controls who the agent can act as, what information it trusts, what it keeps in memory, and which actions it can take—so security cannot simply be bypassed by the agent itself. The goal is not just to restrict agents, but to make it safer to give them broader responsibilities. In this review we look at how AgentKernel brings familiar computer security ideas to AI agents, and why protection built into the foundation could make autonomous systems more capable as well as safer.

How an AI agent selects past experience for a new task

An AI agent can learn from the past and still remember the wrong things for the job at hand. Rather than turning each finished task into a fixed note, the authors keep the full record and select what matters only when a new task arrives. The agent then shapes those past experiences into a concise guide for its current goal, so useful details are not discarded before anyone knows when they might help. Tests across household, shopping, and computer-use tasks show that this just-in-time approach helps agents succeed more often than existing memory methods. In this review we look at how task-specific memory is assembled, why choosing what to remember later can work better, and what the results reveal about helping AI agents learn from experience.

How AI agent self-improvement enhances results and saves tokens

What if an AI agent could improve the way it works—not by changing its core model, but by redesigning the instructions, tools, memory, and workflow around it? The authors propose a more disciplined form of self-improvement that helps an agent test and refine these surrounding components without simply memorizing the tasks it was trained on. The system limits how many changes it makes at once, explores new strategies, and removes edits that are costly, trivial, or useful only for a specific benchmark—leading to more reusable behavior and fewer tokens spent during operation. In this review we look at how regularized self-improvement works, why unconstrained evolution can fail outside familiar tasks, and how the proposed approach builds leaner, more adaptable AI agents.

Viture's $299 Vonder glasses build a memory graph without cameras

Viture is moving from virtual screens to personal memory. The San Francisco company on Tuesday introduced Vonder, its first glasses without a display or built-in cameras, at $299. The device listens through microphones and bone-conduction sensors, builds a graph of the wearer’s memories and interests, and uses several AI models to advise on everything from coffee shops to parenting. That makes Vonder less a hands-free computer than an attempt to turn a person’s daily life into searchable, machine-interpreted context.

How AI Agents Can Save Context and Avoid Failures

Autonomous coding agents can lose their way not because they cannot write code, but because they run out of room to remember what they are doing. The authors examine how the surrounding software harness—the tools and rules that guide an AI agent—affects its ability to solve long, complicated programming tasks. They compare ways to manage context, plan work, and choose actions, showing that the best setup depends on the model’s strengths and on how much memory it has available. The results reveal when planning prevents mistakes, when it mainly saves time and cost, and why simpler command-line control can sometimes work better than a larger toolbox. In this review we look at how these design choices shape an agent’s entire problem-solving path, and what they suggest for building coding systems that stay effective without wasting context.

OpenAI and rivals buy Mac minis by the tens of thousands

OpenAI and competing labs are buying Mac minis in the tens of thousands to train computer agents. Apple's smallest desktop has become a default machine for local AI work — strong chips, unified memory and cooling that holds up under long sustained loads — with interest in OpenClaw adding to the pull. A product line sold to developers and home studios is now being bought by the rack.

CXMT produces its first HBM3E while rivals mass-produce HBM4

CXMT has produced HBM3E chips for the first time, putting a Chinese manufacturer into the memory generation that AI accelerators are currently built around. It remains one generation behind Samsung, SK Hynix and Micron, all of which are already mass-producing HBM4. Two domestic customers are lined up: T-Head, Alibaba's chip unit, and Cambricon, both testing HBM3E now and planning to use it in…

SK Hynix may rent space in Intel's Ohio fab for memory chips

SK Hynix is reportedly in talks with Intel about manufacturing memory chips in the United States. One scenario under discussion has the South Korean company leasing space inside Intel's planned Ohio fab and producing memory there. Another is a joint venture that could include cloud providers as participants. SK Hynix told TechCrunch on Wednesday that nothing has been approved: it is studying…

PrismML shrinks a 27B reasoning model to 5.9 GB, keeps 98%

PrismML, a Caltech spinout with a $22.25 million seed round behind it, released Bonsai 2 27B on Thursday. The model is a compressed version of Alibaba's open Qwen3.8 27B that fits in 5.9 GB, nine to ten times less memory than the original needs, and scores 98% of the original's aggregate benchmark results. The number to watch is not 5.9 GB. It is 98 — up from 95% for the first Bonsai, which…

Kepler Computing says it can build HBM without EUV lithography

Kepler Computing, a San Jose company founded in 2018 by a team of physicists and computer scientists, has spent more than seven years in stealth rearranging the architecture of computer memory. It now says its approach can ease the global shortage of memory chips — if it can be manufactured in volume. The pitch is a supply argument rather than a performance one: Kepler claims a high-bandwidth…

DeepSeek's V4.1-Flash targets the memory bill, not the leaderboard

DeepSeek has released V4.1-Flash, a multimodal model whose pitch is a memory bill rather than a benchmark. The KV cache — the buffer that holds already-processed context so the model does not recompute it at every step — now occupies roughly a quarter of the fast GPU memory that DeepSeek-V4-Flash needed, and the portion permanently offloaded to SSD or host memory falls to about an eighth. The…

DeepSeek's 160,000-chip Huawei cluster waits on Chinese memory

DeepSeek is planning the largest Huawei-powered cluster yet: 160,000 Ascend 950DT chips in Inner Mongolia. Huawei probably cannot fill the order inside a year, held up by its own limited manufacturing capacity and by a shortage of memory chips. The order is part of a broader Chinese government program to build out a domestic semiconductor industry and keep the country's AI development moving…

Daniel Susskind wants every subject taught twice, with AI and without

Daniel Susskind has spent 15 years studying what AI does to work and society, and he has now written the essay a father of three writes. Its argument is that the standard policy response to technological change — work out the skills of the future and teach them to children — has already failed once in living memory. He has the receipt. In 2013 the UK government announced that England would be…

An editable graph of next steps beats memory for long-horizon agents

AI agents have a recurring problem. While the task is short, everything looks fine. The model reads the history, picks the next step, calls a tool, moves on. But once the horizon gets long, things start to break: the agent confuses the order of actions, repeats useless steps, forgets what it has already tried, and does too late what it should…

Multi-agent systems need explicit graphs, not smarter agents

Over the past year the industry has picked up an odd habit: when an agent fails a task, you give it more tools, more memory, a bigger context window and another try. Sometimes that works. But only up to the point where the task stops looking like a conversation with a smart model and starts looking like the work of a small team. Fixing a bug in a…

Evolving the scaffolding around a frozen model adds 17 points

AI agents come with an awkward truth: quality doesn't depend on the model alone. The scaffolding often decides everything — the system prompt, the tools, memory, the rules for choosing the next step, the checks before answering. The same LLM can behave like a careful engineer or like a chaotic intern purely because of how the pipeline is…

85.7% of what AI cites about a brand comes from someone else's site

Who actually tells AI about your brand When you ask an LLM about a company, the model rarely answers from memory. It pulls sources off the web first, then assembles an answer. And that is where it gets interesting: a brand's reputation with AI is shaped less by the company's own site than by everyone else's. 💬 A question about a company …

Finding the causal step fixed three times as many failed agent runs

There is an uncomfortable truth about AI agents: the failure almost never happens where you see it. The final answer can be wrong because 20 steps earlier the agent missed a constraint in the task, pulled the wrong record out of memory, or handed off to another agent without the context that mattered. The logs show you the symptom. The cause sits…

LightMem-Ego recalls your day well, your lost keys poorly

Today's AI assistants are good at answering questions about the here and now. But ask something like "where did I leave my badge this morning?", "what did my colleague ask me for after the meeting?" or "what do I usually do once I get to the office?" and the magic runs out fast. Answering those takes more than understanding text. It takes memory…

Teaching a 32B agent to keep better notes beat a bigger model

AI agents have an old and very mundane problem: they forget fast. Not in the sense that the context window runs out — everyone knows that one. The worse version is this: give a model external memory and it often keeps that memory like a bad set of student notes. Writes down the wrong things. Searches in the wrong place. Duplicates the obvious…

None of 12 agent memory systems wins across every workload

Memory is no longer a minor detail in an AI agent. It is usually the place where it gets decided whether the agent is useful or starts getting confused. If you follow the progress of AI agents, you have probably noticed something odd. Models have gotten better at writing code, holding a conversation, calling tools, even running long chains of…

Agent quality comes from parallel reasoning and merging, not orchestration

A near-cult of engineering complexity has grown up around modern LLM agent systems. Orchestrators, sub-agents, memory, skill libraries, tool calls — it all looks impressive, and it leaves the important question unanswered: what actually produces the gain in quality? The authors of HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harness…

What coding agents transfer across domains is discipline, not code

Coding agents share one weakness: they write code well, but they keep repeating the same mistakes, like an intern who rediscovers every time that running the tests before committing is a good idea. Over the past year researchers have been busy teaching these systems to use memory — to store the moves that worked and the ones that didn't, then…

Agents need less memory when the environment keeps the traces

In AI we are used to thinking of memory as something that sits inside the agent: in an RNN's hidden state, in the network's weights, in a replay buffer, in the KV cache. But what if part of that memory can literally be moved outside — into the environment itself? Not in a metaphorical sense but a formal one: so that the agent genuinely needs less…

Bidirectional memory: how agents evolve by remembering past steps

Today's "deep" AI agents do more than continue text: they run searches, call tools, gather facts from different sources and work through hard questions step by step. Today's "deep" AI agents can search, call tools, gather facts from different sources and work through hard questions step by step. In practice, though, an agent like this behaves…

Taking notes, not reasoning, separates the agents that can run a startup

Agents do well on short tasks. Over a long horizon they are undone by memory, inconsistency and an inability to stick to a strategy. AI agents have gotten decent at problems that take a dozen actions, a couple of tool calls and an answer. Stretch the task to hundreds of steps — the length of real work — and it gets interesting. Early mistakes…

Sophia gives agents autobiographical memory and cuts reasoning steps by 80%

Today's AI agents can plan, call tools, run chains of actions, even operate inside a multi-agent system. But most of these setups share an awkward property: they are fundamentally reactive. An agent can answer well in the moment, yet after deployment it rarely changes its own habits, rarely revisits its strategies, and almost never modifies itself. When the environment shifts — new interfaces,…

Agents with memory and a forum beat centralized search on science benchmarks

Most of today's approaches to "AI for science" look like a familiar pipeline. There is a central controlling algorithm, a metric, and a short loop: generate an improvement, run the test, keep the best one, repeat. Broadly, it works — but it also strips out what makes science science: long memory of past attempts, the exchange of ideas, argument, and the sudden transfer of methods between fields.

DeepCode rebuilds a paper's codebase by managing memory, not context

Over the past year, LLM coding agents really have learned something new: they now handle tests, run commands and get through relatively long sessions. But raise the difficulty — ask an agent to "ship the repository for this paper" — and reality arrives fast. A paper carries a lot of load-bearing detail, scattered across its sections, while the finished project is dozens of files, dependencies,…

Teaching an agent to fold its own memory cuts context 92% by step 100

Tasks that require tool use and repeated web search tend to produce long trajectories that break most LLM agents: either the agent piles up the entire history, which is punishing to carry in context, or it compresses the entire history at every step, which means forgetting details that mattered.

Natural-language memory beats fine-tuning on long agent tasks

Large language models do well on short reasoning and coding benchmarks. But real work stretches over dozens or hundreds of steps, demands switching between applications, careful context management and the ability to catch your own mistakes. The core problem is that at test time most agents stay static: they accumulate no experience and get no better from one attempt to the next. The authors of…

BSC-Nav's three-layer memory lifts robot navigation to 78.5% success on HM3D

Most AI agents today are reactive: they see a frame and act, see the next frame and act again, and never build a coherent picture of the space around them. Hence the trouble with long routes, with reusing past experience, with flexibility. Biology solved this elegantly: the brain keeps landmarks, route knowledge and survey maps. BSC-Nav carries that principle over to robots and gives them a…

Case-based memory lets an agent improve without touching its weights

When we ask a large language model (LLM) to solve a hard problem, one well-crafted prompt no longer carries the job. In practice the work is a sequence of actions: search, read, write code, check, fix. The agent has to plan its steps, use tools and remember what it did before. Yet most agents today are either hardwired into rigid scripts that adapt badly to new conditions, or they demand…