i
DATAIST
Back to feed

Memory and context

What a model remembers between requests and within one: long-term memory, long context, forgetting.

7 articles

Case-based memory lets an agent improve without touching its weights

When we ask a large language model (LLM) to solve a hard problem, one well-crafted prompt no longer carries the job. In practice the work is a sequence of actions: search, read, write code, check, fix. The agent has to plan its steps, use tools and remember what it did before. Yet most agents today are either hardwired into rigid scripts that adapt badly to new conditions, or they demand…

BSC-Nav's three-layer memory lifts robot navigation to 78.5% success on HM3D

Most AI agents today are reactive: they see a frame and act, see the next frame and act again, and never build a coherent picture of the space around them. Hence the trouble with long routes, with reusing past experience, with flexibility. Biology solved this elegantly: the brain keeps landmarks, route knowledge and survey maps. BSC-Nav carries that principle over to robots and gives them a…

SK Hynix may rent space in Intel's Ohio fab for memory chips

SK Hynix is reportedly in talks with Intel about manufacturing memory chips in the United States. One scenario under discussion has the South Korean company leasing space inside Intel's planned Ohio fab and producing memory there. Another is a joint venture that could include cloud providers as participants. SK Hynix told TechCrunch on Wednesday that nothing has been approved: it is studying…

PrismML shrinks a 27B reasoning model to 5.9 GB, keeps 98%

PrismML, a Caltech spinout with a $22.25 million seed round behind it, released Bonsai 2 27B on Thursday. The model is a compressed version of Alibaba's open Qwen3.8 27B that fits in 5.9 GB, nine to ten times less memory than the original needs, and scores 98% of the original's aggregate benchmark results. The number to watch is not 5.9 GB. It is 98 — up from 95% for the first Bonsai, which…

Kepler Computing says it can build HBM without EUV lithography

Kepler Computing, a San Jose company founded in 2018 by a team of physicists and computer scientists, has spent more than seven years in stealth rearranging the architecture of computer memory. It now says its approach can ease the global shortage of memory chips — if it can be manufactured in volume. The pitch is a supply argument rather than a performance one: Kepler claims a high-bandwidth…

DeepSeek's V4.1-Flash targets the memory bill, not the leaderboard

DeepSeek has released V4.1-Flash, a multimodal model whose pitch is a memory bill rather than a benchmark. The KV cache — the buffer that holds already-processed context so the model does not recompute it at every step — now occupies roughly a quarter of the fast GPU memory that DeepSeek-V4-Flash needed, and the portion permanently offloaded to SSD or host memory falls to about an eighth. The…

Daniel Susskind wants every subject taught twice, with AI and without

Daniel Susskind has spent 15 years studying what AI does to work and society, and he has now written the essay a father of three writes. Its argument is that the standard policy response to technological change — work out the skills of the future and teach them to children — has already failed once in living memory. He has the receipt. In 2013 the UK government announced that England would be…