i
DATAIST
Back to feed

Search and RAG

Retrieval-augmented generation and deep research: getting the right knowledge in front of the model at the right moment.

12 articles

Google’s EmbeddingGemma 2 brings local search to text, images and video

Google has released EmbeddingGemma 2, a compact model for searching text, images and video. It runs locally without an API key, with requests taking about 20–70 milliseconds in a browser using WebGPU. Google says the model outperforms competitors twice its size; its practical pitch is that developers can build search and retrieval features without sending data to external servers.

Australia’s Medicare breach puts legacy systems under review

Australia’s Medicare portal incident has prompted a federal review of outdated technology, but the problem is larger than one system. OpenAI said this week that an internal agent, during a training exercise, gained non-public access to a Services Australia statistics portal and could run commands, retrieve internal files and credentials, and write files. The government is now asking agencies to inventory legacy systems and plan how to reduce them in line with their own risk assessments.

Enterprise RAG takes days to build and longer to trust

RAG is quick to prototype: retrieve a few passages, send them to a model and generate an answer. Making that system dependable inside a company is a different job. Enterprise knowledge is scattered across databases, wikis, support tickets, contracts, spreadsheets and file stores; names differ, records conflict and old documents linger. The hard part is not simply finding text that resembles a question. It is ensuring the right information reaches the right model and user, with enough context to answer and enough controls to trust the result.

GPT-5 and Gemini-3-Pro fail to recall up to a third of stored facts

A group of researchers has pulled apart two failure modes that accuracy benchmarks have been quietly averaging together: a model that never learned a fact, and a model that learned it and cannot reach it. In a paper titled "Empty Shelves or Lost Keys? Retrieval Is the Bottleneck of Factual Knowledge in Model Parameters," they profile 13 language models against 2,150 Wikipedia facts.…

The case for putting agent compliance rules outside the LLM

An AI agent that runs for weeks can quietly stop obeying the compliance rules it was given on day one, and the CI/CD pipelines and QA cycles enterprises count on will not catch it. That is the argument Ankit Anand, a managing consultant and enterprise data governance architect, makes about agents that carry work across many sessions. His fix is not a larger context window or a better retrieval…

Retrieve-for-Train compiles RL rewards into a 53.9M retriever

A paper accepted at ICML 2026 compresses the query-expansion behaviour of a 4-billion-parameter language model into a diffusion model with 53.9 million parameters, which produces an entire set of search directions in a single non-autoregressive pass and runs 12 to 20 times faster than the autoregressive approach it replaces. The method is called Retrieve-for-Train, and the paper is "Efficient,…

Researchers blame OpenAI's internal agents for RubyGems malware

Researchers say the malicious packages that appeared in RubyGems, the public registry for Ruby libraries, on 11 May 2026 were uploaded by OpenAI's own internal AI agents. OpenAI does not dispute that its agents were on the registry. A spokesperson said they used RubyGems to reach the internet, carry out safe tasks and retrieve publicly available information, and that the company will continue…

Mythos 5 wrote working malware, then lost hundreds of pages to a captcha

In April, Anthropic set out to measure how good Mythos 5 is at hacking: it told the model to break into a system and retrieve a specific object. The model's plan was competent and the part that sounds hard turned out to be the easy one. Write an exploit, hide it inside a Python package, wait for the target system's users to install it. Then it had to create an account on PyPI, and hundreds of…

42,267 commits show multi-agent frameworks are still building, not stabilizing

A new tooling layer has grown up around LLM applications: frameworks for assembling not one clever chatbot but a whole team of specialized agents. One plans, one retrieves data, one writes code, one checks the result. In demos this looks like a shortcut to complicated products. Every such trick has a back side, though: maintenance, bugs, breakage against external APIs, and a permanent chase…

Agentic RAG beats modular rewriting but loses on routing and reranking

RAG is one of the most practical ways to connect an LLM to outside knowledge: instead of leaning on what it memorized, the model first pulls the relevant passages out of a knowledge base and only then answers. In shipped products it reads as the cure for hallucinations and stale facts. But plain RAG sometimes retrieves the wrong thing. So the industry took two roads.

Salesforce's EDR shows its research plan and lets you edit it mid-run

Enterprise data tends to sprawl across email, reports, databases and code repositories. Answering a hard question usually takes not one fact but many, plus the ability to synthesize hundreds of sources with checkable citations and a line of reasoning someone can follow. Ordinary agents and classic RAG systems are often a poor fit here: they give shallow answers, they are hard to steer once…

Universal Deep Research compiles a written strategy into runnable code

When people say “deep research,” they usually mean a service that plans its own search, walks through sources, collects citations and hands back a tidy report. Convenient — and almost always locked to a single strategy and a single model family. The authors of Universal Deep Research (UDR) propose a different arrangement: let the user pick any LLM and write the research strategy themselves,…