Why agents strain old safeguards
The report grew out of a Google CAPS workshop held in New York in late 2025. Its authors say agents create three problems for conventional software defenses:
The tension is practical: agents may need access to personal information and permission to take consequential actions to be useful. The more flexible they become, the less well a fixed list of permissions describes what they should be allowed to do.
Source: research.google
Privacy depends on the situation
The report uses contextual integrity, a theory that treats privacy as appropriate information sharing under justified social norms—not merely secrecy or control over data. Those norms depend on who is sharing information with whom, whose information it is, what kind of information is involved and the terms of its transfer, such as confidentiality or reciprocity.
A person might share a gift list with a virtual shopping assistant but not with family or friends. The authors extend the idea beyond information flows: an agent’s actions, too, should be judged for whether they fit the social situation.
That shifts the question from whether an agent has permission in general to whether a particular action is appropriate in context. The report’s example is a request to protect someone’s data while arranging a conference trip. A system would need to translate that broad instruction into specific rules for booking travel, applying for a visa and communicating with organizers.
From norms to enforceable rules
The authors argue that large language models make it possible to create machine-readable rules that account for context. They propose adding a contextual policy engine to the system’s oversight layer. It would monitor actions, restrict them when needed and update rules as a user’s request or working context changes, including when the agent discovers new tools and capabilities. Before sending data, the system could check whether that transfer is appropriate.
The engine is one part of a broader set of safeguards:
Source: research.google
The hard part is proving it works
The report also calls for standardized, multi-user benchmarks in dynamic “Agent Gym” environments. Open-source sandboxes, the authors say, could let researchers safely simulate complex, long-running interactions and build a shared basis for evaluating privacy, security and reliability.
I think the policy engine is the report’s most concrete proposal—and also where its hardest unanswered question sits. The report describes rules that can change with context, but does not establish how systems should resolve disagreements over what counts as an appropriate norm, or how reliably an agent can recognize a situation before acting. Without answers to those questions, contextual safeguards risk becoming another layer of rules whose limits are hardest to see precisely when an agent is operating on its own.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X