i
Research
Review · 2026-09-28

How to Give an AI Agent More Freedom with Less Risk

Cover: How to Give an AI Agent More Freedom with Less Risk

When an AI Agent Needs Its Own Operating System

An AI agent reads a web page, searches a repository, saves its findings to memory, then runs a command or sends a message. Each step seems routine. But together, they form a chain in which someone else’s text can become an action with real consequences.

For example, a code comment might hide a prompt telling the agent to ignore previous instructions and run a malicious command. The agent could mistake it for a legitimate instruction, save it to memory, and later use it while writing code. An input filter won’t help if the rest of the system can’t tell where information came from or what the agent plans to do with it.

The authors of the AgentKernel paper propose closing this gap at the infrastructure level. Their idea: agents need an operating-system-like layer that checks not only commands sent to the computer, but also the meaning of an agent’s actions, the provenance of its data, and the permissions of whoever performs those actions.

For now, this is an architectural proposal, not a finished product backed by test results. The paper describes how such a layer could work and the rules it should use to protect AI agents.

Why Existing Safeguards Aren’t Enough

Modern AI agent frameworks make it easier to connect models to tools, memory, and other agents. Separate systems can validate tool calls, log actions, or run code in an isolated environment. But these safeguards often work as optional modules inside the same program as the agent.

That’s where the problem begins. If an agent can call a tool without going through a check, the policy won’t protect it. A sandbox may restrict code execution, but it still can’t tell whether malicious information has made its way into memory. And a computer’s built-in security tools can see that a program is writing a file, but not whether the action came from a user’s request or from text hidden on a web page.

The authors describe this as a missing infrastructure layer. A conventional operating system restricts programs’ access to files, processes, and the network. In their view, an AI agent system needs a similar mandatory intermediary for a different set of operations: delegating tasks, reading data, using memory, and calling tools.

Imagine a chain of agents reviewing a code change and, if certain conditions are met, releasing the software. If each agent can simply claim an identity, it’s hard to verify who assigned it the work. If input data aren’t labeled, a malicious comment can end up in the model’s context. If memory records only vague information about where entries came from, a false summary may later look like a reliable fact. And if command execution is controlled only by policy on paper, a malicious tool may exceed its granted permissions.

The Four Pillars of AgentKernel

The proposed system checks an AI agent’s work at four levels: identity, perception, memory, and execution. Together, they’re meant to track how data and permissions move from input to action.

Identity: Who’s Acting, and Who Granted Their Permissions?

AgentKernel proposes issuing agents digital credentials that link an agent to its creator, code, owner, and launch conditions. The key an agent uses to prove its identity is kept in a protected part of the system, not in the agent’s own code.

This matters in multi-agent systems. One agent may delegate a task to another, but the second agent shouldn’t end up with broader permissions than the first. Each time authority is delegated, the system takes the intersection of the permissions: a child agent receives only the permissions allowed to both it and its parent. Permissions can be narrowed down the chain, but not expanded.

Perception: What Makes It Into the Model’s Context?

Incoming data go through several checks. First, the system labels their source—for example, a user, a tool, or an external web page. It then looks for known attack patterns, checks suspicious wording with a language-model-based classifier, and analyzes the conversation for attacks that build up gradually.

The idea is not to rely on a single filter. If a malicious phrase slips past pattern matching, another check may catch it. And passing the filters doesn’t make the content trustworthy: its source information is retained for later stages.

Memory: How to Keep Falsehoods from Becoming Long-Term Knowledge

Every memory entry gets provenance and trust labels. When retrieving information, the system should take those labels into account, not just how closely the entry matches the query in meaning.

The authors suggest labeling individual entries rather than treating the entire memory as suspect after finding one dubious fragment. That way, useful data can be preserved without treating information from an untrusted source as verified. By default, agents can’t share memory; that access must be granted separately.

There’s a risk here: to label an entry correctly, the system must understand which source fragments support it. If it makes a mistake and mixes trusted and untrusted sources, the resulting label may assign too much trust.

Execution: From Intent to Computer Action

Before a tool is run, the system checks the call against the agent’s permissions and the tool’s description. For ambiguous or dangerous actions, it can bring in a language-model-based checker or ask a person for approval.

The main restriction should be enforced at the operating-system level. AgentKernel proposes defining an allowlist of files, network addresses, and other resources in advance, then applying it to a process and all its child processes. That way, a malicious tool shouldn’t be able to gain more access than it was granted, even if it tries to launch another program.

After execution, the system compares the planned actions with what actually happened. This can help catch extra commands—or cases where an agent claims to have done something that doesn’t appear in the action log.

What Matters Most in This Design

The architecture connects all four pillars through provenance labels. Identity helps establish who passed along the data. Input checks assess how much to trust it. The memory system carries those labels over to saved information. And the execution system uses them when deciding which actions to allow.

AgentKernel’s four pillars connect identity, input, memory, and action checks.

AgentKernel’s key rule is that the intermediary must not be bypassable. To enforce it, the system restricts an agent’s access to three resources: the language model, tools, and memory. Every request to these resources must go through AgentKernel interfaces. If the agent still has direct credentials, file access, or another way around those interfaces, the safeguards lose their purpose.

This is much like a conventional operating system, which controls a program’s access to resources instead of relying on the program’s promise to behave. But the proposal still leaves many questions open. The paper reports no performance or protection-accuracy results from real-world tasks. The authors also acknowledge that attack detection can fail, and that memory labels depend on how well the system can distinguish one information source from another.

The safeguards also assume that the underlying operating system and AgentKernel itself are trustworthy. If an agent can bypass the required interfaces, or the system kernel is compromised, the promised guarantees no longer hold. In practice, verifying that every possible bypass has been closed is difficult.

Conclusion

AgentKernel proposes treating security as a required part of AI agent infrastructure. The system should check who is acting, where data came from, how they’re stored in memory, and which real-world actions are allowed.

This approach could help make agents more autonomous: when permissions are enforced by the system, rather than left to a language model’s promises, operators may feel more comfortable giving agents access to useful tools and data. But for now, AgentKernel is an architectural design—not proof that this kind of protection is ready to deploy. The next step is to test it against attacks, measure its latency, and verify that every action really does pass through the mandatory intermediary.

AI reviews in simple way

Every day we read fresh AI papers and retell the essentials in plain human language — no hype, no fluff. If you want to see where AI agents are heading before everyone else, subscribe.

New reviews — every day

Follow on X