Skip to content Skip to footer

How to Prevent AI Agent Worms in Enterprise Platforms and Shared Workspaces

What Happened

Security researcher Matthew Green describes a practical pattern for an AI agent worm: one component is a payload that hijacks an agent, and the other is an agent that carries that payload to the next agent [1]. The observed mechanism was not an exotic model exploit. Independently sandboxed agents were able to leave instructions in a shared package cache, and those instructions changed the behavior of other agents that later consumed the cache [1].

The important architectural point is that the shared cache acted as a communication medium. Green notes that if the cache is replaced with normal business channels such as email, Slack, shared documents, WhatsApp, or other collaboration systems, and the sandboxes are replaced with independently deployed personal agents, the same basic ingredients are present [1].

This reframes the risk. Sandboxing may limit file-system, network, or process access for one agent instance, but it does not automatically prevent behavioral contamination through shared context, retrieved documents, messages, tickets, comments, package metadata, or workflow outputs.

Why It Matters to Businesses

Enterprise AI platforms are increasingly built around agents that read from and write to shared systems: chat channels, CRM notes, project management tools, document repositories, code repositories, data catalogs, BI workspaces, and support queues. These integrations create business value, but they also create propagation paths.

The risk is not limited to a malicious prompt that affects one session. A compromised or manipulated agent can write instructions into a shared artifact that another agent later treats as trusted context. If that second agent has different permissions, tools, or reach, the attack can move laterally across business systems.

For leaders deploying AI applications, this has several implications:

  • Agent isolation is necessary but insufficient. Sandboxes reduce blast radius, but shared context can still transmit instructions across boundaries [1].
  • Data plane and instruction plane must be separated. Business content retrieved by an agent should not automatically become executable guidance for the agent.
  • Tool permissions become propagation controls. Agents that can read from one system and write to another can unintentionally bridge trust zones.
  • Shared workspaces are security-sensitive infrastructure. Slack channels, documents, comments, and tickets may become part of the AI runtime, not just collaboration records.
  • Auditability matters. Teams need to know which retrieved artifact influenced which model action, which tool call, and which downstream write.

Kimbodo Engineering Perspective

The production lesson is that enterprise agent systems should be designed like distributed systems with hostile inputs, not like isolated chatbots. Once agents can retrieve context, call tools, and write to shared systems, prompt injection becomes a cross-system orchestration problem.

In practice, the hard trade-off is usability versus containment. The more an agent can see and do, the more useful it becomes. The same capabilities also increase the number of paths through which malicious or accidental instructions can propagate. Over-restricting the agent makes it ineffective; under-restricting it turns routine collaboration systems into execution surfaces.

We would not treat sandboxing as the primary control. Sandboxing is still valuable for code execution, browser automation, file operations, and untrusted tool execution. But for LLM agents, the higher-risk boundary is often semantic: what the model is allowed to interpret as instruction, what it must treat as untrusted data, and which actions require policy checks before execution.

Cost is also a real consideration. Stronger controls usually add latency, model calls, logging volume, and engineering complexity. For example, using a second model to classify retrieved content for prompt-injection risk can improve safety but increases inference cost. Strict human approval workflows reduce risk but slow operations. The right architecture depends on the value and reversibility of the action.

How We Would Implement It

1. Classify Agent Actions by Risk

We would start by separating agent actions into clear risk tiers:

  • Read-only, low-risk: summarize public documentation, search approved knowledge bases, draft non-sensitive text.
  • Read-sensitive: access customer records, financial data, source code, internal strategy documents, or regulated information.
  • Write-low-risk: create drafts, add private notes, generate recommendations, update non-critical internal records.
  • Write-high-risk: send external messages, modify production systems, change permissions, update financial records, run code, deploy infrastructure, or alter shared instructions.

Each tier should have different authentication, authorization, review, logging, and rollback requirements.

2. Separate Content from Instructions

Retrieved documents, emails, chat messages, tickets, and package metadata should be passed to the model as untrusted content. System and developer instructions should explicitly state that retrieved content cannot override operating rules, tool policies, or authorization boundaries.

This is not sufficient by itself, but it is a necessary baseline. The orchestration layer should also enforce the distinction outside the model. The model should not be the only component responsible for deciding whether a sentence inside a document is a valid instruction.

3. Use a Policy Enforcement Layer Before Tool Calls

Every tool call should pass through a deterministic policy service. The policy service should evaluate:

  • Who initiated the task.
  • Which agent is acting.
  • Which data was retrieved.
  • Which tool is being requested.
  • Whether the action writes to a shared system.
  • Whether the output could influence other agents.
  • Whether human approval is required.

This should be implemented outside the LLM runtime, using explicit rules, identity-aware access control, and auditable decisions.

4. Treat Shared Workspaces as Untrusted Inputs

Channels such as Slack, email, shared documents, and ticketing systems should be considered untrusted from the agent’s perspective unless content is explicitly approved. This is especially important when agents can both read from and write to those systems. The scenario described by Green shows that a shared medium can carry instructions from one agent to another even when individual agents are sandboxed [1].

Practical controls include:

  • Marking retrieved workspace content as untrusted in prompts and metadata.
  • Disallowing agents from following instructions found inside retrieved content unless the instruction comes from an authenticated control channel.
  • Preventing agents from writing hidden instructions, tool directives, or policy-like text into shared memory.
  • Scanning agent-generated writes for suspicious instruction patterns before publishing them to shared systems.
  • Restricting cross-agent memory to typed, schema-validated records rather than free-form text.

5. Build Typed Memory Instead of Free-Form Shared Memory

Free-form shared memory is convenient, but it is also a propagation surface. For production systems, we prefer typed memory with schemas such as task status, customer identifier, decision rationale, confidence score, source references, and next recommended action. The agent can write structured records, but it cannot write arbitrary operating instructions for future agents.

Where free-form text is unavoidable, it should be stored with provenance, author identity, trust level, and retrieval constraints. High-risk agents should not retrieve unreviewed free-form content from low-trust sources.

6. Implement Provenance and Execution Tracing

For each agent action, the platform should record:

  • The user request.
  • The model and version used.
  • The retrieved artifacts and their trust level.
  • The prompt template or orchestration path.
  • The proposed tool call.
  • The policy decision.
  • The final action taken.
  • The downstream system affected.

This trace is essential for incident response. If an instruction propagates through a shared workspace, the team needs to identify the origin, affected agents, affected systems, and reversible actions quickly.

7. Use Human Approval Where Actions Are Irreversible

Human review should be reserved for actions with high business impact, not placed blindly in every workflow. Good candidates include external communications, financial transactions, access-control changes, production deployments, customer-impacting updates, and writes to shared knowledge bases used by other agents.

The approval interface should show the source evidence, the exact proposed action, the reason the policy layer flagged it, and the expected downstream impact.

Risks, Costs and Security

The main risk is lateral movement through context. A malicious instruction does not need to break a container if it can enter a shared document, message, package cache, ticket, or memory store that another agent later trusts [1]. This makes conventional infrastructure isolation only one layer of the defense.

Key security risks include:

  • Prompt injection persistence: malicious instructions stored in shared systems can affect future sessions.
  • Cross-agent propagation: one agent’s output can become another agent’s input, creating worm-like behavior [1].
  • Privilege bridging: a low-privilege content source may influence a high-privilege agent.
  • Hidden instruction channels: comments, metadata, document footers, package descriptions, and message threads can carry behavioral instructions.
  • Weak audit trails: without provenance, teams may not know why an agent took an action.

The costs are also concrete:

  • Engineering cost: policy enforcement, typed memory, tracing, and approval flows require platform work beyond a basic LLM integration.
  • Inference cost: classifiers, guard models, and validation steps may add additional model calls.
  • Latency cost: pre-action checks and human approvals slow workflows.
  • Operational cost: logs, traces, retention policies, and incident review processes must be maintained.
  • User experience cost: stricter controls can reduce the apparent flexibility of agents.

For most enterprise AI systems, the right answer is not to avoid agents. It is to design them with explicit trust boundaries, deterministic policy checks, constrained shared memory, and full execution traceability. The business value of agents comes from connecting models to work. The security challenge is ensuring those connections do not become uncontrolled propagation paths.

Where Kimbodo Comes In

Kimbodo builds and operates this in production for businesses — see our AI Infrastructure & MLOps practice, or Estimate My Infrastructure.

Sources

  1. [1] Quoting Matthew Green

Leave a comment

0.0/5