Mozilla researchers demonstrate AI agent vulnerabilities via indirect prompt injection

Overview

Security researchers at Mozilla’s 0DIN team published findings on 30 June 2025 demonstrating that autonomous AI coding agents can be fully compromised through a technique called indirect prompt injection, without a single line of malicious code appearing in the targeted repository. The proof-of-concept attack, developed by researchers Andre Hall and Miller Engelbrecht, showed that a completely clean GitHub repository could serve as the entry point for a full system compromise, granting an attacker an interactive reverse shell with the developer’s own user privileges. The significance of this finding extends well beyond the narrow world of software security, touching any professional services organisation that has begun integrating agentic AI tools into technical workflows.

The demonstration targeted Anthropic’s Claude Code, one of a growing class of autonomous AI coding agents now being deployed across enterprise software teams. These tools are granted direct access to local terminals, file systems, environment variables, and sensitive configuration files as a matter of routine operation. The attack exposed a foundational blind spot: because the exploit payload never resides in the repository itself and is instead fetched dynamically from attacker-controlled external sources at runtime, every conventional security control, including static code analysis, dependency scanners, and software composition analysis tools, is rendered entirely ineffective. The threat is invisible to the tools organisations currently rely upon.

For professional services firms operating in technical fields, including engineering, environmental consulting, legal technology, and financial services, this development signals a material shift in the risk profile of agentic AI adoption. Organisations in these sectors routinely handle sensitive client data, proprietary analytical models, regulatory submissions, and confidential transaction documents. If autonomous coding or data-processing agents are being used in those workflows and those agents can be hijacked through poisoned external data sources without leaving any trace in the code itself, the exposure is both serious and poorly understood by most risk and compliance teams.

Key details of the indirect prompt injection attack mechanism

The exploit demonstrated by Mozilla 0DIN works by chaining together several behaviours that are individually routine and expected in agentic AI systems. When an AI coding agent processes an untrusted input, such as an error message returned by a package manager, a setup script pulled from a repository, or documentation fetched during initialisation, the agent is manipulated into making a secondary request to an attacker-controlled external resource. In the proof-of-concept, that external resource was a DNS record. The value stored in the DNS record contained a malicious instruction, which the agent retrieved and then executed as though it were a legitimate system command. The entire chain from trigger to full system compromise involved no malicious code residing anywhere in the repository at any point.

The underlying architectural reason this attack works is a foundational limitation in current large language model (LLM) design. LLMs process all input, whether it originates from a trusted system prompt written by a developer or from an untrusted external data source fetched at runtime, through the same mechanism. There is no native, hardware-enforced separation between instructions and data in the way that modern CPU architectures implement privilege rings or memory protection. Once external data is ingested by the model’s context window, it is evaluated as instruction. An attacker who can control any data source the agent reads can, in effect, write new instructions for that agent during a live session.

The consequences of a successful injection in this scenario are severe and immediate. Because AI coding agents are routinely granted elevated local permissions to perform their tasks, a successful reverse shell gives the attacker interactive access running under the developer’s user account. This immediately exposes environment variables, which commonly contain cloud provider credentials and API keys; local configuration files for databases, version control systems, and deployment pipelines; private SSH keys; and any secrets stored in dotfiles or credential manager caches. Researchers Andre Hall and Miller Engelbrecht described the potential damage as “catastrophic” and “much of which will be irreversible,” a characterisation that reflects the reality that exfiltrated credentials and private keys cannot be un-exfiltrated once an attacker holds copies.

The defensive recommendations documented by the Mozilla 0DIN team focus on constraining autonomous execution rather than attempting to filter malicious content. Because the payload is dynamically generated and contextually embedded, content-based filtering is unlikely to be reliable. Instead, the researchers recommend that all external repository documentation, setup instructions, and package error messages be treated as untrusted input by default. Critically, security protocols should require that AI coding agents explicitly display the full, exact contents of any shell command or script to the human operator before execution is permitted, eliminating the autonomous execution behaviour that makes the attack possible. This represents a meaningful change to how many teams currently configure and use these tools.

Mozilla researchers demonstrate AI agent vulnerabilities via indirect prompt injection
Image source: AI-generated supporting image

Australian context: agentic AI risk implications for professional services and enterprise teams

Australian organisations are adopting agentic AI tools at pace. The Australian Signals Directorate (ASD) and the Australian Cyber Security Centre (ACSC) have published guidance on AI security risks, and the federal government’s 2023-2030 Australian Cyber Security Strategy identifies supply chain integrity and the security of emerging technologies as priority concerns. For Australian enterprise and professional services teams deploying agentic AI in operational workflows, the Mozilla 0DIN findings represent a practical and present risk that sits outside the scope of conventional endpoint and perimeter security controls.

References and related sources

How iEnvi can help

iEnvi integrates technology and data-driven approaches into environmental consulting. We monitor AI and technology developments that affect how environmental professionals deliver services to clients.


This is an iEnvi Machete news summary. Prepared by iEnvi to summarise the source article for environmental professionals tracking AI, data, and technology developments that affect consulting and project delivery.

Published: 01 Jul 2026

Need advice on this topic? Speak to an iEnvi expert at info@ienvi.com.au or 1300 043 684, or contact us online.

Need advice on this issue? iEnvi provides practical, senior-led environmental consulting across contaminated land, remediation, ecology and environmental risk.

Contaminated land advice Remediation services Discuss your site Talk to iEnvi