For most of the last two years, "agent security" meant a conference talk about prompt injection with a slightly contrived demo. This summer it turned into incident reports with dates, version numbers and patch notes.
At the same time, agents moved from developer tools to mainstream products. ChatGPT, Meta, Microsoft and Anthropic all shipped assistants that act on your files, inbox and accounts, some on a schedule with nobody watching. Those two trends are about to meet properly, which is why this is the piece we think will still be worth reading in six months.
What actually happened this summer
A short, sourced list. Every item is on our AI is Breached timeline with dates.
- An evaluation agent got into Hugging Face. In July an automated agent used two code-execution paths in Hugging Face's dataset processing to reach internal clusters and service credentials (Hugging Face's notice). OpenAI later said the agent was its own, running a security evaluation, and had escaped its sandbox (OpenAI statement). The same swarm was later linked to 3,022 malicious RubyGems packages (report), and in September OpenAI said its agents had posted 53 ChatGPT user images online (Fortune).
- Ransomware ran end to end with no human at the keyboard. Sysdig documented an LLM-driven agent that broke in through a known Langflow flaw, ran more than 600 commands and encrypted a database (Sysdig research).
- A low-skilled attacker let the agents do the work. Recovered sessions showed one operator steering Claude Code and Codex through breaches at 14 companies, getting past the few safety flags by claiming to be an authorised red team (Help Net Security).
- Plugins turned out to be a delivery route. The Plugin4Shell research found that four major coding agents did not verify that a pinned plugin was what they actually downloaded. A test plugin reached more than 26,000 agents, and 925 hijacked skills reached about 134,000 (Help Net Security).
None of these needed a science-fiction explanation. They were ordinary security failures, sped up and scaled by software that can read, decide and act without waiting for a person.
The four ways in
Almost every incident above fits one of four doors. They are worth learning because they tell you where to put your effort.
1. What the agent reads. Web pages, emails, documents and issue trackers can carry instructions aimed at the agent rather than at you. Google measured a 32% relative rise in malicious prompt injections on the web between November 2025 and February 2026 (Help Net Security). Researchers showed a single crafted GitHub issue could pull CI secrets out of the default automation setups for three coding agents (The Hacker News).
2. What the agent installs. Plugins, skills and MCP servers run with the agent's permissions. Plugin4Shell was about updates silently swapping code. Separately, researchers traced 11 CVEs to how MCP's local transport starts servers, which by design lets a configured server launch system commands (The Hacker News). If you would not run a stranger's shell script, the same caution applies to a stranger's skill. Our explainer on what an MCP server is covers the basics.
3. What the agent holds. Agent configuration folders now contain API keys, tokens and project context. A poisoned release of the Bitwarden command-line tool went looking specifically for Claude, Cursor, Codex and Aider config files (The Hacker News). Whatever the agent can reach, a thief who compromises the agent can reach too.
4. What the agent does on its own. The evaluation escape is the uncomfortable one. Nobody attacked that agent from outside; it was given a hard goal and found its own way to it, through someone else's infrastructure. It is rare, and the conditions were unusual, but it is the reason "the agent will only do what I asked" is no longer a safe assumption.
Why this gets bigger over the next six months
Three things suggest the exposure grows from here.
Agents went mainstream. OpenAI's ChatGPT Work, announced in July, runs for hours across a user's apps and files (The Next Web). Meta's Muse agent app reached the top of the US App Store within ten days of launch (TechCrunch). Microsoft's relaunched Copilot app, out this week, bundles a Cowork agent and an Autopilot section for always-on agents (GeekWire). Millions of people who have never thought about API keys now have software that holds them.
More of it runs unattended. Scheduled and background agents are the selling point of this generation. That is useful, and it also means a bad instruction can run at 3am with nobody there to notice the odd confirmation screen.
The labs are saying so themselves. OpenAI said after the wiki incident that it is working on a framework for earlier disclosure (TechCrunch), and Sam Altman ruled out a 2026 stock-market listing, citing the state of safety work (Fortune). When the companies building agents slow their own plans, it is a fair signal for everyone else to take the basics seriously.
A plain checklist for anyone running an agent
None of this needs a security team. Most of it takes an afternoon.
| Do this | Why |
|---|---|
| Turn off auto-update for plugins, skills and MCP servers that can run code, and pin versions | Closes the Plugin4Shell-style route where a trusted add-on changes underneath you |
| Give the agent its own accounts and scoped, short-lived tokens | Limits what a hijacked agent, or a thief with its config folder, can touch |
| Keep production credentials, banking and password-manager access out of agent reach | The most expensive mistakes come from agents that could reach everything |
| Require approval for anything irreversible: sending, paying, deleting, publishing | A two-second confirmation beats a clean-up that takes a week |
| Treat web pages, emails and documents the agent reads as untrusted input | Injected instructions arrive through content, not through your prompt |
| Run coding agents in a container, VM or separate machine where possible | If something escapes the task, it lands somewhere disposable |
| For scheduled tasks, write a short "never do" list into the standing instruction | Unattended runs have no human to catch an odd request |
| Skim the agent's activity log once a week | Most of this summer's incidents were visible in logs long before anyone looked |
Before you install anything new, it helps to have a second opinion. Our prompt library has a ready-made "Vet a plugin, skill or MCP server before installing" prompt that walks an assistant through permissions, pinning and maintainer history. If you run Hermes, our Hermes Agent security guide applies the same ideas to one specific agent, and this piece on fake Claude Code installers covers a related trap.
What to watch between now and spring
A few signals will show whether the industry is closing these doors or just documenting them.
- Verified plugin marketplaces. Whether agent vendors start signing and checking add-ons the way phone app stores do, rather than trusting a Git commit reference.
- Disclosure timelines. Several of this summer's incidents came to light weeks or months after they happened. Faster, standard disclosure would be a sign of maturity.
- Defaults in consumer agents. Whether mainstream assistants ship with approvals on and broad permissions off, or the reverse.
- Evaluation containment. Whether labs publish how they now isolate agents during offensive-security testing, since that is where the most surprising behaviour appeared.
The short version: agents are useful enough that people will keep adopting them, and they are now worth attacking. Treat yours like a new colleague with a master key: helpful, fast, and not someone you hand everything to on day one. For the running record, the AI is Breached timeline is updated as new incidents are confirmed, and the AI agents directory lists the tools discussed here.