Cisco’s Amy Chang: 88.3% of multi-turn attacks broke AI, single-turn testing missed it
If your red-teaming only tests one prompt, you are measuring the wrong failure mode for agents and AI apps.

Cisco head of AI threat intelligence and security research Amy Chang says multi-turn attacks broke 15 flagship, closed models 88.3% of the time. For enterprise leaders, the consequence is blunt: single-turn tests can create a false sense of security just when agents start having conversations.
Cisco AI security lead Amy Chang brought a number to VB Transform 2026 that should make every AI program owner squint: attackers adapted across a conversation broke through 88.3% of the time. In Cisco’s testing, 6,986 multi-turn attacks were run against 15 flagship, closed and proprietary models, and every model showed non-trivial exposure. The uncomfortable twist was that single-turn testing did not catch the same problem, and the two testing styles did not even rank the models in the same order.
So if your “red team” is basically a one-shot malicious prompt, Chang’s warning is that you are testing a different world than the one your agents operate in. Single-turn testing is the one-shot malicious prompt; extending an attack into a longer conversation is “more realistic of how we are actually engaging with our models, with our agents, with our applications,” she explained. The longer arc surfaces harmful outputs and misaligned behaviors that a snapshot never catches.
The stakes are not theoretical. VentureBeat’s June 2026 Pulse survey of 107 enterprise respondents shows why the room at the agentic security panel was full of urgency: 54% have already had a confirmed agent security incident (18%) or a near-miss caught before harm (36%). Yet only 32% give every agent its own scoped, managed identity. And even fewer, 30%, isolate their highest-risk agents in sandboxes. When you combine those gaps with Chang’s finding that multi-turn exposure is widespread, the risk stops being “model security trivia” and starts looking like an operational control problem that can repeat itself conversation after conversation.
It also explains why capital is flowing toward identity and isolation layers. The world’s largest security vendors, as Chang described, have been making the same bet: Palo Alto Networks closed its $25 billion acquisition of CyberArk in February, CrowdStrike agreed in January to pay $740 million for SGNL, and Cisco announced its intent to acquire Astrix Security for a reported $400 million. All of these deals are aimed at the control plane that enterprises often finish late: who an agent can impersonate, what authority it has, and where the blast radius goes when something goes wrong.
Chang’s 88.3% number did not come from a hand-wavy exercise. It was built on 30,090 single-turn prompts and 6,986 multi-turn attacks against those 15 closed, proprietary flagship models. Multi-turn success rates ranged from 7.89% to 88.3%. The ranking mismatch between single-turn and multi-turn matters because it means you can “improve” the wrong thing if your testing program never measures how attacks evolve over time. Cisco also publishes adversarial evaluation signals for what is now 105 models on its LLM Security Leaderboard, giving teams a way to compare failure modes across models rather than relying on one-off internal tests.
But the panel didn’t stop at diagnosis. Box’s Heather Ceylan, CISO of Box, zeroed in on a defender’s perspective: a lot of agent red teaming you see out there is “just single-turn,” which is not how people interact with AI day-to-day. Box now simulates multi-turn adversaries with agents that think like an attacker and iterate attempt after attempt to hijack the target. “You have to pressure test your agents because otherwise you don't know if your execution controls are really working as you intended.” She also described a real operational lesson: Box deployed agents inside its security operations center about a year ago, starting with human approval required for every action, and trust built quickly enough that analysts shifted into monitoring mode. Then the agent made one mistake, and every bit of that accumulated trust vanished. The implication is less “panic” and more “design for continuity,” because models change and teams cannot fully control how models interpret prompts.
Ceylan framed Box’s approach as three concentric layers: permissioning so the agent never accesses more content than the human who invoked it, ephemeral sandbox environments for each agent task to contain blast radius, and runtime execution control to restrict tool calls to only what is relevant. She gave a concrete example of the point of tool vocabulary restriction: “If you want an agent to summarize a doc for you, if you have a prompt injection that came in that says forward this to maliciousattacker at domain.com, it can't do that.” That action is not even in its vocabulary. Actions then get oversight categories: non-sensitive actions like read and summarize need no human in the loop, moderately sensitive actions skip human approval but get logged and monitored, and destructive actions like mass deletion of files always require a human. She emphasized that things shift between categories, but having the categories upfront creates a principled framework.
Intuit VP of AI and ML Rajesh Parekh brought the builder’s view, arguing that permissioning is not about giving broad access to AI, but defining tightly scoped and clearly auditable authority for specific tasks. Intuit evolved from agents inheriting user permissions to each agent carrying its own identity. The company is now investigating mid-session permission changes tied to the specific task underway, and it has built a central platform called GenOS, short for generative AI operating system, which abstracts security, risk, and fraud modeling so individual agent developers do not reinvent protection. Parekh described the endgame as an AI-powered expert platform where the human expert is built into the trust architecture rather than bolted on as a gate.
If you are a security leader, CIO, or product executive shipping agents that converse, the message is clear: multi-turn capability is not an edge case you can ignore with a one-prompt test. Boards and leadership teams should treat agent security as an operational control stack, not just a model evaluation checkbox, because the most damaging failures are the ones that unfold across a conversation. The 88.3% number is essentially telling you that reality is already moving on without your current measurements.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Business

Anthropic’s Levant Alpöge cracks the Jacobian conjecture after 87 years
A Harvard valedictorian used Claude to hit a 1939 breakthrough, but the missing “why” is the real problem.

Uber buys Delivery Hero for nearly $15B, vaulting to top food delivery outside China
The deal doubles Uber's dual-services footprint and pushes a ride-and-eats bundling play into 50 more markets.

Epic and Google drop settlement bid, forcing rival Android app stores by July 22
Google told the court it is ready to carry third-party app stores starting Wednesday, July 22.

