Wiz Atlas and Microsoft MDASH reach 90%+ CyberGym bug success using routed models
Two multi-agent, multi-model systems beat rival scanners on CyberGym, arguing security needs the right model for each step.

Wiz says its Project Atlas bug-hunting agent hit a 90.9 percent success rate on CyberGym, while Microsoft’s MDASH harness scored 95.95 percent. The implication for decision-makers: continuous, evidence-based vulnerability coverage may depend on model routing, not one-size-fits-all scanning.
Two AI bug-hunting systems, one from Wiz and one from Microsoft, are posting 90 percent-plus success rates on CyberGym. Wiz’s Project Atlas achieved a 90.9 percent success rate on CyberGym, outperforming Anthropic’s Mythos Preview and OpenAI’s GPT-5.5 Cyber, and Wiz says it uncovered more than 200 zero-day security holes in widely used open-source code. Microsoft’s MDASH harness landed even higher at 95.95 percent on the same benchmark, again beating Mythos, Gemini, and GPT models evaluated on real vulnerability-finding tasks.
Both vendors point to the same practical idea: you do not get the best security outcome by picking a single model and throwing it at everything. Wiz and Microsoft describe their systems as “multi-model” and “routed” architectures, where different stages of the security workflow use the model best suited to that specific job. In other words, the breakthrough is not just intelligence, it is workflow design. If the goal is to find, validate, and remediate real vulnerabilities, the right reasoning model for the right step appears to matter as much as raw capability.
Wiz’s Atlas is not commercially available yet. Wiz says it is used internally, and it stems from the company’s effort to understand how frontier models can be used for advanced code scanning. The headline result here is straightforward: Atlas reached 90.9 percent success on CyberGym, and in the process Wiz says it uncovered more than 200 zero-day security holes in widely used open-source code. Wiz’s Nir Ohfeld, head of vulnerability research at Wiz, told The Register that Atlas uses Claude Opus 4.6 with GPT-5.5. Wiz also says it is working to incorporate Gemini, which Ohfeld frames as timed with Wiz’s recent work with DeepMind on Gemini Flash Cyber.
On the Microsoft side, MDASH is a two-part effort. Microsoft says MDASH combines “red-team” agents that find and simulate real, exploitable vulnerabilities and attack paths, and “green-team” agents that remediate the issues. Under the hood, MDASH combines MAI-Cyber-1-Flash, based on Microsoft AI’s internally developed MAI-Thinking-1 reasoning model, and GPT-5.4. Microsoft describes MAI-Cyber-1-Flash as handling up to 90 percent of all tasks, with MDASH detecting, patching, and validating vulnerabilities before handing the remaining 10 percent of more complex tasks to GPT-5.4. Microsoft executive vice president of Microsoft Security Hayete Gallot said on Monday that within its harness, this multi-agent and multi-model implementation achieved the best results it could get.
The comparisons on CyberGym show how uneven this field can be. OpenAI’s GPT-5.5 Cyber scored 85.6 percent on CyberGym, and its GPT-5.6 Sol scored 83.6 percent. Anthropic’s Mythos 5 reproduced the target vulnerability on 83.8 percent of CyberGym challenges. Google’s Gemini 3.5 Flash Cyber in CodeMender achieved an 83.2 percent success rate. The point is not that those systems are weak. The point is that Atlas and MDASH are routing work in a way that appears to reduce the mismatch between what a model is best at and where it is used.
Wiz argues the “why” is measurable using its internal benchmarking tool, Cyber Model Arena. Wiz says it evaluates every new model by scoring success across security-investigation tasks like threat modeling, hunting, validation, and proof generation. According to Wiz and co-author Yuval Avrahami, “results are rarely uniform”: the model that reasons best through a complex exploit chain is often not the one that triages most precisely. They say Atlas routes each stage to whichever model wins on that task. This matters because security teams usually do not suffer from a single failure mode. They suffer from workflow brittleness, expensive deep scans, staleness as code changes, and missing “continuous coverage” across repositories.
That “continuous and economical” framing is where the business stakes get real. Ohfeld tells The Register that pointing a frontier model at a codebase once is not a sustainable security strategy. He argues deep scans are expensive, results become stale as code changes quickly, and a point-in-time analysis cannot provide the continuous coverage organizations need. Microsoft and Wiz also claim a cost advantage from multi-model routing. Microsoft says combining its much smaller in-house model with GPT-5.4 halves customers’ costs, and Mustafa Suleyman, CEO of Microsoft AI, said that as models hand off between each other, they deliver performance better than all other models combined at 50 percent of the cost. For boards and CISOs, this is the part that can move budget conversations: higher success rates plus lower marginal cost per remediation attempt.
Zoom out, and the strategic stakes apply beyond these two vendors. Many enterprises today face a regulatory and contractual tightening around software supply chain risk, patch timelines, and demonstrable security controls. Even when regulations do not explicitly dictate “how to scan,” they raise the bar for evidence, responsiveness, and ongoing assurance. Wiz says its bet is “frontier-model depth where expert reasoning is required,” an architecture that improves as models evolve, and rigorous validation so each finding arrives with evidence, not just a plausible answer. If security leaders are choosing between pilots and scaled programs, the question shifts from “which model is best” to “how does your system take advantage of the best model available today, continuously and economically, and what continues to work when a better one arrives,” which Ohfeld frames as the central challenge behind Atlas.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology
Blacklisted Inspur Still Got Nvidia's Best AI Chips via a Subsidiary
Washington blacklisted Inspur for military ties, but its subsidiary kept shipping Nvidia's top AI chips to China's leading firms - exposing a compliance gap with huge stakes.
Cyborg cockroaches can now carry cameras and inject medicine on command
A WIRED report shows electrodes, cameras, and injection devices turning live roaches into remote medics for disaster rescue.
Isar Aerospace's Spectrum reaches orbit on second flight, a European commercial first
The German startup's second-flight success lands days before Macron's Paris summit, giving Europe a homegrown launch option as SpaceX and Blue Origin bow out.



