Microsoft’s MDASH Tops Claude Mythos and GPT-5.6 Sol on CyberGym at Half the Cost
Microsoft's MDASH multi-agent system scored 88.45% on CyberGym, outpacing Claude Mythos and GPT-5.5 while claiming half the cost of its prior setup.
Microsoft’s MDASH multi-agent system scored 88.45% on the CyberGym cybersecurity benchmark, outpacing Anthropic’s Claude Mythos at 83.1% and OpenAI’s GPT-5.5 at 81.8%, while the company claims its latest configuration finds software vulnerabilities at half the cost of its previous best setup. The results, reported by Decrypt and detailed in a Medium/AIGuys writeup, cite Microsoft’s own assertions — meaning the headline numbers come directly from the company whose product won.
That provenance matters. CyberGym, which tests automated vulnerability discovery, is not a neutral third-party audit. The cost-reduction claim — half the expense of the prior MDASH configuration — is also Microsoft’s own, unverified by any independent research available in the published record. The competitive framing, in other words, is being shaped largely by the competitor holding the trophy.
What is verifiable is the architecture. MDASH deploys more than 100 AI agents working in parallel to find software flaws — a swarm approach that sits in direct contrast to the single-model strategies of Anthropic and OpenAI. The logic is straightforward: throw specialized agents at different facets of a vulnerability — reconnaissance, exploit generation, verification — and aggregate their findings. Whether that structure genuinely outperforms a frontier model on real-world bug hunting, or merely benchmarks well on CyberGym’s specific test suite, is a question the available sources do not resolve.
The lead is not entirely new, either. A Forbes piece from May 15, 2026 noted that MDASH had already outperformed an earlier “Mythos Preview” on CyberGym, suggesting Microsoft’s edge has been widening across configurations rather than appearing overnight. The latest results extend a trend the company has been building for months.
The competitive picture is messier than a clean podium suggests. A Penligent.ai analysis from July 13, 2026 found that GPT-5.6 nearly matches Claude Mythos on advanced cybersecurity benchmarks, with each model leading in different task categories — meaning a single benchmark score can obscure where a model actually excels or falls short. A separate we0.ai article published four days ago reported that GPT-5.6 SSOL$73.16▼4.02% achieved the highest average result in an AISI cyber challenge, with Mythos 5 close behind. That result cuts against the CyberGym ranking and shows just how much benchmark outcomes depend on which test suite you pick.
The escalation is unmistakable regardless of who tops a given leaderboard. Mindfort.ai describes GPT-5.6 Sol as “OpenAI’s first model family rated High for cyber” — a classification that signals the lab is deliberately pushing into offensive security capabilities. Anthropic’s Mythos line is positioned similarly. Microsoft’s swarm-of-agents approach is a different bet on architecture, not just a different bet on scale, and the cost claim — if it holds under scrutiny — would matter more than a few percentage points on any single benchmark.
Broader industry coordination is moving in parallel. Microsoft, Nvidia, and IBM recently launched the Open Secure AI Alliance to strengthen AI cybersecurity, per the same Decrypt story — a signal that even as labs compete on benchmarks, they are pooling resources around shared defensive infrastructure.
The cost question is the one most likely to reshape procurement decisions. If MDASH genuinely delivers comparable or superior vulnerability discovery at half the price of its prior configuration, the economics of AI-driven security testing shift — not because of a benchmark score, but because of the unit economics underneath it. That claim, though, rests entirely on Microsoft’s word. The next independent audit of MDASH against Mythos and GPT-5.6 Sol on a neutral test suite — not CyberGym, not AISI, but something neither lab designed — will tell whether the swarm approach holds up where it counts.