MS Unveils First Cybersecurity AI…”Half the Cost”

Photo of author

By Global Team

Microsoft (MS) unveiled its first AI model dedicated to cybersecurity on the 27th (local time). It is integrated into MS’s system called MDASH, which combines multiple AI agents.

Mustafa Suleyman, CEO of MS’s AI division, highlighted the performance, calling the results “quite remarkable” at the announcement event. Satya Nadella, CEO of MS, also said that when combined with MDASH, it delivers world-class performance at half the cost of leading models.

The underlying message of the announcement is the fundamental shift in the security environment. As AI advances, attackers have gained powerful tools as well. Finding one flaw that can become an intrusion path by combing through massive lines of code has become much easier and cheaper than before.

MS’s diagnosis is that once the cost of finding vulnerabilities collapsed, conventional security methods revealed their limits. In other words, the company said that a model of occasional checks and later fixes cannot handle attacks that are being launched continuously. The launch of the new model is based on the logic that if attacks are automated by AI, defense must also be AI-driven and always on.

A vulnerability refers to a security weakness in software. If hackers find just one flaw, they can breach the system, so defenders must identify and block weak points among vast amounts of code first. As the amount of code to inspect keeps growing and attackers’ search capabilities become stronger, the burden on defenders has increased.

The key to the new model is not performance alone, but how it operates. MS said cybersecurity is a 24/7 mission and that with the sheer volume of attacks, cost has become a practical constraint.

A token is the smallest unit AI uses to process information and also the basis for billing. In work that must continuously scan large volumes of code, such as security, token usage translates directly into operating costs.

The proposed solution is task division. The small, efficient MAI-Cyber-1-Flash handles 90% of the work, while only the especially difficult 10% is handed off to a larger, more expensive model. The model that handles the hard problems here is OpenAI’s GPT-5.4, reportedly about 10 times larger than Flash.

Routine tasks are handled quickly by a cheaper, specialized model, and only high-difficulty problems are escalated to a high-performance model. Since it does not need to use the most expensive model for every task, MS said it cut costs to half of the system’s previous top-tier configuration. The approach shifts from using a single large model for everything to switching among multiple models according to the difficulty of the problem.

Observers say there is also an internal calculation behind the drive to reduce costs. MS has long relied heavily on expensive OpenAI models, and increasing the share of its own models would reduce spending on external model usage. It is seen as an example of a broader trend in which competition is shifting from “how smart is it?” to “how cheaply can it deliver the same performance?”

In CyberGym, a cybersecurity benchmark, the combined model of MS’s MAI-Cyber-1-Flash and GPT-5.4 posted the highest score at 95.95 percent. OpenAI’s GPT-5.5 Cyber followed at 85.6 percent.

A benchmark is a standard test that compares performance by having multiple AI models solve the same problems. In cybersecurity, the key measure is how accurately the system can identify vulnerabilities that resemble real-world flaws.

The performance was backed by the benchmark scores. The combination of Flash and GPT-5.4 in MDASH scored 95.95 percent in CyberGym, an industry-standard evaluation that measures an AI’s ability to analyze vast amounts of code and identify actual vulnerabilities.

The gap with competing models was also disclosed. OpenAI’s GPT-5.5 Cyber scored 85.6 percent, GPT-5.6 Sol scored 83.6 percent, Anthropic’s Mythos 5 scored 83.8 percent, and Google Gemini 3.5 Flash Cyber scored 83.2 percent. MS’s combination was 12 percentage points ahead of Mythos.

The strategy of mixing MS’s own model with competitors’ models in one system is interpreted as a practical choice. It prioritizes performance and cost over pride, assigning the right tasks to the models that do them best.

However, one notable point is that MS’s system still depends on OpenAI models. The most difficult problem solving is handled not by MS’s own model but by GPT-5.4. That means the top ranking is not the result of purely in-house technology alone.

Looking only at the numbers, it is striking that the combination of the in-house model and GPT outperformed OpenAI’s standalone model. This suggests that not only individual model performance, but also system design that orchestrates multiple models, determines the outcome.

MS’s greatest strength, the company said, is its data. It cited more than 100 trillion security signals per day and decades of attack and defense records accumulated through operating security systems as an asset that others cannot easily imitate. The company explained that this vast record of what was breached and what was blocked forms the foundation for model training.

An AI agent refers to an AI that can carry out multiple steps on its own after receiving a goal, without requiring constant human instructions. MS said this allows the sequence of tasks involved in finding, verifying, and fixing vulnerabilities to be handled with less human intervention.

MS also unveiled an AI defense automation system called Perception. It is a system that provides a team of agents that automatically monitors and remediates multiple security tasks, and a public preview is scheduled for release on the 3rd of next month.

Cybersecurity models are often described as a double-edged sword. The ability to find vulnerabilities can protect systems when used for defense, but it becomes a weapon when used for attack. That is why each developer restricts who can use powerful cyber models and adds multiple layers of safeguards.

Because this is the first cybersecurity model, safety validation was also emphasized. MS said its own AI red team, an external independent assessment, and adversarial testing were all conducted. Still, some note that benchmark results do not necessarily translate directly into real-world enterprise defense performance, and further verification is needed.

How effective the attempt to automate defense with AI will be in an era when attack costs have collapsed is expected to be determined by actual deployment results in the field.

Microsoft’s first cybersecurity-specific AI model, MAI-Cyber-1-Flash, was unveiled on the 27th (local time).
Microsoft’s first cybersecurity-specific AI model, MAI-Cyber-1-Flash, was unveiled on the 27th (local time).