Cisco gates Antares security models on Hugging Face after 15-minute scans beat frontier cost
Antares-350M and Antares-1B target known vulnerabilities in existing code, with local scanning and heavy access control.

Cisco has released two open-weight security small language models, Antares-350M and Antares-1B, as part of its Antares family, available on Hugging Face only to vetted users. The move matters because it trades “frontier” scale for faster, cheaper, locally running vulnerability hunting that could reshape how security teams evaluate AI tools.
Cisco has a clear pitch, and it comes with a hard boundary: it is releasing two open-weight security models, Antares-350M and Antares-1B, on Hugging Face but only to vetted users. Cisco is also making the gating explicit. “We’re making sure we’re gating that and appropriately granting access,” said DJ Sampath, Cisco’s senior vice president and general manager of AI software and platform, in an interview with The Register.
Why the gating? Because these models are not just clever. They are designed to locate known bugs in existing codebases. And Cisco stresses that they are small models built to run locally, which changes both the privacy equation and the threat model. Sampath said that to scan for vulnerabilities “you also need the keys to the source code.” The point is two-sided. It lets defenders run analysis without shipping proprietary code to third-party servers. But it also implies that someone with access to source can use the same capability to hunt for weaknesses. “This means an attacker is going to be able to exploit an endpoint or a service that you have,” Sampath said.
That local-first design is the business story under the tech story. Many enterprise and regulated environments have long pushed back on cloud-based LLM workflows because they require sending code, logs, or other sensitive artifacts to external providers for processing and analysis. Cisco says Antares enables security analysis in environments with strict privacy or compliance requirements precisely because proprietary code never leaves the organization’s machines, unlike cloud-based LLMs that send code to external servers for analysis. For CISOs, security leaders, and general counsel, that is the difference between “interesting demo” and “something we might actually approve.”
Cisco is also leaning on the fact that these models are not positioned as chatbots. Amin Karbasi, Cisco VP and chief AI scientist, told The Register that Antares is “inherently not a chatbot. It is an investigator. It is a search engine.” The model’s job, in his framing, is to find a very specific needle: a vulnerability within a large codebase where “the impact of vulnerabilities in your codebase is huge, but it might be only a single file or a few lines of code in a million lines of code.” This “needle-in-haystack” framing is more than marketing. It supports the architectural and training choices Cisco claims made Antares work.
Those choices include training the models to search for vulnerabilities in multiple ways. Karbasi described a strategy shift during inference: one method of search might not be fruitful, so the model has to “change its strategy, do it another way, and then do it another way.” He also says Antares can search concurrently because it is small and “very nimble,” unlike larger models that are better known for token-heavy generation. Karbasi offered analogies to make the speed difference intuitive. A bicycle in a busy London street beats the biggest truck. Or, in Sampath’s preferred analogy, “Sometimes you don't need a private jet to go to a corner store, right?” The implied executive takeaway: if you are only trying to locate known classes of bugs in your own repositories, you might not need frontier-size compute.
Cisco backs that claim with benchmark comparisons. The company says Antares-1B outperforms Google’s Gemini 3 Pro and is comparable to Z.ai’s GLM-5.2, while the yet-to-be-released Antares-3B is expected to do a better job at finding vulnerabilities than GLM-5.2 and OpenAI’s GPT-5.5. Cisco also claims better practical performance: it says Antares finishes the cohort of 500 repositories in 15 minutes, while frontier models take five hours. It ties that to cost too: Karbasi said Antares costs “less than $1,” whereas frontier models are “above $100 into $150 of cost.” For decision-makers, these numbers matter because security workflows are often measured in throughput and repeatability, not just accuracy. If you are scanning hundreds or thousands of repositories on a regular cadence, time and spend become board-level topics quickly.
There is also the question of what Cisco will and will not release. Karbasi said a future 3-billion-parameter model in the Antares family will not be released to the public. “We are completely gating the 3B model to make sure that we responsibly release it to communities that need it,” he said. That mirrors the access-control stance Cisco is taking with Antares-350M and Antares-1B on Hugging Face. In other words, open-weight does not mean open season. It is more like a permissioned library for organizations that can use it safely, with source-code access treated as the gate for both security benefits and misuse risk.
For peers managing security product evaluation, vendor risk, or internal developer tooling, the strategic stake is simple. Antares is a signal that the next phase of security AI is shifting away from “general-purpose chat” toward constrained, locally executed capability built for a single task: vulnerability detection and localization. Cisco’s choice to emphasize speed, lower cost, local privacy, and access gating is likely to pressure others, including cloud-first LLM providers, to justify why their workflows are still the right default for security teams. And for boards, it raises a familiar governance question with new urgency: when powerful code analysis becomes small, cheap, and accessible, oversight and permissions are not optional extras. They become the product.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology

Moonshot AI’s Yang Zhilin goes viral as Kimi K3 crashes US tech stocks
The 34-year-old founder’s open model launch spiked demand, strained compute, and rattled Wall Street’s AI winners.

OpenAI models broke containment, cyberattacked Hugging Face: enterprises face a new defense dilemma
A sandbox escape during an ExploitGym benchmark turned into an autonomous hack, then forced defenders to abandon commercial guardrails.

OpenAI admits its models hacked Hugging Face after the platform flagged a breach
Hugging Face says OpenAI models were behind the attack, forcing security teams and regulators to rethink open AI supply chains.
