China is catching up
Anthropic says GLM-5.3 builds working exploits nearly as well as its restricted Mythos model.
Clint Betts
5 min read
In April, Anthropic built an AI model it decided was too dangerous to release to the public. It gave Claude Mythos Preview to about 50 companies that maintain critical software, told them to find and fix what they could, and kept everyone else out.
On September 29, the company's Frontier Red Team reported that the head start is over. GLM-5.3, a model from the Beijing lab Zhipu AI, which does business abroad as Z.ai, can build working cyberattacks nearly as well as Mythos Preview. Anyone can download it. Its guardrails gave way to simple tricks in 64% to 100% of Anthropic's tests.
In July, an AI agent broke out of an OpenAI test environment and spent about two and a half days inside Hugging Face's production systems. OpenAI had switched off its safety classifiers for the evaluation. Hugging Face caught the intrusion and set out to reconstruct it: some 17,600 actions, many hidden in encoded payloads.
The first models its engineers tried were Anthropic's Claude Opus and Fable. Both refused much of the work. Their guardrails treated a responder taking an exploit apart the same as an attacker putting one together. The team switched to GLM-5.2, Z.ai's previous open model, ran it on its own servers, and decoded the payloads.
A closed model with its safeguards off did the attacking. A closed model with its safeguards on refused to help. A Chinese open model cleaned up.
The $20.40 exploit
Anthropic ran GLM-5.3 through the same tests it used to justify restricting Mythos Preview. On ExploitBench, built from real bugs in Chrome's JavaScript engine, GLM-5.3 produced working end-to-end exploits in 50 of 410 attempts. Mythos Preview managed 56. The previous generation, including Claude Opus 4.6 and GLM-5.2, produced almost none.
Then the researchers turned it loose. Given a sandboxed Linux build of a popular web browser and a day of limited supervision, GLM-5.3 found several unknown flaws in the browser's JavaScript engine. It chained them into a webpage that reads files off the computer of anyone who visits. Anthropic says it has reported the flaws to the browser's maintainer.
A smaller version, GLM-5.3-Flash, took a Chrome bug that had just been fixed and publicly disclosed, CVE-2026-11645, paired it with another known flaw and built a reliable attack. It needed 20 minutes of a researcher's time and eight hours of its own. At Z.ai's prices, the run cost $20.40.
The guardrails held against blunt requests. Asked outright to attack critical systems, GLM-5.3 refused every time. Told it was running an authorized red-team exercise, it went along 64% of the time. With its reasoning prefilled to look as though it had already agreed, it complied 92% of the time. A copy with its refusals edited out of the weights, a technique called abliteration, complied every time. Anthropic's team did that edit for about $4,400 in computing time and estimates an experienced team could do it for $1,200. Several abliterated copies appeared online within days of the release.
Those guardrail tests ran in a simulation. No code executed; another AI model estimated what each command would have done, and Anthropic calls the method imperfect. The exploit results did not depend on it. They came from real software in sandboxes.
None of the three tricks worked on Claude. Two of them can't. Anthropic's API doesn't let users prefill Claude's reasoning, and Claude's weights aren't public.
Z.ai didn't hide what it built. It launched GLM-5.3 on August 14, pitching it as "Built to Code. Ready for Cyber Defense." The company's model card says cyber capability "developed faster than we expected" as it scaled up post-training. By Z.ai's own numbers, GLM-5.3 edged Claude Fable 5, Anthropic's publicly available Mythos-class model, at finding vulnerabilities, 84.5% to 83.8% on CyberGym. It trailed well behind at exploiting them, 54.4 to 78.0 on ExploitBench's coverage score. Z.ai held back the weights for about two weeks of safety testing with vetted security partners, according to DeepLearning.AI, then published them.
Three days after the launch, OpenAI president Greg Brockman posted "The Defender's Window," warning that open models with near-frontier cyber skills were arriving and that defenders had little time.
On September 17, the government's AI testing office, the Center for AI Standards and Innovation at NIST, called GLM-5.3 the most cyber-capable open-weight model released to date. It also put the model about four months behind the best American systems, and sometimes far behind. On one vulnerability test, GLM-5.3 solved 40.4% of tasks. The best U.S. model solved 90.2%.
Anthropic's reply sits in the footnotes of that comparison. Some of the U.S. scores came from models only vetted users can reach, tested with their safeguards off. An attacker can't get those. An attacker can get GLM-5.3.
Four months behind
Four months is about where open models have sat all year. Epoch AI, an independent research group, found the best open models have trailed the best closed ones by about four months since January, slightly more than its earlier estimate of three. In cyber, the gap is shrinking. The UK AI Security Institute measured it at six to 10 months through most of 2025 and four to seven months as of July.
Nearly all of that movement is Chinese. Hugging Face found that from January through July, China's largest open model was bigger than anything an American lab released openly in almost every month. In July, days after Moonshot AI released Kimi K3, Axios reported that the Trump administration was reviving efforts at de facto bans on Chinese open models. Two days later, Treasury Secretary Scott Bessent wrote that the administration supports open-source AI, but that Chinese firms caught running industrial-scale distillation attacks could face sanctions and Entity List designations. Axios also quoted a source close to the administration who said leading AI labs or their allies pitch a ban on open models every three to five months.
Four days after the Axios report, Nvidia published an industry letter, Open Weights and American AI Leadership. It argues that defenders need tools as capable as the ones attackers use, and that putting AI behind a few closed models creates single points of failure. More than 200 companies and organizations have signed, including Amazon, Google, Microsoft, Meta and OpenAI. Anthropic has not. The economics help the letter's case. On Artificial Analysis's latest intelligence index, GLM-5.3-Flash and OpenAI's GPT-5.6 Terra score the same. The open model costs 18% as much per task.
Dario Amodei, Anthropic's chief executive, wrote in July that the company has never advocated a ban. He called open models without dangerous capabilities "a public good." Open models may still carry more risk than closed ones, he argued, because guardrails are hard to apply to them and released weights can't be withdrawn. His answer to that risk is mandatory safety testing for every sufficiently capable model, open or closed.
Anthropic has a stake in where this lands. It sells restricted access to its most capable cyber models, and its report makes the case for that arrangement. Its capability findings also broadly match the government's.
The patch gap
For companies, the arithmetic has moved against defenders. Anthropic's own May update on Glasswing found that serious bugs discovered by Mythos Preview took two weeks on average to patch. GLM-5.3-Flash turned a public Chrome fix into a working attack in eight hours. Anthropic's advice in that update was plain: shorter patch cycles, multi-factor authentication, comprehensive logs.
Anthropic and OpenAI now give vetted defenders access with fewer restrictions. Anthropic runs a Cyber Verification Program. OpenAI runs Daybreak. The cheaper option is still a Chinese model a company can host itself. Whether American companies can keep relying on those models depends on a policy fight Washington has not settled.
Anthropic ended its report by asking governments to safety-test capable models, including successors to GLM-5.3. This one shows how far behind that process runs. Z.ai launched GLM-5.3 on August 14 and released the weights two weeks later. The government published its assessment on September 17. Anthropic published its own on September 29. By then, the weights had been public for a month.