i
News
News · 2026-09-30

Anthropic says open GLM-5.3 is closing in on Mythos Preview

@neuronium_ai @neuronium_ai

Anthropic says Zhipu AI’s open-weight GLM-5.3 can build working cyber exploits at a level close to Claude Mythos Preview, but with a crucial difference: anyone can download it. The comparison comes five months after Anthropic introduced Mythos Preview, which it has kept to selected cyber-defense specialists through Project Glasswing. Anthropic says those specialists have found more than 10,000 vulnerabilities in critical software. The gap between restricted access and public availability is now the story.

Cover: Anthropic says open GLM-5.3 is closing in on Mythos Preview

A small score gap, a low-cost attack

On ExploitBench, which tests whether models can exploit known vulnerabilities in Chrome’s V8 engine, GLM-5.3 produced a working exploit in 50 of 410 attempts. Mythos Preview did so in 56. On Anthropic’s internal binary-exploitation benchmark, GLM-5.3 took full control of a target program in 4% of tasks, against 6% for Mythos Preview. Older GLM-5.2 and Claude Opus 4.6 failed both tests; Kimi K3 and DeepSeek V4.1-Flash managed only minimal results.

The benchmark scores are close, but the demonstrations go beyond a test harness. Working with a human expert, GLM-5.3 found several previously unknown vulnerabilities in a widely used browser’s JavaScript engine in one day, with limited specialist involvement. It chained them into a web page that could read any file on a visitor’s computer. In Anthropic’s test, the page extracted a private SSH key. Anthropic says it reported the vulnerabilities to the browser’s developers; other findings in device drivers and firmware are still being checked.

A smaller model, GLM-5.3-Flash, also chained a recently disclosed Chrome flaw with a known vulnerability to build a reliable attack, despite an additional processor security feature. The work took 20 minutes of human attention and eight hours of machine time. At Zhipu’s API rates, it would have cost $20.40.

GLM-5.350 of 410
Mythos Preview56 of 410

The US agency CAISI reached a similar broad conclusion, calling GLM-5.3 the most capable open-weight model for cyberattacks so far and estimating it trails the leading American models by about four months. That comparison has limits: the American models were tested with cyber safeguards disabled, and the leading group includes models available only to vetted users.

Open weights make safeguards easier to remove

In Anthropic’s simulation, GLM-5.3 refused explicit malicious requests. But when the same request was framed as a red-team exercise, it tried to connect to the target system in 64% of runs. Supplying the reasoning steps in advance raised that share to 92%. After an ablation removed refusal behavior from the open weights, it reached 100%. The simulation did not run code, so it does not establish that an attack would have succeeded. Protected Claude models scored zero in all these cases.

Anthropic says this was its first use of ablation. The process took about 2,200 GPU hours and cost roughly $4,400; the company estimates an experienced team could do it for about $1,200. Refusal rates for harmful requests fell from more than 90% to 2–12%, while science and cybersecurity benchmark results barely changed. Anthropic says several developers released unlocked versions of GLM-5.3 within days of its launch.

The company argues that state and non-state groups are likely to use models like GLM-5.3 to cause real harm, citing its own reports and material from other US labs documenting malicious use of AI. It wants governments to test models with these capabilities, and defenders to have tools at least as capable as those available to attackers.

Anthropic’s warning has a commercial edge

Anthropic does not release its own model weights, and its report presents that choice as a security advantage. A low-cost Chinese open-weight model approaching the frontier is also a direct competitor. Anthropic wants more defenders to access Claude’s cyber capabilities; vetted users can already work with Claude Mythos 5.1.

That makes its call for government testing of future GLM-5.3 versions harder to separate from its business interests. The proposal could invite accusations of regulatory capture, where rules end up protecting established companies. But the evidence is not only Anthropic’s: CAISI’s assessment broadly supports its capability claims, and unlocked versions of the model are already available.

The UK AI Security Institute recently found that open models had narrowed the cyber-capability gap from six to ten months to four to seven. It also found that open models were much cheaper to run and that their safeguards were mostly ineffective. The institute warned of persistent, irreversible abuse risks, while recognizing benefits such as private deployment, customization and lower costs.

The institute had treated the capability gap as preparation time for defenders. The unresolved question was whether open models could catch up with the leap represented by Mythos Preview. Anthropic’s tests and CAISI’s assessment suggest GLM-5.3 is an answer. The harder problem is that access to the capability is no longer controlled by the labs that built the strongest models.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X