Anthropic says Zhipu AI's open-weight model GLM-5.3 nearly matches Claude Mythos Preview in developing cyber exploits. The analysis comes five months after Anthropic unveiled Claude Mythos Preview and its signature capability reached competitors. Zhipu AI operates as Z.ai outside China. According to Anthropic, GLM-5.3 can build complete cyber exploits on its own just like Mythos Preview. Unlike every other model with comparable skills, GLM-5.3 shipped without effective safeguards and is available for anyone to download.
Benchmark tests show similar performance
Anthropic deliberately held Mythos Preview back, giving access only to select defenders through Project Glasswing so they could get a head start. Those defenders have since found more than 10,000 vulnerabilities in critical software according to the company. OpenAI is taking a similar approach with Daybreak. ExploitBench measures how well models exploit known bugs in Chrome's V8 engine. GLM-5.3 built a working exploit in 50 of 410 attempts on that benchmark while Mythos Preview managed 56.
Anthropic runs an internal binary exploitation benchmark based on open-source projects from Google's OSS-Fuzz. GLM-5.3 took full control of the target program in 4 percent of tasks in that test, compared to 6 percent for Mythos Preview. Older models like GLM-5.2 and Claude Opus 4.6 failed both tests, and Kimi K3 and DeepSeek V4.1-Flash barely got off zero. Frontier-level exploits now cost about as much as lunch.
Human pairing uncovers new vulnerabilities
Anthropic paired GLM-5.3 with a human expert. The model found several previously unknown vulnerabilities in the JavaScript engine of a widely used browser within a single day and with little human attention. It chained them into a web page that can read any file on a visitor's computer and pulled a private SSH key in the test. Anthropic reported the vulnerabilities to the browser's developers, while other findings in drivers and device firmware remain under review.

Anthropic used the smaller GLM-5.3-Flash to test how quickly a freshly disclosed vulnerability turns into a working attack. The model combined a recently disclosed Chrome bug with another known vulnerability and built a reliable attack with little guidance. The attack bypassed an extra security feature built into the processor. The job took 20 minutes of human attention and eight hours of model time, costing $20.40 at Zhipu's API prices.
US agency reaches similar conclusions
The US agency CAISI reached similar conclusions in its own assessment. CAISI calls GLM-5.3 the most cyber-capable open-weight model to date and puts it about four months behind the best US models. That comparison comes with caveats because CAISI tested the US models with their cyber safeguards turned off, and the top tier includes models that only vetted users can access.
Open weights make safeguards easy to remove. In an Anthropic simulation, GLM-5.3 refused openly malicious attack commands. When the request was dressed up as a red-team exercise, the model tried to connect to the target system in 64 percent of runs. Pre-filled reasoning steps raised that number to 92 percent. After abliteration, a technique stripping refusal behavior out of open weights, it reached 100 percent. The simulation does not execute code so it cannot show whether an attack would actually have succeeded, while protected Claude models stayed at zero.
Read nextOpenAI Nears $70 Billion Annualized Revenue Rate Amid Rapid GrowthAnthropic spent about 2,200 GPU hours costing roughly $4,400 on its first abliteration test, estimating an experienced team could do it for around $1,200. The refusal rate for harmful requests fell from over 90 percent to between 2 and 12 percent, while scores on science and cyber tests barely moved. Anthropic says several developers had already released unlocked versions within days of the model's launch.
Anthropic concludes that state and non-state actors will likely use models like GLM-5.3 to cause real harm. The company points to its own reports and those from other US labs documenting attackers who already use AI, arguing governments should test capable models and defenders need tools at least as good as their adversaries. The analysis also serves Anthropic's business since the company does not release model weights and presents that choice as a security advantage while competing with cheap Chinese open-weight models.



