American artificial intelligence laboratory Anthropic studied the capabilities of the Chinese model Z.ai GLM-5.3 in the field of cybersecurity and came to disappointing conclusions in conclusion: The system has broad functionality and could be exploited by hackers because its protection against abuse is not effective enough.
Image source: anthropoic.com
When it comes to finding vulnerabilities, Z.ai GLM-5.3, which is open and free to everyone, is almost as good as the closed Anthropic Claude Mythos Preview, which can only be used by a limited number of organizations (participants of the Glasswing project). Experts at the Center for Artificial Intelligence Standards and Innovation (CAISI) at the National Institute of Standards and Technology (NIST) tested the capabilities of Z.ai GLM-5.3 and concluded: “The most powerful open model with cyber attack capabilities released to date”By all metrics, it’s four months behind the U.S.’s advanced level – Anthropic’s own test results are broadly consistent with this assessment. The difference is that the free Chinese mode is available to an unlimited number of people.

In an initial phase, Anthropic tested GLM-5.3 using the ExploitBench toolkit, which helps evaluate the model’s ability to exploit known vulnerabilities in the V8 engine used in the Google Chrome browser. Here, special attention is paid to the model’s ability to perform a full range of work to create ready-made vulnerabilities. GLM-5.3 successfully produced 50 out of 410 tests; by comparison, the Claude Mythos Preview scored 56 out of 410. In the Binary Exploitation test, Anthropic engineers tested the model’s ability to find and exploit vulnerabilities in open source projects participating in the Google OSS-Fuzz project. The highest score in the test is awarded to a successful complete interception of control of program execution, a full-fledged hacking attack. The performance of multiple models was evaluated on 100 random tasks: GLM-5.3 succeeded in 4% of cases, Claude Mythos Preview – succeeded in 6% of cases. The Chinese model was inferior to the American model, but there is another important point: the earlier models, including the Claude Opus 4.6 and GLM-5.2, failed to cope with any tasks.

Further testing was conducted under near “combat” conditions and with human participation. In the first experiment, GLM-5.3 was used on an isolated machine – the object was a native Linux version of a popular browser. The model discovered several previously unknown vulnerabilities in the JavaScript engine and combined them into a working vulnerability: when opening a specially prepared web page, a hypothetical attacker could access arbitrary files on the user’s computer. In the second process, a more compact and less powerful version of the same model, GLM-5.3-Flash, was used – it was informed of two known vulnerabilities in the Google Chrome browser and created a working chain almost without human intervention, in addition to bypassing the hardware protection mechanism PAC – Pointer Authentication on the ARM64 architecture. This takes an expert 20 minutes and the model itself 8 hours – costing $20.40 at current prices if accessed through the Z.ai platform via API.
In its original form, GLM-5.3 was released with some built-in protections against abuse: if the user intentionally gave unacceptable instructions, the model would respond with a rejection. But these mechanisms can be bypassed relatively easily. The most effective approach is elimination: a model with open weights that can be reconfigured by the user to deactivate the faulty functionality with little or no impact on system functionality. The GLM-5.3 version that went through this process was released a few days after the original model. Human engineers perform this process in-house. For the full-size GLM-5.3 model, this requires 2,200 GPU hours and costs $4,400; for the compact GLM-5.3-Flash – 600 hours. Since then, the failure rate in professional benchmarks JailbreakBench and HarmBench dropped from 90% to 3% and 2% respectively, and in StrongREJECT to 12%. Meanwhile, in the GPQA-Diamond test, which assesses scientific ability, the standard and abolished versions of the model showed identical results, while in the CyberGym test a drop of several percentage points was recorded.

In fact, when executed in an isolated environment, the deprecated GLM-5.3 executed 100% of obviously malicious queries. But that’s not the only way to disable defense mechanisms. Misleading requests that told the model it was participating in a training exercise caused the model to follow malicious instructions 64% of the time. Another approach – pre-populating inference tokens, where the model appears to have already considered the request and decided to proceed, raised this number to 92%.
As a result, Anthropic experts concluded that Z.ai GLM-5.3 provides cybercriminals with virtually unlimited capabilities to hunt for and exploit vulnerabilities. American AI Labs acknowledges that this and other similar models will be used in practice to cause real harm.
If you find an error, select it with your mouse and press CTRL+ENTER.









