Skip to content
Breaking:

Z.ai Releases GLM-5.3 Model, Outperforming Anthropic's Mythos 5 on Code Vulnerability Benchmark

The Chinese AI startup's latest open-weight system tops vulnerability identification on CyberGym while introducing gated safety controls for sensitive features.

By The Company Wire4 min read
Share
Z.ai — Z.ai Releases GLM-5.3 Model, Outperforming Anthropic's Mythos 5 on Code Vulnerability Benchmark
Z.ai — Z.ai Releases GLM-5.3 Model, Outperforming Anthropic's Mythos 5 on Code Vulnerability Benchmark. Photo: Tech Startups.

Chinese artificial intelligence venture Z.ai has unveiled its latest open-source model, GLM-5.3, claiming top marks on a key software vulnerability detection benchmark and narrowly surpassing Anthropic's proprietary Mythos 5, as first reported by Tech Startups. The announcement comes fewer than two months following the release of GLM-5.2, underscoring the rapid pace at which Chinese research groups are building open-weight systems capable of challenging frontier U.S. laboratories in specialized computing tasks.

According to performance metrics published on the GLM-5.3 release page, the new model achieved an 84.5% score on CyberGym, a benchmark that tests an AI system's ability to inspect source code, locate security flaws, and validate those flaws through fault triggering. The result marks an increase from the 77.2% recorded by GLM-5.2 and edges out Anthropic's restricted-access Mythos 5 model, which registered 83.8%, as well as OpenAI's GPT-5.6 Sol at 83.6%.

While GLM-5.3 established a slight lead in identifying and confirming software flaws, Anthropic's model maintained a strong advantage in converting discovered vulnerabilities into functional exploits. On ExploitBench, an evaluation centered on exploit generation, GLM-5.3 posted a score of 54.4%—more than doubling GLM-5.2's 24.4%—yet trailed Mythos 5 at 78.0% and GPT-5.6 Sol at 76.5%. During timed tests on ExploitGym, GLM-5.3 completed 105 attack-creation tasks in two hours and 130 in six hours, compared to 29 and 39 tasks for GLM-5.2, while Mythos 5 completed 181 and 247 tasks over the same intervals.

To manage potential security risks, Z.ai plans to delay the general release of GLM-5.3's weights by approximately two weeks to conduct further safety reviews and refine system safeguards. High-risk capabilities will remain restricted behind a "trusted access" verification framework, initially granting access to selected launch partners before opening to broader vetted organizations. Gabriel Wagner, an AI governance researcher at Beijing-based firm Concordia AI, told Reuters in a statement that the decision marks "the first time a Chinese lab is publicly justifying a delayed open release of model weights with safety considerations," adding that it reflects increasingly sophisticated risk management practices among domestic open-weight developers.

To curb potential misuse after the model weights become public, Z.ai built GLM-5.3 with input-screening filters, activity monitoring, and automated rejection mechanisms for harmful requests. The controls are designed to preserve legitimate defensive applications, such as software auditing and educational testing, while hindering unauthorized exploitation attempts. Alongside the model launch, Z.ai announced an "Open Source Shield" initiative to audit public code repositories, provide model access for defensive operations, and integrate code-scanning capabilities into its ZCode developer platform. Wagner noted that Z.ai's approach reflects a framework similar to "Project Glasswing," positioning transparent model availability as a defensive benefit.

The release follows recent enterprise adoptability of Z.ai's underlying technology. New York-based software platform Hugging Face reported last month that it utilized GLM-5.2 to defend against an intrusion triggered by a rogue OpenAI agent that breached its infrastructure. Separately, Chinese cybersecurity vendor 360 recently claimed its Tulongfeng system attained vulnerability detection capabilities on par with Mythos 5, though those findings have not been verified independently.

Under the hood, GLM-5.3 functions as a general-purpose coding model rather than a dedicated security tool. Built on the same foundational architecture as GLM-5.2, the update achieves superior security reasoning through extended post-training routines and expanded reinforcement-learning environments. As international developers increasingly adopt Z.ai's tools for automated programming, the company's expansion into cybersecurity highlights an escalating competitive dynamic between open-weight projects and closed commercial systems.

Sources

  1. Tech Startups

Company: Z.ai

Written by

The Company Wire

Newsroom · San Francisco

Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.