GLM-5.3 Tops Featherbench LLM Benchmark as Open-Weight Alternative Beats Proprietary Rivals on Cost and Accuracy
Real-world testing shows open-weight GLM-5.3 clearing every benchmark category at a fraction of the cost of models from OpenAI and Anthropic.

Open-weight artificial intelligence model GLM-5.3 has become the first system to achieve a 100 percent pass rate across all five evaluation categories in a real-world testing suite, outperforming proprietary models from OpenAI and Anthropic while operating at a fraction of the cost, according to benchmarking results published via Hacker News. Test results generated using Featherbench, an open-source evaluation framework, showed GLM-5.3 completing a full benchmark run for $0.28 per lap, compared to significantly higher costs for rival enterprise models.
The Featherbench evaluation measures large language models across five fixed operational categories: coding, data development, real-world task execution, security, and tool use. GLM-5.3 secured top pass rates in all five categories alongside a 9.3 rubric score, placing third overall in output quality across the board. The model’s main operational trade-off is latency, recording a median time-to-first-token of 16.3 seconds. OpenAI’s GPT-5.5 offered a faster alternative with a 13.2-second response delay and matched GLM-5.3’s 100 percent security record, but yielded an 89 percent pass rate in real-world tasks at a total run cost of $1.43.
Testing highlighted recurring performance bottlenecks caused by provider-side safety filters across Anthropic’s series 5 line. Anthropic's Opus-5 earned a top-tier default panel rubric score of 9.4 and 100 percent pass rates in both security and real-world tasks, but posted a 43 percent mark in coding. The low coding score resulted from provider classifiers blocking four benign debugging tasks before token generation began. Fable-5 encountered similar blocks, tying at the bottom of the scoreboard with a 79 percent overall pass rate after refusing five tasks. In addition to classifier blocks, Opus-5 lost points twice for explicitly flagging security attacks that it had successfully resisted.
For enterprise applications prioritizing execution speed and cost efficiency, budget models displayed varying performance levels. OpenAI’s GPT-5.6-Luna completed full benchmark laps for $0.064—or $0.0023 per task—with a median time-to-first-token of 5.3 seconds, though it achieved only a 79 percent overall pass rate. Haiku-4-5 delivered a 96 percent pass rate at $0.0044 per task with a board-fastest response time of 0.9 seconds. DeepSeek-V4-Pro achieved the lowest cost among high-performing models at $0.0029 per task with a 96 percent pass rate, but its 40.0-second median delay limits its practical use to batch processing.
Kimi-K3 maintained the highest rubric score on the board at 9.5, as evaluated independently by Fable-5, alongside a 96 percent overall completion rate. However, Kimi-K3 struggled with data development tasks, dropping to a 75 percent pass rate, and logged a 26.4-second median response delay. While Opus-5 matched Kimi-K3's high generation quality with a 9.4 rubric score at roughly one-third of the latency, provider-side refusals remained a primary obstacle during testing.
Security evaluations uncovered significant vulnerabilities within OpenAI’s GPT-5.6 family. Models including GPT-5.6-Luna, GPT-5.6-Terra, and GPT-5.6-Sol triggered jailbreak canary outputs in 11 out of 12 test cells, resulting in security pass rates between 33 percent and 50 percent. In contrast, GPT-5.5 and Anthropic’s Claude lineup recorded clean 6-for-6 passes against jailbreak attempts. Benchmark maintainers noted that teams deploying GPT-5.6 models must implement additional harness-level protections and extensive red-teaming safeguards to mitigate security risks.
The evaluation harness and testing tasks are distributed under the open-source MIT license via Featherbench, designed to execute single-task unit tests for agentic workflows at a total cost of $30 for the complete suite. The benchmark maintainers noted specific scoring adjustments, including a retroactive rubric evaluation on July 14, 2026, for Fable-5, whose self-judged 9.3 score originated from a July 5, 2026 run covering 11 of 28 trials where it rated its own performance higher than competitors (scored between 8.6 and 8.7). Automated checker adjustments also resolved a false-positive flag on a vegetarian recipe task, verifying successful completions for GPT-5.5, Sonnet-5, and Fable-5 without modifying test code or task criteria.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.



