An industrial-scale distillation of models, or subtle benchmaxxing: What developers really think of GLM-5.3
Chinese frontier model outfit Z.ai released GLM-5.3 on Friday, a model hewn from the same codebase as its predecessor GLM-5.2, The post An industrial-scale distillation of models, or subtle benchmaxxing: What developers really think of GLM-5.3 appeared first on The New Stack .
Chinese frontier model developer Z.ai unveiled GLM-5.3 on Friday, a descendant of its previous GLM-5.2 model, which has undergone enhancements through post-training optimization. The latest version is touted as significantly superior in tackling intricate coding and long-term tasks. Post-training model optimization processes, encompassing reasoning alignment, supervised fine-tuning, and Reinforcement Learning from Human Feedback (RLHF), are deemed crucial in this context.
Z.ai disclosed that they fine-tuned the model by augmenting the training environment with more diverse tasks and environments over the past month. The company further explained that these models now encompass a broader spectrum of production workflows, with tasks designed around actual engineering and research practices. For instance, a model could be given an engineering environment, equipped with access to compute clusters, storage systems, internal documentation, codebases, and experiment results, and tasked with diagnosing bottlenecks, implementing optimizations, running experiments, and delivering measurable end-to-end speedup while maintaining correctness.
The company asserts that its newly introduced Z.ai Code Bench indicates a 50% improvement in GLM-5.3's coding abilities compared to GLM-5.2. They emphasize the use of private benchmarks over public ones to minimize the risk of external contamination and provide a more accurate reflection of real-world user experience. Other public benchmark placements are also disclosed, including TerminalBench 3.0, DeepSWE, Agents’ Last Exam, AutomationBench, HLE w/ Tools: Humanity’s Last Exam (HLE), an independent academic benchmark, as well as OpenAI’s GDPVal-AA v2.
In response to these developments, industry professionals have varying opinions. Nishant Soni, co-founder of NonBioS.ai, a company focusing on autonomous long-horizon software engineering, advises caution regarding Z.ai's claims. Soni doubts the benchmarks' capacity to objectively demonstrate the alleged superior long-horizon capability of GLM-5.3 in real-world tasks, questioning the feasibility of the model achieving differentiated performance with the proposed dataset.
He believes the company may be deflecting attention from the model's true source of frontier capability, suggesting it's an industrial-scale distillation of Anthropic models. Soni also notes a striking similarity between the outputs of GLM-5.3 and Claude from NonBioS's internal testing, contrasting it with the greater diversity observed in Gemini and Grok outputs.
Sherif Higazy, founder of AI benchmarking specialist Megaton, acknowledges the value of internal evaluation systems that measure tokens and spend against tasks. He acknowledges the importance of training models within diverse and realistic environments to enhance task performance. Higazy suggests that Z.ai's approach of shifting the burden upstream into model training could yield better results.
However, he cautions that consistent and useful work from an agent requires active participation from teams to adapt their working environments and ensure machine-readable and verifiable inputs and outputs. Higazy emphasizes the need for model agnosticism to hedge against an AI market where the underlying model's performance can change rapidly.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.