- thekb.eu
- Knowledge graph
- Organization
- Artificial Analysis
Artificial Analysis
Artificial Analysis — Organization. role: Only third-party evaluator cited in the post: produces the GDPval-AA v2 and FrontierSWE scores in the comparison table, and supplies the per-model throughput figures used to normalize ExploitGym time budgets
Artificial Analysis published a comparison on 22 June 2026 placing Z.ai's GLM-5.2 at 1524 Elo on GDPval-AA, third overall and first among open weights models, behind Claude Fable 5 (1783) and Claude Opus 4.8 (1615) and level with GPT-5.5 at its xhigh setting (1509). The margin inside the open camp was wide: MiniMax-M3, the next open model, scored 1408.
The method is what makes the number legible. GDPval-AA scores long-horizon, multi-turn tasks written as professional exercises (a store supervisor's daily task list, an IEC technical document). Artificial Analysis gave identical briefs to GLM-5.2 and three proprietary frontier models, then rendered every deliverable exactly as produced. GLM-5.2 averaged about 31 turns per task across 1,999 matches. The verdict Artificial Analysis attached to the result: an open weights model at that price competing with the proprietary frontier on genuinely useful agentic work is "a real step for open models".
Two months later, in Z.ai's unsigned GLM-5.3 announcement of 14 August 2026, Artificial Analysis is the only outside evaluator present. The GDPval-AA v2 and FrontierSWE figures in the comparison table are the sole scores attributed to a third party, and its per-model throughput numbers are what normalize the time budgets for ExploitGym. The rest, including the Z.ai Code Bench results carrying the "most capable open-weights model for coding" claim, is in-house and private. That leaves one external anchor holding the credibility of a release whose cyber results Z.ai itself calls emergent.
- Type
- Organization
- role
- Only third-party evaluator cited in the post: produces the GDPval-AA v2 and FrontierSWE scores in the comparison table, and supplies the per-model throughput figures used to normalize ExploitGym time budgets
- relations
- 3
- Cited in
- 2 fiches
Neighborhood
→ publishes