Skip to content
Technology

Claude Sonnet 4.5

Claude Sonnet 4.5 — Technology. category: Anthropic LLM · even-handedness score: 94%

Claude Sonnet 4.5 scored 94% on Anthropic's even-handedness measure published in November 2025, fourth in a field of six behind Gemini 2.5 Pro (97%), Grok 4 (96%) and Claude Opus 4.1 (95%), ahead of GPT-5 (89%) and Llama 4 (66%). The more telling fact is that the same model did the scoring. Anthropic used Claude Sonnet 4.5 as the automated grader across 1,350 prompt pairs, 9 task types and 150 topics of American political discourse, with a subsample re-graded by Claude Opus 4.1 and GPT-5 as a validity check. Agreement reached 94% with Opus 4.1 and 92% with GPT-5, against 85% between human raters.

Its own profile is uneven. It declines rarely (3% refusal, second only to Grok 4's near zero) but raises counter-arguments least often of the four models measured on that axis: 28%, against Opus 4.1's 46%. Ready to engage, less inclined to qualify.

The second appearance is practical rather than evaluative. Ethan Mollick handed it a full economics paper and its replication data archive. The model read the article, sorted the files, converted STATA code to Python and verified the findings including the complex interactions, in minutes. Mollick spot-checked the output, then had GPT-5 Pro replicate the replication.

Grader and subject, verifier and thing verified. Anthropic open-sourced the evaluation on GitHub and listed grader dependence among the eight limits it acknowledged.

Type
Technology
category
Anthropic LLM
even-handedness score
94%
relations
4
Cited in
2 fiches

Neighborhood

Ethan Mollick réplication d'articl… réponses autres modè…

← uses

Ethan Mollick PERSONNE high confidence stable Source ↗

→ enables

réplication d'articles de recherche scientifique CONCEPT high confidence stable Source ↗

→ measures

réponses autres modèles CONCEPT high confidence evolving Source ↗

Cited in (2)