Skip to content
Technology

GPT-5

GPT-5 — Technology. category: OpenAI language model · even-handedness score: 89% · publisher: OpenAI · evaluation: METR concludes catastrophic risk is low but trends are concerning

GPT-5 generated the first commit of an empty repository at OpenAI, run through Codex CLI and guided by templates. That repository became, over five months, an internal beta product of roughly one million lines with zero lines written by hand: three engineers, later seven, whose Codex agents opened, evaluated and merged about 1500 pull requests, an average of 3.5 per engineer per day. OpenAI's account of that work locates the limit outside the model. The bottleneck on agent performance is usually the design of the environment (dependency-layer rules, custom linters, versioned architecture documents) rather than the intelligence of the model driving it.

The measurements point the other way. Anthropic's paired-prompts evaluation, 1,350 prompt pairs across 150 topics, scored GPT-5 at 89% on political even-handedness, behind Gemini 2.5 Pro (97%), Grok 4 (96%), Claude Opus 4.1 (95%) and Claude Sonnet 4.5 (94%), ahead of Llama 4 (66%). GPT-5 sat on both sides of that study: Anthropic used it to check its own grader, and its scores agreed with Claude Sonnet 4.5's on 92% of the sampled pairs, correlation 0.86.

In Microsoft's Magentic Marketplace, 100 simulated customers against 300 businesses, GPT-5 and Gemini 2.5 Flash were the exceptions to the pattern of agents settling for a handful of sellers. Against six manipulation strategies, only Claude Sonnet 4 resisted every attempt. METR's assessment of GPT-5: low catastrophic risk, concerning trends.

Type
Technology
category
OpenAI language model
even-handedness score
89%
publisher
OpenAI
evaluation
METR concludes catastrophic risk is low but trends are concerning
relations
3
Cited in
4 fiches

Neighborhood

← measures

METR ORGANISATION high confidence stable Source ↗

← uses

Codex TECHNOLOGIE high confidence evolving Source ↗

Cited in (4)