Skip to content
Methodology

GDPval

GDPval — Methodology. definition: Expert tasks benchmark — models reach 40-49% of human expert level, but require *extensive human framing* — the metric itself is underspecified · method: Experts with 14 years of experience, blind evaluation · scope: Finance, law, retail, software development

Experts averaging 14 years of experience build the tasks: realistic projects that take a human four to seven hours. Several AI models and human experts then complete them, and a third panel of experts grades the results blind, spending over an hour per question. Ethan Mollick, writing in November 2025, holds this design up against MMLU-Pro and its peers, whose public answer keys can leak into training and whose questions ("average cranial capacity of Homo erectus?") measure something nobody can name.

The scores that come out of it point in two directions. Jasmine Sun, in the NYT in April 2026, reports OpenAI's benchmark covering 44 occupations and reaching over 80% win rate compared to human professionals within months, and places it inside the Silicon Valley argument about white-collar displacement. Dan Shipper, a month later, cites 40 to 49% of human expert level, and adds the condition attached to that figure: extensive human framing. The metric itself is under-specified.

What the task-by-task results show is unevenness rather than a single frontier. Mollick reports AI beating humans on software development and financial advice while pharmacists, industrial engineers and real estate agents beat AI, and models diverging from each other: ChatGPT the better sales director, Claude the better financial adviser.

So one benchmark, spanning finance, law, retail and software development, gets read as proof that the models have arrived and as proof that someone still has to set the frame.

Type
Methodology
definition
Expert tasks benchmark — models reach 40-49% of human expert level, but require *extensive human framing* — the metric itself is underspecified
method
Experts with 14 years of experience, blind evaluation
scope
Finance, law, retail, software development
relations
5
Cited in
3 fiches

Adoption measures

Neighborhood

44 occupations humai… Jagged Frontier

→ measures

44 occupations humaines CONCEPT high confidence stable Source ↗
Jagged Frontier CONCEPT high confidence stable Source ↗

Cited in (3)