In two years, AI went from 1-in-8 to better than the expert.
One contest, run 1,320 times: an expert’s work against AI’s, judged blind by a second expert. This is the result.
Source: OpenAI GDPval benchmark (Sep 2025) and subsequent model-launch reports. The line shows the share of 1,320 real tasks across 44 professions where a model’s deliverable is judged at least as good as a human expert’s, blind-graded. 50% means parity with the expert; it measures one finished deliverable, not the whole job.
How they do it.
Three acts, 1,320 times
Act 01 · The expert
A real problem, solved.
A professional, 14 years into their career, takes a genuine task from their own field, with all the usual files, and produces a finished deliverable. One of 44 occupations across the nine biggest US industries.
Act 02 · The AI
The exact same task.
The model gets the identical brief and identical resources. No hints, no easier version. It produces its own deliverable for the same problem.
Act 03 · The judge
Blind verdict.
A second expert in the same field compares the two side by side, not knowing which is which, and rules on the better one. That verdict is the line on the chart.
Read it straight.
What it means, and what it doesn’t
What it means
Parity is behind us. Above the dashed 50% line, AI beats the average expert more often than not. The frontier crossed it in late 2025.
The slope is the story. Every release lands higher, months apart. Whatever today’s number is, the direction has held for two years.
~100× faster, ~100× cheaper. Even where the human still wins, AI drafting plus expert review beats the expert working alone.
What it doesn’t
One deliverable, not the job. Judgement, relationships and accountability are not on the chart. Beating an expert on a task is not replacing the expert.
The tasks are well briefed. Every task comes with full context attached, which is exactly what most people fail to give AI. Generic in, generic out.
It is OpenAI’s benchmark. Independent re-runs, like the Artificial Analysis leaderboard, are the cross-check. They show the same picture.
The takeaway.
The tool is no longer the constraint. On a well-briefed task, frontier AI performs at or above expert level. The constraint is how well you brief it, challenge it and review it. That is a skill, and it is yours to build.