AI·Potential
AI vs Human Experts · Explained
How good AI has got

In two years, AI went from
1-in-8 to better than the expert.

One contest, run 1,320 times: an expert’s work against AI’s, judged blind by a second expert. This is the result.

AI win rate against human experts on real professional tasks, May 2024 to April 2026 A rising line chart from 12.5 percent for GPT-4o in May 2024 to 84.9 percent for GPT-5.5 in April 2026, crossing the 50 percent human-parity line in late 2025. 100% 75% 25% 0% Human-expert parity (50%) May ’24 Apr ’25 Aug ’25 Dec ’25 Apr ’26 Win-or-tie rate vs a human expert Frontier model, by release date Above the dashed line, AI beats the average human expert. GPT-4o12.5% o4-mini29.1% o335.2% GPT-539.0% Claude Opus 4.147.6% GPT-5.270.9% GPT-5.584.9%
Source: OpenAI GDPval benchmark (Sep 2025) and subsequent model-launch reports. The line shows the share of 1,320 real tasks across 44 professions where a model’s deliverable is judged at least as good as a human expert’s, blind-graded. 50% means parity with the expert; it measures one finished deliverable, not the whole job.

How they do it.

Three acts, 1,320 times
Act 01 · The expert
A real problem, solved.
A professional, 14 years into their career, takes a genuine task from their own field, with all the usual files, and produces a finished deliverable. One of 44 occupations across the nine biggest US industries.
Act 02 · The AI
The exact same task.
The model gets the identical brief and identical resources. No hints, no easier version. It produces its own deliverable for the same problem.
Act 03 · The judge
Blind verdict.
A second expert in the same field compares the two side by side, not knowing which is which, and rules on the better one. That verdict is the line on the chart.

Read it straight.

What it means, and what it doesn’t
What it means
  • Parity is behind us. Above the dashed 50% line, AI beats the average expert more often than not. The frontier crossed it in late 2025.
  • The slope is the story. Every release lands higher, months apart. Whatever today’s number is, the direction has held for two years.
  • ~100× faster, ~100× cheaper. Even where the human still wins, AI drafting plus expert review beats the expert working alone.
  • What it doesn’t
  • One deliverable, not the job. Judgement, relationships and accountability are not on the chart. Beating an expert on a task is not replacing the expert.
  • The tasks are well briefed. Every task comes with full context attached, which is exactly what most people fail to give AI. Generic in, generic out.
  • It is OpenAI’s benchmark. Independent re-runs, like the Artificial Analysis leaderboard, are the cross-check. They show the same picture.
  • The takeaway.
    The tool is no longer the constraint. On a well-briefed task, frontier AI performs at or above expert level. The constraint is how well you brief it, challenge it and review it. That is a skill, and it is yours to build.

    Sources.

    Check it yourself
    Primary
    The original benchmark: methodology, first results and the ~100× speed and cost finding.
    Paper
    Full detail on task sourcing, grader selection and per-occupation results.
    Data point
    The 70.9% score: the first clear jump past human-expert parity.
    Data point
    The 84.9% figure at the top of the chart.
    Independent
    Independent re-run across all major labs, updated as models ship. The best place for the current number.
    Context
    Independent annual review putting GDPval alongside the wider benchmark trend.
    Figures checked against these sources, July 2026. This field moves quickly; the leaderboard is the live number.
    AI·Potential
    AI vs Human Experts
    July 2026