Transcript · Brainqub3

AI Benchmarks vs Real Work (GDPVal Explained)

11:401,735 words9 min read
0:00

Speaker A

We ought to be more skeptical about how we evaluate AI, because the benchmarks we've been using are failing us. Every few weeks, another frontier model drops, another leaderboard gets topped and the marketing teams declare victory. But here's the problem. These exam style benchmarks have become a game unto themselves. The scores don't translate to real world performance the way you'd expect. OpenAI recently released research on something called GDPVAL. That's a benchmark built from actual professional work, not academic puzzles. And what it reveals challenges a lot of the narrative around AI capability. I'll walk through why the current benchmark race is creating perverse incentives. What GDPVAL actually tells us about model performance on real tasks and what this means for how we should think about AI in professional work. Whether you're a user or engineer trying to pick the right tools for your workflow and an executive making AI strategy decisions, or an investor trying to separate signal from noise, this matters.

0:58

To understand why GDP Vial matters, you need to understand how it was built. Take Humanities Last Exam One of the most talked about benchmarks for frontier models. It was created by academics, researchers and professors who submitted exam style questions. Nearly 1,000 contributors across 500 institutions. But here's the key Design choice questions were specifically selected because they stumped frontier AI models. Questions that models could answer were filtered out during the review. There was a $500,000 prize pool incentivizing the hardest possible questions. And grading is automated multiple choice or exact match. Short answers. No ambiguity, no subjectivity. It's an examination, a very hard exam, but it's still just an exam. Gdpval takes a fundamentally different approach. It was built by recruiting industry professionals, not academics designing test questions Practitioners doing professional work outside of academia. The average contributor had 14 years of professional experience. They came from places like Goldman Sachs, JP Morgan, IBM, Boeing, Johnson and Johnson, the Department of Justice, and hundreds of other organizations.

2:12

And here's what matters. These experts didn't write exam questions. They contributed tasks based on work they actually did in their jobs. Real deliverables A financial analyst submitting an actual competitor landscape analysis. A registered nurse submitting an actual skin lesion consultation report. A manufacturing engineer submitting an actual 3D CAD model design. Each task was then validated through multiple rounds of review. First by generalist reviewers, then by occupation specific experts who checked whether the task was genuinely representative of real work in that field. Then through an iterative feedback loop until the task met the quality standards required. On average, each task received five human reviews. The result is tasks that take an average of around seven Hours for a professional to complete the gold subset averages to closer to nine and a half hours. Some tasks span multiple weeks. This isn't about stomping the models. It's about measuring whether models can do the actual work that professionals do.

3:17

And the grading is different too. Blinded pairwise comparisons where human industry experts look at the two deliverables, one from the human and one from a model, without knowing which is which. Then they judge which is better, considering correctness, but also structure, style, format, aesthetics and practical usability. The same subjective factors that matter in real professional work. So when we talk about GDP vowel versus Humanities Last Exam, we're not just comparing different questions. We're comparing different philosophies of what it means to evaluate AI. One asks, can the model pass the hardest academic test? The other asks, can the model do the actual job? There are four things that really matter when we're talking about AI performance on real work. First, your benchmark type. Are we measuring exam style, reasoning or actual professional output? These are fundamentally different things. Second, human collaboration. Does the workflow assume full automation or human oversight and review?

4:22

This dramatically changes the economics. Third, context engineering. How much effort goes into prompt design, scaffolding and providing the right information to the model? Fourth, task coverage. How much of the actual long tail of real work does any given benchmark represent? These four axis explain why your experience with a model might diverge wildly from what the leaderboards suggest. Let's look at how different players sit on this board. The Frontier Labs, OpenAI, Anthropic Google. They're locked in a benchmark arms race. Today they're optimizing heavily for popular benchmarks like Humanities Last Exam, RKGI2 and Suitebench. These benchmarks have become table sticks. No Frontier Lab can afford to fail them publicly. The reputational damage would be severe. But this creates a conflict of interest. Training to top exam style benchmarks may not translate to real world utility and may actually obfuscate what these models can genuinely do. If we shift towards real world benchmarks like GDPVAL, the picture changes.

5:27

On GDPVAL's gold subset, Claude Opus 4.1 was judged better than or equal to human experts in 47.6% of pairwise comparisons. G5's strict win rate where the model was explicitly preferred over the human was 39%. That's progress, but not decisive superiority. The upside for Labs is that real work benchmarks could better demonstrate genuine value to enterprises. The risk is that the numbers look less impressive than the marketing friendly exam scores for enterprises and practitioners. The current position is One of confusion. You're trying to pick models based on benchmarks that don't reflect your actual workflows and the story is told in the divergence. When Opus 4.5 and GPT 5.1 dropped, I found 5.1 to be a regression and moved my workflow to Claude. Others found the opposite. How is that possible if both models beat their predecessors on the benchmarks? If GDP valve style evaluation becomes standard, enterprises get clearer signal on what actually works.

6:35

The upside is better purchasing decisions. The risk is discovering that the productivity gains require more engineering effort than expected. Fast forward a couple of years and a few paths stand out to me. In the first scenario we'll say the benchmark theater continues. Labs keep optimizing for exam style benchmarks because that's what drives headlines and enterprise procurement. Real world performance continues to diverge from benchmark scores. Practitioners remain frustrated and the gap between marketing and reality widens. In the second scenario, real work benchmarks win. Something like GDPVAL becomes the industry standard labs shifts resources towards performance on actual professional tasks. This surfaces more honest capability assessments but potentially slower looking progress because real work is harder than exams. In the third scenario, hybrid equilibrium, we end up with a tiered system exam benchmarks for raw capability, raw work benchmarks for applied performance. Sophisticated buyers learn to read both while the marketing race continues at the exam level.

7:41

Whatever scenario plays out, I think a few principles will hold context. Engineering is not optional, but GDPVAL research showed that when prompts were shortened to 42% of the original length, deliberately omitting guidance on approach and formatting, GPT5 performed worse. The model struggled to figure out requisite context on its own. Human oversight remains economically necessary. The GDPVAL research models a realistic workflow. Try the model, review the output and if it's still unsatisfactory, do it yourself. Under this setup, GPT5 offers roughly 1.1 to 1.4 times speed up and 1.2 to 1.6 times cost reduction. Weaker models like GPT4O actually came out slower and more expensive than just letting the expert do the work. And these figures don't account for serious mistakes. When GPT5 lost to the human expert, about 29% of those failures were rated bad or catastrophic, while roughly 3% rated strictly catastrophic. The long tail of serious errors is non trivial.

8:44

The long tail is longer than any Benchmark. GDPVAL covers 44 occupations across the top nine GDP sectors in the US. It represents jobs earning around $3 trillion annually, yet still only covers 14% of the O NET task types. Onet by the way, is the US Department of Labor's database of occupational information. Even this ambitious benchmark is sampling a fraction of real professional work. If you're an engineer or technical practitioner, expect that model selection will remain empirical. Benchmarks give signal, but your specific workflow is the real test. Your edge is in context and and prompt engineering. The GDPVAL research showed that adding proper scaffolding, forcing correctness checks, rendering outputs as images for visual inspection, and using best of n sampling improved win rates. That's meaningful margin that comes from technical expertise, not model swaps. Focus on building robust evaluation for your specific use cases. Understand that different models fail differently.

9:46

The GDPVAL research found that Claude and Gemini lost most of them on instruction following, while GPT5 lost mainly on formatting errors. Pick accordingly based on what matters for your work. If you're an executive making AI strategy decisions, expect that the productivity gains from AI are real but modest. Once you include human review, we're talking tens of percent improvement, not orders of magnitude. Your edge is in building the right human AI workflows rather than chasing full automation. Focus on investing in the technical leadership to do context engineering properly. Don't outsource model selection to benchmark leaderboards and build in the review and oversight processes that make the economics actually work. If you're an investor, expect continued benchmark theater at the marketing level with real work performance being the actual differentiator for enterprise value. Your edge is in identifying which companies understand this distinction and are building accordingly.

10:47

Focus on looking past the headline benchmark numbers, ask how companies are measuring real world performance, and be skeptical of automation narratives that don't account for the review and oversight costs. To recap my main structural argument, exam style benchmarks reward exam style training. Real work has a long tail that no benchmark can fully capture, and the economic gains from AI depend heavily on human oversight and context engineering. We're early in the shift towards more honest evaluation. If you focus on building genuine expertise in working with these models on the prompt engineering and the scaffolding that actually moves the needle, you'll be on the right side of it. Assuming the agentic era will be built on large language models, remember it will be built by humans who know how to make them useful. The agents are partners for human experts, not replacements.