🔍 Read the full analysis: How The Referee Shortage Shapes AI’s Economics on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
The source report points to a widening gap between AI-generated work and the human capacity to verify it, citing examples in mathematics, software and contract review. The figures suggest review may become a constraint on how much AI output organisations can safely use, though several cited studies have limitations and the report’s broader economic conclusions remain analysis, not settled fact.
AI systems are producing work faster than human experts can verify it, according to an analysis citing a new batch of mathematical manuscripts, software pull-request studies and an AI contract-review evaluation. The report argues that this imbalance could make qualified reviewers—not generation capacity—a constraint on how much AI output organisations can use.
The analysis says OpenAI published 722 mathematical manuscripts this week, with an average result taking about three hours of compute to produce. The manuscripts came from work on roughly 4,000 problems and were grouped into 372 families, according to the source. Some results have been formally checked using Lean, a proof-assistant system; OpenAI cautioned that some results not formalized in that way could have issues.
The report contrasts that output with the response to an earlier result from the same programme: a counterexample to an Erdős conjecture that received careful verification from five leading mathematicians. That example illustrates a distinction between producing a candidate result and establishing that it is correct, significant and relevant. The report describes the former as increasingly abundant and the latter as dependent on limited expert time.
Software data cited in the analysis points to a similar tension. Faros AI reported that teams in high-AI-adoption periods merged 98% more pull requests, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes took 4.6 times longer to begin review and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found that 61% of AI-agent pull requests received no human review before being merged or closed. These figures measure different things and should not be treated as directly comparable.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Shapes AI’s Value
If generating drafts, code or research becomes cheap while checking remains time-consuming, the practical value of AI depends on whether organisations can review and adopt its output responsibly. More generated material does not automatically mean more usable work: teams still need people able to identify errors, judge whether a result addresses the right problem and accept responsibility for decisions.
The analysis predicts a potential premium for experienced reviewers, including senior engineers, auditors, specialist lawyers, scientists and safety assessors. That is an economic interpretation rather than a measured labour-market outcome. The data cited does, however, raise a concrete operational concern: if review queues lengthen or work is merged without review, organisations may either slow down adoption or accept higher risk.
The report also points to a workforce challenge. Junior staff often develop judgement by drafting code, contracts or research and receiving feedback. If AI takes over much of that early work, employers may need deliberate ways to provide practice and supervision. Otherwise, the same tools that reduce routine production could weaken the pipeline of people qualified to check more complex work.
As an affiliate, we earn on qualifying purchases.
Evidence Across Three Workflows
The report brings together examples from mathematics, software and contract work, but they do not establish one uniform measure of an AI-driven review shortage. In mathematics, formal proof tools can verify that a proof follows from its stated premises. They do not by themselves determine whether the theorem is important or whether the formal statement captures the intended question.
In software, automated tests can check specified behaviour, but their coverage depends on what developers thought to test. The source notes that Faros AI and LinearB sell tools related to software development or review, so their findings warrant care. The report says the direction of the results is consistent across sources, but differences in samples, methods and definitions make the exact figures difficult to combine.
The contracting example is an evaluation of GPT-6 Astra, developed through OpenAI’s partnership with contract-software company Ironclad. According to the source, Astra met an average of 55% of evaluation criteria across 11 tasks, an improvement over the prior model. That score indicates progress on the evaluation, but it also leaves criteria unmet; it does not, on its own, establish how the system performs across all legal workflows or whether its output is ready for use without professional review.
mathematical proof assistant software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Much Review Is Missing?
The cited evidence does not establish the size of a general shortage of qualified reviewers or prove that AI adoption caused every reported change. The software findings come from different datasets and measures; several sources have commercial interests in review tools, and the source provides limited detail on study methods. The peer-reviewed study’s 61% figure concerns AI-agent pull requests in its sample, not all software development.
It is also unclear how often unreviewed work caused defects or other harm, and whether review time will fall as tools and workflows improve. The contract evaluation’s 55% average across 11 tasks does not show which criteria were missed or how consequential those gaps were. The source’s economic prediction that reviewers will command a premium remains a forecast, not a reported wage trend.
As an affiliate, we earn on qualifying purchases.
Track Review and Training Practices
The next useful evidence will be whether organisations can expand review capacity without lowering standards. That means tracking review delays, defect rates and review coverage alongside measures of how much AI-generated work is produced or merged. The cited figures alone do not show whether teams are changing staffing, requiring different levels of review or using automated checks to reduce routine reviewer workload.
Employers and professional bodies will also need to watch how junior staff gain experience. The report argues for protecting training routes into expert judgement, but it does not identify a single established solution. Whether AI systems take on more verification—and whether institutions accept that verification as sufficient—remains an open question across the fields discussed.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development described?
The source report argues that AI-generated output is growing faster than expert review capacity, drawing examples from mathematics, software and contract evaluation. It is an analysis of a trend, not an announcement of a single new policy or industry-wide measurement.
Did OpenAI publish 722 verified mathematical proofs?
The source says OpenAI published 722 mathematical manuscripts produced from work on roughly 4,000 problems. It reports that some results were formally checked in Lean and quotes OpenAI warning that some unformalized results could have issues. The 722 should not be described as 722 fully verified proofs.
What does the software data say about review?
The cited sources report longer review delays and, in one study, a high share of AI-agent pull requests receiving no human review. The measures come from separate samples and methods, so they should not be combined into a single estimate for all software teams.
Does the report prove that AI reviewers will become more valuable?
No. A growing need for expert review is the report’s economic interpretation. The cited material does not provide direct evidence of rising reviewer pay or establish how labour demand will change.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
