A recent study raises concerns about the unreliability of artificial intelligence in predicting the impact on various job sectors. Researchers from Northwestern University and American University, in a working paper from the National Bureau of Economic Research, discovered significant discrepancies in assessments made by different advanced AI systems when evaluating their effect on employment.
The study evaluated four cutting-edge AI models – GPT-4, ChatGPT-5, Gemini 2.5, and Claude 4.5 – using a standardized framework to analyze nearly 19,000 job tasks. Results revealed substantial disparities among the models, with mean exposure scores varying from 0.14 (GPT-4 and Gemini) to 0.51 (Claude), representing a 3.6-fold difference. Pairwise agreement between models dropped to as low as 57%, indicating only a “fair” level of consensus.
The most significant discrepancies were observed in occupations blending cognitive and physical responsibilities, such as management, teaching, and sales. Management roles, for instance, displayed exposure scores ranging from approximately 0.08 (Gemini) to 0.83 (Claude). Similarly, computer and mathematical professions ranged from 0.42 (Gemini) to 0.95 (Claude). Occupations like educational instruction, life sciences, and sales exhibited exposure variations of 0.30 or higher among the models. While AI models generally agreed that physical labor jobs like construction were secure, they differed in their assessments of coding positions and white-collar roles.
These inconsistencies led to varied real-world implications. At the county level, Claude 4.5 indicated a statistically significant negative correlation between AI exposure and employment levels. Conversely, GPT-4, ChatGPT-5, and Gemini 2.5 did not find any significant impact, with Gemini even suggesting a positive – albeit insignificant – association. At an individual level, all models pointed to a negative effect, but the magnitude differed significantly, with Gemini showing the most substantial impact, 2.4 times greater than the initial GPT-4 estimation.
The researchers highlighted the critical role of selecting the AI model tasked with rating the job tasks, emphasizing the circular nature of relying on AI to assess its own influence. They advised policymakers, economists, and labor agencies to approach current job exposure assessments cautiously, advocating for metrics based on tangible AI usage data instead.
