Newsroom
Jobs

New Research Shows AI Can Complement Workers, Boost Wages, and Expand Opportunity

A
Alice Thornton
May 13, 202614 min readUpdated August 18, 2026
Share:
New Research Shows AI Can Complement Workers, Boost Wages, and Expand Opportunity

New Research Shows AI Can Complement Workers, Boost Wages, and Expand Opportunity

TL;DR

A growing body of evidence from MIT, Stanford, Goldman Sachs, and the IMF is shifting the labor market consensus: AI isn't purely a job-destroyer, and in controlled experiments, it measurably raises productivity and output for the workers who use it. The gains are most pronounced for less-experienced workers — an unusual finding that opens a real policy window. Whether those gains survive structural displacement pressures, and whether they reach workers outside high-income economies, remains genuinely unresolved.

Key Takeaways

  • MIT economists documented that AI assistance raised customer-service worker productivity by an average of 14%, with the steepest improvements among newer, less-tenured employees, according to Brynjolfsson, Li, and Raymond's 2023 NBER working paper on generative AI in the workplace
  • Noy and Zhang documented in Science in 2023 that ChatGPT reduced professional writing task completion time by approximately 40% while improving output quality scores — suggesting AI is compressing the performance gap between workers, not just lifting the top
  • The IMF estimated that roughly 40% of global employment is exposed to AI, rising to 60% in advanced economies, with the Fund's January 2024 staff discussion note arguing that roughly half of that exposure has a complementary rather than substitutive character
  • Goldman Sachs economists projected generative AI could raise global GDP by approximately 7% while also displacing an estimated 300 million full-time jobs globally, per their March 2023 research — a finding whose headline growth figure and displacement estimate are rarely cited together
  • The World Economic Forum's Future of Jobs Report 2023 projected 83 million job roles displaced and 69 million created by 2027, a net loss of 14 million, with sharp variation by sector and national income level
  • David Autor's MIT research argues AI may partially reverse labor-market polarization by augmenting judgment-based tasks in middle-wage occupations — a meaningful reversal of the hollowing-out thesis that dominated labor economics for two decades
  • Bloomberg Intelligence's AI investment tracker logged over $200 billion in corporate AI commitments in 2024, but deployment in labor-intensive industries lags knowledge-sector adoption by several years, meaning today's wage effects are sector-specific, not economy-wide

What the Research Is Actually Saying — and Why Timing Matters

The debate on AI and work has oscillated between two poles: displacement apocalypse and productivity miracle. What distinguishes the evidence published since 2022 is that it's empirical. Controlled experiments. Actual firms. Actual workers. Actual output data.

That shift matters for how the labor market conversation gets framed. For years, the debate rested on task exposure models — theoretical counts of what AI could automate — rather than observations of what happens when workers actually use AI tools. The current wave of evidence closes some of that gap. Not all of it, but enough to change the policy conversation.

The critical nuance buried in several studies: AI appears most complementary for workers at the lower end of the skill distribution within a given role. Newer employees, workers still building competence, people without established professional networks — these are the groups seeing the largest productivity gains. That's genuinely unusual. Most general-purpose technologies historically benefited experienced workers first, compounding existing advantages. This one appears to run in the other direction, at least in the settings studied so far.

Whether this holds at scale — across industries, across national income levels — is the open question. Current evidence is concentrated in knowledge work, customer service, and software development. Manufacturing and physical labor are underrepresented in the data.

The Specific Evidence: Data, Benchmarks, and Findings

How the 14% Productivity Gain Plays Out — and Who Captures It

The Brynjolfsson, Li, and Raymond study is the most-cited controlled experiment on AI and labor productivity. The researchers tracked over 5,000 customer service agents at a large technology firm over 18 months after the company deployed an AI-powered conversation assistant. Average productivity rose 14%. That number is significant on its own. The distribution, though, is the story: agents in the top performance decile saw almost no gain, while agents in the bottom decile improved by over 30%.

AI was compressing the performance distribution from below — raising the floor without raising the ceiling much.

This has a direct implication for wages. If AI allows mid-tier workers to perform near the level of senior colleagues, wage compression follows. For employers, that's a cost reduction. For workers, it's a credible basis for higher base pay: "I produce what a senior employee produces; the compensation should reflect that." Whether that argument wins in negotiation depends entirely on labor market structure — which brings us to the gap in the data.

The Wage Effects: Who Gains, and Where the Evidence Is Thin

The wage evidence is less settled than the productivity evidence. The strongest theoretical case comes from Autor's ongoing MIT research, which argues that AI could restore demand for non-routine, middle-skill tasks — the occupational category that automation hollowed out from the 1990s onward. His argument: prior automation eliminated routine tasks (data entry, repetitive analysis); AI is uniquely better at augmenting judgment-based work than replacing it.

That framing is hopeful. Daron Acemoglu has argued the opposite — that substitution effects in service and administrative roles will outpace augmentation gains, and that productivity estimates from AI deployments are narrower than advocates claim. Neither position is fully settled by current data.

What can be said with more confidence: in firms that deployed AI tools alongside structured training programs, early wage data shows modest upward pressure in high-exposure occupations. In firms that deployed AI alongside layoffs, wages for remaining workers rose, but the total wage bill fell. Both outcomes are real. Which one dominates across an economy depends on policy choices, not just technology.

The IMF's 40% Exposure Framework — and What It Means for Different Economies

The IMF's January 2024 staff discussion note introduced an AI exposure index that differs from earlier task-exposure models. Rather than asking what tasks AI can automate, it asks which tasks are simultaneously high-skill and cognitively intensive — the zone where AI and human workers are most likely to collaborate productively.

On that measure, advanced economies are highly exposed, but in a way that creates opportunity for workers whose roles have this complementarity. The harder picture is for developing economies, where service and administrative work is a primary employment base. Those jobs are exposed in the substitutive direction: AI can do the task, the worker faces difficulty retraining, and the wage support infrastructure is weaker.

The distributional fairness question this creates is addressed in the broader debate over AI growth acceleration versus distributional fairness — a tension the IMF's modeling acknowledges but doesn't resolve. The Fund's own recommendation — invest in social safety nets alongside AI deployment, not after — is the conservative policy posture for a reason.

What This Changes for Journalists, Policymakers, and Labor Economists

For Journalists Covering Work and Technology

The new evidence complicates both the utopian and dystopian framings that dominate coverage. A headline reading "AI Destroys Jobs" and one reading "AI Boosts Wages" can both be accurate simultaneously — they apply to different sectors, workers, and timescales.

Source credibility now depends on distinguishing controlled experiment evidence from task-exposure modeling. The two methodologies answer different questions and produce different results. Reporters should push sources to specify which they're using before drawing policy conclusions from productivity studies.

For Policymakers Designing Labor Interventions

The complementarity finding creates a policy opening that didn't exist five years ago. If AI genuinely raises the productivity floor — helping lower-performing workers more than high-performers — then public investment in AI access and workforce training is defensible as a labor equity intervention, not just a productivity play.

The risk: policy outpacing evidence. Designing national retraining programs around findings from customer-service pilots at large tech firms may not generalize to healthcare, logistics, or construction. The IMF's timeline recommendation — act on safety nets and transition support now, before displacement accelerates — is the credible hedge.

For Labor Economists and Researchers

The methodological challenge is endogeneity: firms that deploy AI effectively may already be better managed, better capitalized, and better at retaining staff. Attributing wage or productivity gains purely to AI, without controlling for firm quality, is a serious identification problem in the current literature.

The research frontier is longitudinal wage data at the individual worker level — not firm level — following AI adoption. That data is beginning to accumulate. The next three to five years will likely yield the first credible answers to whether AI-driven productivity gains translate into durable individual wage trajectories.

AI Tools Researchers and Analysts Are Using to Study Labor Markets

Journalists, economists, and policy analysts increasingly rely on AI tools to synthesize the literature they're analyzing. A practical comparison of the tools most relevant to this domain:

ToolBest ForLimitationsCost (2026)
Perplexity AIFast literature scanning with inline citationsDoesn't reliably surface paywalled research; citations require manual verificationFree / $20/mo Pro
ElicitSystematic academic paper review; extracts findings from uploaded PDFsFocused on empirical papers; weaker on working papers and grey literatureFree tier / ~$12/mo
ConsensusSynthesizing findings across multiple studies on a single claimLimited to published research; no real-time data or recent preprintsFree / $9/mo
Claude (Anthropic)Long-document analysis, policy brief drafting, comparative synthesisNo real-time web access in base mode; doesn't pull live labor statisticsFree / $20/mo Pro
NotebookLM (Google)Analyzing uploaded PDFs — reports, working papers — with Q&AWorks only with documents you upload; no independent web searchFree
ChatGPT + BrowseBroad research queries with live web accessInconsistent citation accuracy; prone to confident-sounding errors$20/mo

For labor economists: Elicit and Consensus are the most rigorous for academic synthesis. For journalists on deadline: Perplexity with manual citation verification is the most practical workflow.

Checklist: How to Evaluate AI Labor Market Claims Before You Cite Them

Before using AI-productivity research in reporting, policy briefs, or academic work, run through these questions:

  • Is this task exposure or empirical measurement? Task exposure models estimate potential disruption. Controlled experiments measure actual outcomes. They're not interchangeable, and conflating them produces bad policy.
  • What sector is the evidence from? Knowledge-work findings don't generalize directly to manufacturing, physical labor, or the informal economy.
  • Who funded the research? Vendor-funded productivity studies consistently show larger gains than independent research on the same interventions.
  • What's the time horizon? Short-term productivity gains can coexist with medium-term job displacement. Both can be documented from the same deployment.
  • Is the comparison group meaningful? "Workers with AI versus workers without AI at the same firm" is strong evidence. "Countries with more AI investment grew faster" is weak.
  • Are wages or just output being measured? Productivity gains do not automatically become wages. The transmission depends on union coverage, market power, and competitive dynamics.
  • How long after deployment was the measurement taken? Early productivity data often understates long-run effects — in both directions. The internet productivity paradox ran for nearly a decade before resolving.

Where This Is Heading

The complementarity thesis will face its first real stress test in mass-market services. Controlled enterprise experiments with deployed tools and structured training are not the same as millions of freelancers and small-business workers all gaining AI capability simultaneously. When every entry-level analyst, copywriter, or paralegal has AI assistance — and so does every competitor — the individual productivity gains may be real but competed away at the market level.

Wage negotiation will incorporate AI productivity evidence. This is already visible in some U.S. tech-sector labor negotiations, where worker groups have cited productivity research to argue for wage floors tied to AI-assisted output benchmarks. Expect this to spread into higher-unionization sectors within three to five years, reframing AI deployment as a shared-gains negotiation rather than a unilateral employer decision.

The developing-economy divergence will widen before institutions can respond. The IMF's complementarity finding disproportionately applies to advanced economies with knowledge-intensive labor structures. For countries where service and administrative work is a primary employment base, substitution risk is higher and policy infrastructure weaker. International labor institutions are behind the deployment curve on this.

Individual-level wage data after AI adoption is the next empirical frontier. Several national statistics agencies and research institutions are building longitudinal datasets that track individual workers through AI tool adoption at their employers. When those datasets mature — likely 2026 to 2028 — they will either confirm or challenge the complementarity thesis at a scale that current experiments cannot.

The "task augmentation" hypothesis needs sector-level validation. Autor's argument that AI restores middle-skill demand is theoretically coherent but not yet validated at scale outside white-collar professional work. Healthcare and logistics are the next test cases: both sectors are in active AI deployment, outcome data is thin, and the stakes are high.

FAQ

Does AI actually raise wages, or just productivity for the employer? Both, in different configurations. In the Brynjolfsson et al. study, productivity gains accrued primarily to the employer as output volume — wages weren't explicitly tracked. In some firms, AI-driven productivity improvements have been used as the basis for negotiated wage increases. The transmission from productivity to wages depends on labor market power structures, and there's no automatic mechanism. Employers capture gains by default unless workers can organize around them.

Who benefits most from AI assistance in the current evidence base? Consistently: less-experienced workers within a given role. AI appears to compress performance distributions, helping the bottom quartile of performers reach closer to median output. Whether this translates into individual wage gains — or simply lowers employer training costs while pay stays flat — is institution-specific and contested.

Is there evidence AI creates new jobs, or only that it displaces them? Both effects are documented, but they're not synchronized. Job creation in AI-adjacent roles (model evaluators, AI safety reviewers, prompt engineers, data curators) is real but small in absolute numbers. Displacement risk in administrative, clerical, and routine cognitive roles is larger and moving faster. The WEF's net estimate of -14 million jobs by 2027 is the most widely cited summary, but the methodology for counting "new roles" is a legitimate target for scrutiny.

Should policymakers wait for more evidence before acting? No — but they should be specific about what evidence would change which policy. Retraining investment is defensible under current evidence. Labor market insurance redesign (portable benefits, income smoothing) is defensible as a hedge against displacement scenarios the evidence doesn't rule out. Waiting for certainty before any action is the highest-risk posture, given that deployment timelines are not waiting for the evidence.

How credible is Goldman Sachs' 7% GDP growth projection from AI? It's a long-run projection under assumptions that are individually debatable. The $7 trillion headline figure gets repeated without its confidence interval or the timescale attached to it. Goldman's methodology assumes productivity gains from experimental settings are sustained over a decade at scale — a strong assumption with no historical precedent at this speed. Treat it as a directional signal, not a forecast.

What does Bloomberg's AI investment tracking tell us about labor outcomes? Investment data predicts deployment timelines, not outcomes. High AI capital expenditure correlates with faster tool rollouts, which is a leading indicator of where labor market effects will appear. The connection from AI deployment to measurable wage and employment changes has historically lagged by three to five years in comparable technology cycles. Bloomberg's data tells you where to watch; it doesn't tell you what has happened.

What's the single most important thing labor economists should track in the next 24 months? Individual-level wage trajectories after AI tool adoption — not firm-level productivity aggregates, not sector-level headcounts. The key empirical question is whether the worker who uses AI tools sees measurable wage growth relative to a comparable worker who does not. That data is beginning to exist. Whoever publishes it first will significantly shape the next round of policy decisions.

A
> Editor in Chief **20 years in tech media**, the first 10 in PR and Corporate Comms for enterprises and startups, the latter 10 in tech media. I care a lot about whether content is honest, readable, and useful to people who aren’t trying to sound smart. I'm currently very passionate about the societal and economic impact of AI and the philosophical implications of the changes we will see in the coming decades.

Related Articles