No Result
View All Result
SUBMIT YOUR ARTICLES
  • Login
Wednesday, August 26, 2026
TheAdviserMagazine.com
  • Home
  • Financial Planning
    • Financial Planning
    • Personal Finance
  • Market Research
    • Business
    • Investing
    • Money
    • Economy
    • Markets
    • Stocks
    • Trading
  • 401k Plans
  • College
  • IRS & Taxes
  • Estate Plans
  • Social Security
  • Medicare
  • Legal
  • Home
  • Financial Planning
    • Financial Planning
    • Personal Finance
  • Market Research
    • Business
    • Investing
    • Money
    • Economy
    • Markets
    • Stocks
    • Trading
  • 401k Plans
  • College
  • IRS & Taxes
  • Estate Plans
  • Social Security
  • Medicare
  • Legal
No Result
View All Result
TheAdviserMagazine.com
No Result
View All Result
Home Market Research Startups

The best evidence on AI and productivity does not agree with itself: Harvard and BCG found consultants using GPT-4 finished skilled tasks 25 per cent faster, while a separate trial found experienced coders using AI tools were 19 per cent slower, and did not notice

by TheAdviserMagazine
3 weeks ago
in Startups
Reading Time: 4 mins read
A A
The best evidence on AI and productivity does not agree with itself: Harvard and BCG found consultants using GPT-4 finished skilled tasks 25 per cent faster, while a separate trial found experienced coders using AI tools were 19 per cent slower, and did not notice
Share on FacebookShare on TwitterShare on LInkedIn


Ask anyone who has spent the last two years wiring an AI assistant into their workday whether it has made them faster, and the answer is almost always yes, delivered with total confidence. The best evidence available does not agree with itself, and worse, it suggests that confidence is exactly the thing not to trust.

Start with the result that made AI look like the productivity story of the decade. Researchers at Harvard Business School, working with the Boston Consulting Group, gave 758 BCG consultants a set of 18 realistic business tasks and split them into groups with and without access to GPT-4. For tasks that sat within what the model was actually good at, the AI-assisted consultants were 12.5 per cent more likely to complete the task successfully, finished it 25.1 per cent faster, and produced work independent evaluators rated as roughly 40 per cent higher in quality. Those are not modest numbers. They are the kind of result that gets cited in every AI strategy deck written since.

The frontier is jagged, and stepping off it is expensive

The same study, led by Fabrizio Dell’Acqua and colleagues, also found the part that rarely makes the slide deck. When consultants used GPT-4 on tasks that sat just outside the model’s actual capability, a distinction the researchers call the “jagged frontier” because AI is unpredictably brilliant at some tasks and unpredictably weak at adjacent-looking ones, performance did not just fail to improve. It got worse than the group that used no AI at all. The consultants tended to accept plausible, confident-sounding AI output without checking it closely enough, a pattern the paper’s authors nicknamed falling asleep at the wheel. The tool did not fail loudly. It failed quietly, and the human in the loop did not notice in time.

That would already complicate the simple story that AI makes knowledge work faster. A second study, published by the AI evaluation organisation METR in 2025, complicates it further, this time for a task where confident predictions of an AI productivity boom have been loudest: software engineering.

Slower, and certain they were faster

METR ran a randomised controlled trial with 16 experienced developers, each a maintainer on a large open-source project, working through 246 real coding issues, bug fixes, features and refactors on codebases they already knew well. Each issue was randomly assigned to be done with AI tools allowed or not. The developers used Cursor with a frontier Claude model, and were paid $150 an hour to work as they normally would, screen recorded throughout.

Before the study, the developers expected AI assistance to speed them up by 24 per cent. What the trial actually measured was the opposite: tasks done with AI took 19 per cent longer to complete. The gap between expectation and result is large enough on its own. What happened next is the more useful finding. After the study, having lived through the slower outcome themselves, the developers still estimated that AI had sped them up, by about 20 per cent. Direct, recent, personal experience of a real slowdown did not update their belief that it had been a speed-up.

METR’s own analysis does not settle on one tidy explanation. Among the factors it points to are that most of the developers had only around 50 hours of experience with the specific tool by the time the study ran, and that the review standards on mature open-source projects are unusually exacting, the kind of environment where a plausible-looking suggestion needs more scrutiny, not less. Neither factor makes the result an indictment of AI coding tools in general. Both suggest the gap between feeling faster and being faster is not a fluke of one bad study design.

What the two findings have in common

Put next to each other, the BCG and METR results are not actually a contradiction. They are the same warning from two different angles. AI assistance can produce a real, measurable, sizeable productivity gain, the consulting study’s numbers are large by the standards of any workplace intervention, but only inside the boundary of what the tool is genuinely capable of on that specific task, and only when the person using it is scrutinising the output rather than accepting it. Once a task drifts outside that boundary, or once the human reviewing the work relaxes because the last ten suggestions were good, the same tool can make things measurably worse while still feeling helpful.

Broader industry data, gathered together in one recent review of surveys and datasets, points the same way. A DX survey of around 121,000 developers at more than 450 companies found 92.6 per cent using AI coding tools at least monthly, a figure close to the 90 per cent adoption reported separately in the 2025 DORA report on software delivery. Adoption, in other words, is close to universal. Yet a review of six independent studies on organisational productivity found the gains converging on roughly 10 per cent at the level of a whole team or company, nowhere near the 25 per cent the BCG consultants achieved on tasks squarely inside the frontier. One dataset from Faros AI, tracking more than 10,000 developers, found teams with the heaviest AI use merged 98 per cent more pull requests, while review time rose 91 per cent, reported bugs rose 9 per cent, and the broader delivery metrics organisations actually care about stayed flat. Near-universal adoption and a large, well-documented capability boost have not, so far, translated into a large measured gain once the work is checked all the way through.

That last part is the awkward one for anyone hoping to manage this with a simple rule. Feeling faster turned out to be a poor guide to being faster in the one study that checked. The BCG consultants who wandered off the frontier presumably also felt they were being helped, right up until an evaluator scored their output. None of this argues for abandoning these tools, the productivity gains inside the frontier are real and worth having. It argues against trusting your own sense of how well a session went as the measure of whether it actually worked, which is precisely the measure most people are relying on.

I wrote a while back about how being busy and being productive are routinely confused, and this feels like the same mistake wearing a new tool. The confidence that something worked is cheap to produce and easy to feel. Whether it actually did is a separate question, and on the evidence so far, it is one that generally needs an outside check to answer honestly.



Source link

Tags: agreeBCGcentCodersconsultantsevidenceExperiencedfasterFinishedGPT4HarvardNoticeProductivityseparateskilledSlowerTasksToolstrial
ShareTweetShare
Previous Post

Apple Pulls Telegram After Illegal Content Appears in Public Group

Next Post

The Syllabus and the Scar Tissue: What Leadership Preparation Programs Leave Out and Why It Matters – Faculty Focus

Related Posts

edit post
Psychology says procrastination about retirement may be less about discipline than identity — brain scans found people often represent their future selves more like strangers than like themselves, and experiments using age-progressed faces made tomorrow’s person feel real enough for participants to save more money for them

Psychology says procrastination about retirement may be less about discipline than identity — brain scans found people often represent their future selves more like strangers than like themselves, and experiments using age-progressed faces made tomorrow’s person feel real enough for participants to save more money for them

by TheAdviserMagazine
August 8, 2026
0

Retirement procrastination is usually described as a failure of discipline. The forms have been sitting in a folder for months....

edit post
The self-improvement industry sells becoming your best self through grit and mindset, but the large meta-analyses are deflating: grit turns out to be mostly conscientiousness renamed, and growth-mindset programmes move academic results only slightly

The self-improvement industry sells becoming your best self through grit and mindset, but the large meta-analyses are deflating: grit turns out to be mostly conscientiousness renamed, and growth-mindset programmes move academic results only slightly

by TheAdviserMagazine
August 8, 2026
0

Become your best self is one of the most profitable sentences in modern culture. It sells books, courses, apps, and...

edit post
Most advice on becoming happier assumes it is up to you, but the evidence is humbler: much of the variation is dispositional, the claim that 40 per cent sits within your control does not hold up, and what works best points outward, towards other people

Most advice on becoming happier assumes it is up to you, but the evidence is humbler: much of the variation is dispositional, the claim that 40 per cent sits within your control does not hold up, and what works best points outward, towards other people

by TheAdviserMagazine
August 7, 2026
0

The advice on how to become happier is endless, confident, and mostly built on the same assumption: that your happiness...

edit post
Malachyte Raises M to Solve E-Commerce’s Biggest Blind Spot: the Visitor Who Never Logs In – AlleyWatch

Malachyte Raises $10M to Solve E-Commerce’s Biggest Blind Spot: the Visitor Who Never Logs In – AlleyWatch

by TheAdviserMagazine
August 7, 2026
0

E-commerce brands now spend roughly 40% more to acquire each new customer than they did in 2023, yet the website...

edit post
Psychology says the ability to eat alone in public is a sign of quiet confidence because it takes a certain kind of security to occupy a table built for two and not feel the empty chair as an accusation

Psychology says the ability to eat alone in public is a sign of quiet confidence because it takes a certain kind of security to occupy a table built for two and not feel the empty chair as an accusation

by TheAdviserMagazine
August 7, 2026
0

Somewhere in your city tonight, a person is going to walk into a restaurant they have wanted to try for...

edit post
The barrier to starting a conversation is usually a bad forecast, not a lack of charm: people underestimate how much others like them, and commuters made to talk to a stranger enjoyed the trip more than those left alone, the opposite of what they predicted

The barrier to starting a conversation is usually a bad forecast, not a lack of charm: people underestimate how much others like them, and commuters made to talk to a stranger enjoyed the trip more than those left alone, the opposite of what they predicted

by TheAdviserMagazine
August 7, 2026
0

Some people seem to strike up conversations anywhere, with the person next to them in a queue, a stranger at...

Next Post
edit post
The Syllabus and the Scar Tissue: What Leadership Preparation Programs Leave Out and Why It Matters – Faculty Focus

The Syllabus and the Scar Tissue: What Leadership Preparation Programs Leave Out and Why It Matters - Faculty Focus

edit post
Finance of America reaffirms 2026 adjusted EPS of .50-.00 as Onity .2B HECM MSR deal closes (NYSE:FOA)

Finance of America reaffirms 2026 adjusted EPS of $4.50-$5.00 as Onity $5.2B HECM MSR deal closes (NYSE:FOA)

  • Trending
  • Comments
  • Latest
edit post
Judge Who Helped Violent Illegal Alien Evade ICE Faces New Test

Judge Who Helped Violent Illegal Alien Evade ICE Faces New Test

July 31, 2026
edit post
Garbage Trucks Surveillance Florida Neighborhoods

Garbage Trucks Surveillance Florida Neighborhoods

July 29, 2026
edit post
Does a Revocable Trust Protect Your Assets From Lawsuits and Creditors?

Does a Revocable Trust Protect Your Assets From Lawsuits and Creditors?

August 7, 2026
edit post
Montana Puts Democrats in a Bind as Senate Hopes Fade

Montana Puts Democrats in a Bind as Senate Hopes Fade

August 2, 2026
edit post
New Jersey’s PAS-1 Application Opens the Door to Three Senior Tax Relief Programs

New Jersey’s PAS-1 Application Opens the Door to Three Senior Tax Relief Programs

July 31, 2026
edit post
3 Common Cruise Rules I Broke in Alaska

3 Common Cruise Rules I Broke in Alaska

July 29, 2026
edit post
Explained: How BSE traded fewer contracts after CAS but premiums rose 75% in first week

Explained: How BSE traded fewer contracts after CAS but premiums rose 75% in first week

0
edit post
E.W. Scripps Q2 2026 Loss Widens to -.68/Share, Revenue Down 9%

E.W. Scripps Q2 2026 Loss Widens to -$12.68/Share, Revenue Down 9%

0
edit post
Psychology says procrastination about retirement may be less about discipline than identity — brain scans found people often represent their future selves more like strangers than like themselves, and experiments using age-progressed faces made tomorrow’s person feel real enough for participants to save more money for them

Psychology says procrastination about retirement may be less about discipline than identity — brain scans found people often represent their future selves more like strangers than like themselves, and experiments using age-progressed faces made tomorrow’s person feel real enough for participants to save more money for them

0
edit post
Four AI Escapes Just Redefined “Responsible AI”

Four AI Escapes Just Redefined “Responsible AI”

0
edit post
Bill Ackman’s hedge fund made janitors and receptionists millionaires—and its investment team summer together

Bill Ackman’s hedge fund made janitors and receptionists millionaires—and its investment team summer together

0
edit post
Kalshi Predicts Bitcoin Price Could Reach K in August

Kalshi Predicts Bitcoin Price Could Reach $68K in August

0
edit post
Bill Ackman’s hedge fund made janitors and receptionists millionaires—and its investment team summer together

Bill Ackman’s hedge fund made janitors and receptionists millionaires—and its investment team summer together

August 8, 2026
edit post
Links 8/8/2026 | naked capitalism

Links 8/8/2026 | naked capitalism

August 8, 2026
edit post
Wisconsin: The Next Frontier for Socialists

Wisconsin: The Next Frontier for Socialists

August 8, 2026
edit post
Psychology says procrastination about retirement may be less about discipline than identity — brain scans found people often represent their future selves more like strangers than like themselves, and experiments using age-progressed faces made tomorrow’s person feel real enough for participants to save more money for them

Psychology says procrastination about retirement may be less about discipline than identity — brain scans found people often represent their future selves more like strangers than like themselves, and experiments using age-progressed faces made tomorrow’s person feel real enough for participants to save more money for them

August 8, 2026
edit post
Why You Should Be Wary of Aspartame, but Not Totally Rule It Out

Why You Should Be Wary of Aspartame, but Not Totally Rule It Out

August 8, 2026
edit post
Kalshi Predicts Bitcoin Price Could Reach K in August

Kalshi Predicts Bitcoin Price Could Reach $68K in August

August 8, 2026
The Adviser Magazine

The first and only national digital and print magazine that connects individuals, families, and businesses to Fee-Only financial advisers, accountants, attorneys and college guidance counselors.

CATEGORIES

  • 401k Plans
  • Business
  • College
  • Cryptocurrency
  • Economy
  • Estate Plans
  • Financial Planning
  • Investing
  • IRS & Taxes
  • Legal
  • Market Analysis
  • Markets
  • Medicare
  • Money
  • Personal Finance
  • Social Security
  • Startups
  • Stock Market
  • Trading

LATEST UPDATES

  • Bill Ackman’s hedge fund made janitors and receptionists millionaires—and its investment team summer together
  • Links 8/8/2026 | naked capitalism
  • Wisconsin: The Next Frontier for Socialists
  • Our Great Privacy Policy
  • Terms of Use, Legal Notices & Disclosures
  • Contact us
  • About Us

© Copyright 2024 All Rights Reserved
See articles for original source and related links to external sites.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Financial Planning
    • Financial Planning
    • Personal Finance
  • Market Research
    • Business
    • Investing
    • Money
    • Economy
    • Markets
    • Stocks
    • Trading
  • 401k Plans
  • College
  • IRS & Taxes
  • Estate Plans
  • Social Security
  • Medicare
  • Legal

© Copyright 2024 All Rights Reserved
See articles for original source and related links to external sites.