Deel WorksReporting a feature for Deel Works, Deel's publication for HR and people leaders, on why enterprise AI rollouts underperform. The argument: almost nobody measures AI productivity with anything but self-report, and self-report has been shown experimentally to fail here.
METR ran a randomized trial with 16 experienced developers across 246 real tasks in repositories they knew well. They predicted AI would make them 24% faster. Measured, they were 19% slower. Afterward they still believed they had been 20% faster. Separately, Workera's benchmark across 88,753 assessments found verified AI skills diverge sharply from self-reported proficiency, while 85% of L&D leaders say they trust self-reported skills data.
I'm looking for people who measured AI's effect on work with real operational data.
Not looking for: adoption dashboards, license utilization, satisfaction surveys, or vendor-supplied ROI models. If your measurement is people reporting how much time they think they saved, that's what this piece is about — I'd still like to hear from you, but please say so.
What you measured
What did you instrument? Cycle time, throughput, error rates, rework hours, revision counts, something else?
What was your baseline, and how did you establish it before AI was in the workflow?
What did the measurement show? I'm interested in flat and negative results as much as positive ones.
Did the measured result match what employees said about their own productivity? If there was a gap, how big and which direction?
How you did it
Who pushed for real measurement, and did anyone resist? What was the objection?
What did it cost to set up, and what would you do differently?
Did trained users measurably outperform untrained ones at your company, or is that a distinction you can't actually make with your data?
How long after deployment did you start measuring? Was there a usable pre-AI baseline or did you have to reconstruct one?
Did the effect change over time? Some studies find early gains that flatten or reverse.
If you didn't measure
Why not? Was it never raised, raised and dropped, or actively decided against?
What are you using instead to decide whether the rollout is working?
Please include your name, job title, company, company size, industry, and what specifically you measured. I quote sources by name and title, and link your LinkedIn and company site. Partial anonymity is available if the numbers are sensitive — tell me what can and can't be attributed.