In July 2023, two MIT economics PhD students, Shakked Noy and Whitney Zhang, published an experiment in the journal Science involving 453 college-educated professionals — people in marketing, consulting, HR, and grant writing. Half wrote work documents, such as a cover letter or a corporate email, with no help at all; the other half had access to an AI model. Evaluators who didn't know who had used what judged the results: people who used AI took 40% less time and turned in work rated 18% better.
That number became a staple of every presentation about AI at work. The trouble is what usually happens next: a company reads "40% less time" and assumes the same gain repeats for every task, for every person. It doesn't. Two years after the MIT study, another piece of research measured the opposite — professionals getting slower with AI, not faster. Both studies are correct. What differs is the type of task.
Understanding where the gain shows up — and where it disappears — is what separates AI adoption that genuinely gives time back from adoption that just pushes the work later, in the form of review and correction.
When the gain is real and measurable
The MIT study already offered a clue: the gain didn't come from AI writing better than the person. It came from restructuring the work — less time on the first rough draft, more time on idea generation and editing what already existed. And the gain wasn't even: workers who started out weaker improved more than those who were already strong, which narrowed the gap between the two groups.
The same pattern showed up, with even sharper numbers, in a study published by the NBER (the U.S. National Bureau of Economic Research) by Erik Brynjolfsson, Danielle Li, and Lindsey Raymond. They tracked 5,179 customer-support agents at a software company before and after a generative AI assistant arrived that suggested real-time responses during the conversation — the agent still decided; the tool only suggested.
The average productivity gain was 13.8%. But that average hides the real story: novice and lower-performing agents improved by 35%; the most experienced and highest-rated agents saw an effect close to zero, in some cases slightly negative. An agent with just two months on the job, using the tool, performed at the level of a six-month colleague working without it. The researchers' explanation: the AI was, in effect, spreading the working style of the people who already knew how to do the job well to the people still learning it.
The common thread between the two studies: the gain concentrates on tasks that already have a pattern to follow — a draft similar to one that worked before, a support reply that follows a known script, a summary, a comparison, a triage decision. It's work where "what to do" is already mapped out; AI speeds up the execution.
When the gain disappears — or turns into a loss

In July 2025, the research organization METR published a result that contradicts the intuition of anyone who has ever used an AI assistant to write code. Sixteen experienced developers — averaging five years of work on the very projects they knew deeply — completed 246 real bug-fix and maintenance tasks, randomly assigned to use or not use AI tools on each one.
The result: with AI, the developers took 19% longer to finish the same task. Before the study, they had expected to be 24% faster. After actually living through the slowdown, they still believed they had been 20% faster — perception didn't match the stopwatch, on either end.
The explanation isn't that the AI model "got worse" for people who are already skilled. It's that, in a large, mature project, the knowledge that makes someone experienced fast isn't written down anywhere the AI can consult: it's architectural decisions made years earlier, internal team conventions, the reasons a piece of code looks the way it does. The AI suggested something plausible; the developer had to review it, figure out why it didn't fit, and correct it — and that review ate up more time than the ready-made text saved.
It's the same pattern that showed up, on a smaller scale, in the support study: the more a task depends on judgment over context that only exists in the head of whoever is already good at it — and was never written down anywhere the AI could read — the smaller the gain, to the point of turning into a loss.
The risk is concrete for Brazilian and Latin American companies adopting AI quickly, without first organizing their internal knowledge. The gain shows up fast on the simplest tasks — and creates the expectation, inside the company, that it will show up everywhere. Including in the decisions that only someone who has worked there for years knows how to make, and that were never written down anywhere.
What has to be in place
The split between the two sets of results points directly to what a company needs to build before it can expect the MIT-style gain instead of the METR-style loss.
Approved knowledge with sources, not knowledge that only lives in someone's head. The reason the experienced developers got slower was the absence of a place where the project's decisions and conventions were documented and accessible to the AI. When the material that makes an expert fast is written down, reviewed, and citable, the AI stops guessing and starts consulting.
Reusing the request and the agent that already worked. In the NBER study, the gain came from spreading the pattern of someone already good at the job to someone still learning. That only holds up if the script that worked gets saved and stays available for someone else to use — not reinvented in every conversation.
Cost and gain visible by area, not by "AI use" in general. An average of 13.8% hides 35% gain in one group and almost none in the other. Without measuring by task type and by team, a company doesn't know where it's worth investing training time and where the problem is something else entirely.
Human approval proportional to the task's risk. In work that requires judgment — not just execution of a known pattern — keeping a review step before any decision moves forward is what stops a plausible-but-wrong suggestion from becoming a loss disguised as productivity.
This is how Skyller was designed: company knowledge with a citable source, requests and agents the team can reuse, and visible consumption by area — so the gain a company measures looks more like MIT's than METR's.
The compounding effect of a good start

Once a pattern that works gets documented and made reusable, the gain doesn't stop with the first person who discovered it. A response script a senior agent refined over months, once saved and released to the team, speeds up everyone who comes after — that's exactly the mechanism behind the 35% jump the novice agents made in the NBER study.
The reverse is also true, and it's the part that tends to go unnoticed: if everyone uses AI on their own personal account, with nothing documented or shared, the company pays the cost of each person relearning alone what someone else had already figured out — in practice reproducing, on a smaller scale, the same slowdown the METR developers felt in the face of context the AI had no way of knowing.
A starting checklist
Before measuring "how much AI helped" as a single number, it's worth splitting the work into two groups and treating each one differently:
- List the tasks that already have a clear pattern — drafting an email, summarizing a document, triaging a request, answering a recurring question. These are the closest to the MIT and NBER studies, and where the gain tends to show up fast.
- Set apart the tasks that depend on judgment over undocumented context — decisions that only make sense to someone who knows the history of the project or the client. On these, treat the AI's output as a draft to review, not a finished answer.
- Document what already works before asking the team to use more AI. Without this, each person repeats alone the work of figuring out the right path.
- Measure by area and by task type, not by overall usage. A comfortable average can be hiding an entire team with no gain at all.






