McKinsey's global AI survey, with nearly 2,000 respondents across 105 countries in mid-2025, produced two numbers that shouldn't sit this close together. The first is 88%: the share of companies already using AI in at least one part of the business. The second is 39%: the share that can point to any measurable impact on the bottom line — and most of those report less than 5% of a contribution.
The Boston Consulting Group reached a similar picture from the value side: in its October 2025 report, "The Widening AI Value Gap," the firm found that 60% of companies are still capturing little to no material value from their AI investments, even as adoption keeps climbing. And Gartner, in June 2025, went further: it projected that over 40% of agentic AI projects will be canceled by the end of 2027, citing three specific causes — rising cost, unclear business value, and inadequate risk controls.
The pattern across all three studies is the same: the company knows it's using AI, but can't put a number on what that AI is actually giving back. And without that number, any decision to expand, cut, or continue an AI project gets made in the dark.
For Brazilian and Latin American companies entering this stage of adoption now, there's an advantage in skipping the most expensive mistake: instead of buying the tool and only later asking how to measure it, it's possible to decide the metric before turning on the first project — and save a full year of license reports with no results report to match.
The Problem Is the Wrong Unit
When a company measures its AI adoption, the easiest number to get is licenses or active accounts: how many people have access, how many logged in last week. It's a real data point, and it explains why McKinsey's adoption figure keeps climbing every year.
But a license isn't the unit that produces results. Results come from tasks: a billing inquiry answered faster, a report that used to take a day and now takes an hour, a response script that stopped generating rework. Measuring "how many people have access" measures the input. Measuring "how long this task took before and how long it takes now" measures the effect — and that second number is exactly what's missing at most of the companies in McKinsey's survey.
BCG's report points to the same gap from a different angle: companies that don't set a clear financial indicator for each AI initiative from day one tend to end up in the group that can't prove value later. It isn't that the value doesn't exist — it's that nobody defined, on day one, what would count as proof of it.
This mix-up between input and effect also explains why different teams at the same company reach opposite conclusions about the same tool. Whoever tracks adoption says "it's going well"; whoever tracks the budget says "we don't see a return." Both are right, because they're measuring different things — and neither is the question the business result actually needs answered.
Why the Metric-Free Pilot Fails

Gartner's projection helps explain what happens to that measurement vacuum over time. An AI project starts as an experiment, with no defined baseline. Months later, someone asks whether it was worth it, and the answer depends on opinion — "people seem to like it," "I heard it helps" — because there's no before/after comparison and no cost tied to the task.
Without that number, the project becomes an easy target at the next budget review: nobody can defend it with data, and the decision to cancel it is just as arbitrary as the decision to start it was. That's exactly the pattern Gartner describes — rising cost, value nobody can point to, controls that were never designed in from the start.
There's also a third, less-discussed effect: when a project dies without a metric, the company also loses the lesson about why it didn't work. Without a baseline, there's no way to know whether the problem was the task chosen, missing access to the right information, or simply nobody owning the result. The next pilot starts from zero, free to repeat the same mistake.
What has to be in place
Measuring year one of AI in a way that survives the question "was it worth it" takes a few simple mechanisms, decided before the project starts — not a report reconstructed later, under pressure.
A concrete task with a before number. Before turning on any project, record how much time, how many steps, or how much cost that task consumed without AI. Without that baseline, "it got better" is opinion, not measurement.
Transparent cost per conversation. Instead of a single monthly invoice, spend tied to each conversation or task, visible by area. That's what lets you answer, without guesswork, how much it cost to automate that specific process.
AI credits shared by the team. Under a per-seat license model, one team has paid capacity sitting idle while the team next door runs short. A shared budget, with consumption visible by area, fixes both sides at once.
Reuse through permissions, not a lost spreadsheet. When a script works, it needs to be shareable with other people and groups within an allowed scope — instead of dying in the personal history of whoever built it. That multiplies the result of one measured task across everyone who repeats it.
A record of who approved what. When an AI-generated answer needs a correction, that trail shows whether the problem was the instruction, the information used, or the task itself — the data point that separates "AI doesn't work" from "one step needed adjusting."
This is how Skyller was designed: cost per conversation, credits shared by the team, and reuse through permissions built into the environment, not a side spreadsheet someone maintains on their own.
What the Metric Changes

With a baseline and a visible cost per task, the conversation about AI stops being about sentiment and starts being about return. It becomes possible to compare two areas that adopted AI in different ways and say which approach cost less per result delivered — the kind of comparison almost none of the companies in the studies above can currently make.
That also changes how budgets get defended. A project with a before-and-after number survives a cost cut because it can justify itself; a project with no number gets cut for the same arbitrary reason it got approved. And once a script that worked can be reused by other people with permission, the return on one well-measured task stops being limited to whoever measured it first.
There's also a gain when deciding where to invest next. With cost per task in two or three different areas, it becomes visible which kind of process responds best to AI at that specific company — instead of picking the next area by impression or by whoever asked the loudest.
A Roadmap for Year One
Before expanding any AI project to a second area, these steps keep it from becoming another cancellation statistic:
- Pick a real process, not a generic pilot. A process that already exists, with a known owner and frequency, is easier to measure than an initiative created just to "test AI."
- Record the before number on day one. Time, cost, or steps — any of them works, as long as it exists before the project goes live.
- Define access and approval before starting, not after an incident. Fixing governance after something goes wrong costs far more than designing it up front.
- Track cost per task, not total monthly spend. A single invoice hides exactly the comparison the company needs to make.
- Review reuse every quarter. A good script nobody else uses is lost results, not proven ones.






