In June 2025, Gartner published a forecast that changed the tone of the conversation around agentic AI inside companies: more than 40% of projects will be canceled by the end of 2027. Not for lack of ambition — for the opposite reason. According to the firm, the cancellations come from costs that escalate quickly, business value nobody can prove, and risk controls that were never designed for a system that acts on its own.

Gartner senior director analyst Anushree Verma summed up the diagnosis in one line, quoted by CIO: most agentic AI projects today are early-stage experiments, driven more by hype than by proven results. And the market itself helped inflate that hype. Gartner coined a term for vendors rebranding old automation with a new label — and estimates that, of the thousands of vendors claiming to sell AI agents, only about 130 actually deliver on it.

For whoever decides where to invest inside a company, the question that matters isn't "can the agent do this on its own?" It's a more uncomfortable one: what does it cost if it gets it wrong — and who notices in time?

Why the project stalls before it scales

Gartner points to three causes for cancellation, and none of them is a lack of model intelligence. The first is business value: today's systems still lack the maturity to pursue a complex goal on their own for long, so the return promised at the start of the project doesn't show up in practice.

The second is cost. AI usage billing climbs fast once a company stacks coordination and control layers on top of the agent — a pattern that surprises anyone who budgeted based on a simple conversational tool, which is far cheaper to run.

The third, and most structural, is risk control. Verma explains that the governance most companies already have — an approval spreadsheet, an access policy, an audit review — was designed for a system that waits for a human to click at every step. An agent acting on its own, especially when several agents interact with each other, doesn't fit that design without adaptation.

Governance and risk controls aren't really designed, at this time, precisely for agentic systems.

Anushree Verma, Gartner, via CIO

The picture barely changes when you look at day-to-day operations instead of the project budget. A survey by Gravitee of more than 900 executives and technical practitioners, published in February 2026, found that 88% of organizations already reported a confirmed or suspected security incident tied to AI agents in the past year — in healthcare, that rate tops 92%. The pattern repeats: the failure is rarely in the model. It's in who defined — or failed to define — what the agent could do on its own.

For a Brazilian or Latin American company entering this cycle now, the warning arrives at a good time: it's possible to skip the "let the agent decide everything and hope for the best" phase entirely, because other companies have already paid for that lesson.

Full autonomy or manual control don't solve it

Full autonomy or manual control don't solve it

Faced with these numbers, the most common reaction is one of two extremes, and both fail for the same reason: they treat every action the agent takes as if it carried the same weight.

At one extreme, the company lets the agent act on its own across the board, to "capture the promised productivity gain." The problem shows up fast: a wrong decision isn't slower or smaller for having been made by an agent — it just gets made faster, and in more places at once, than a person ever could. That's exactly the scenario behind the incidents Gravitee measured.

At the other extreme, after a scare or a security audit, the company locks everything down: any agent action now requires human review, in any system, for any task. The practical result is that AI becomes a slower form, an approval queue grows, nobody trusts the pending item will be seen in time, and the team goes back to doing manually what the automation was supposed to solve.

Neither route holds up, and Gartner is describing exactly this dynamic when it talks about cost escalating and projects getting canceled: a project with no middle ground lasts until the first major incident, or until the first quarter nobody can prove the return. What's missing in both cases is the same thing — a way to separate, inside the workflow itself, what's routine from what's sensitive.

What has to be in place

A well-scoped agent rests on four verifiable mechanisms, not a promise of "trustworthy AI."

Risk classification by type of action, not by whole project. Sending a routine report, updating a record, answering a question based on an approved document: that runs on its own. Approving a payment, changing a policy, sending something outside the company: that pauses. The classification lives in the action, not in a generic list of "tasks allowed for agent X."

Approval inside the conversation itself, not in another system. When an action is classified as sensitive, the agent stops and asks a person to confirm right there, in the same flow where the task is happening — without forcing anyone to open a second tool to approve what the first already knew needed approval.

One single queue for pending decisions. Whoever approves needs to see, in one place, everything waiting on a decision — from any agent, any department. That queue is what prevents the invisible backlog that broke the second extreme described above.

An audit trail per action. Who requested it, what the agent did, who approved it and when: logged. In an incident investigation or a compliance audit, that's the difference between reconstructing what happened in minutes and not being able to reconstruct it at all.

This is how Skyller was designed: human approval before a sensitive action happens inside the conversation itself, and every decision is logged in a single audit trail, by system area.

The payoff shows up once the limit is clear

The payoff shows up once the limit is clear

The effect of classifying risk by action doesn't just show up in the security section of the internal report — it shows up in how fast the team trusts automation again after a mistake.

When the line between "the agent handles it alone" and "the agent asks for confirmation" is explicit, the team stops treating every new task as a gamble. The person approving knows exactly what's being asked and why, because the pending item arrived already classified, not as a generic alert. And leadership can answer, with data instead of opinion, the question that decides whether the project survives the next budget review: how much of this work is running on its own, safely, today?

That answer — not the promise of full autonomy — is what separates the project that reaches year three from the one that becomes a line in Gartner's cancellation statistics.

Questions before approving the next agent

Before greenlighting a new agent or a new automation, it's worth bringing these questions to the meeting with IT, legal, and whoever will use the tool day to day:

  1. What types of actions will this agent perform, and which of them, if they go wrong, require an explanation to a customer, an auditor, or a regulator? That list separates routine from sensitive before the first use.
  2. Where will the approver see the pending item — inside their own workflow, or on one more screen they'll forget to open? Approval outside the workflow is approval that's delayed or never happens.
  3. If this agent makes a mistake today, will someone know within minutes, or find out weeks later in an audit? The answer measures the audit trail, not the agent's intelligence.
  4. Besides whoever built the agent, who can explain why it's allowed to do what it does? If the answer is "only whoever configured it knows," the scope isn't documented — it's in one person's head.

Discover Skyller