In June 2025, Gartner published a forecast that has become a fixture in budget conversations: more than 40% of agentic AI projects will be canceled by the end of 2027. The three reasons given are not technical — escalating costs, unclear business value, and inadequate risk controls. None of them is "the model got it wrong." All of them are governance.

The market noticed the size of the bill. The FinOps Foundation's State of FinOps 2025 report shows that 63% of organizations now actively manage AI spend — up from 31% the year before. The number doubled in twelve months. But the same survey shows where the attention sits: allocating, measuring, and forecasting consumption. Optimization is still a low priority. In other words: companies have just learned to see the bill. Few have learned to steer the spend.

The pattern behind the surprise is always the same. A company gives a team or an agent permission to use AI. Nobody sets a cap. Each person gets a license. The agent starts running calls back to back: check, summarize, search, generate, repeat. At month-end the invoice doesn't match the forecast — and nobody can explain which part of the consumption came from where.

That's why cost governance stopped being an "optimize later" topic. It's a day-one question: who pays, how much, for what, and who approves before the bill becomes internal news.

The pattern nobody sees: per-person license, idle seat

Per-person licensing is easy to buy and hard to justify later. Zylo's software management index, published in 2026, measures exactly that waste: organizations leave an average of 36% of their software licenses unused. These aren't canceled licenses — they're paid seats nobody opens.

With AI, the problem gains a second floor. The same research shows that 78% of IT leaders faced unexpected charges tied to consumption-based or AI pricing models, and 61% had to cut projects because of unplanned software cost increases. Spend on AI-native apps rose 108% in a single year — and 393% at companies with more than 10,000 employees.

Put the two halves together and the picture appears: part of the budget sits parked in seats nobody uses, while the other part blows past the forecast without warning because real consumption has neither a cap nor visibility. Both failures share one root — the unit you buy (the person) is not the unit you consume (the call the agent makes).

A simple exercise makes it obvious. Picture a team of 15 people with one subscription each. Four never open the tool. Nine use it lightly. Two run agents all day and hit their limit before the month ends. The company paid for 15 subscriptions to sustain the work of two — and still throttled exactly the people who were producing. None of these numbers is a market measurement; it's the arithmetic of per-person licensing applied to any company that has opened its own usage report.

The gain from treating budget as budget isn't "spend less." It's seeing what's happening and having the power to decide while it's still happening.

What has to be in place

What has to be in place

Team-shared credits, not per-person licenses. A single monthly budget is shared across the team. When someone runs an agent, the consumption comes out of that common pool. Nobody holds a "personal" idle license, and whoever needs more volume isn't blocked while four seats sleep next door. It's the direct fix for those idle 36%: no use, no consumption, and the budget flows to whoever needs it.

Human approval before a sensitive action. An agent can summarize a hundred emails on its shift without asking anyone. But it stops and asks for confirmation before sending a company-wide announcement, quoting a customer, or changing a contract. This isn't "every action needs approval" — that would kill the gain. It's "an action that costs a lot or commits the company needs a person confirming, right there, inside the conversation."

Model choice by task, not by default. Classifying a support message, cleaning a spreadsheet, and summarizing an email don't need the same AI model as drafting a commercial proposal. When the agent runs on the cheapest model by default and only scales to the pricier one for the specific call that needs it, cost drops without capability dropping with it. It's the difference between paying for the headroom you use and paying for the headroom you might use.

Consumption reporting by area and by agent, visible every month. Whoever decides spending needs to see when the curve started climbing, which agent consumes the most, and which area accelerated. It isn't about cutting on reflex — it's about asking again whether that agent still pays off after three months running live. That's how Skyller was built: AI credits shared across the team, human approval matched to the risk of the action, automatic routing to the right model per task, and transparent cost per conversation.

The gain the metrics don't show

When budget is shared and consumption is visible, two gains show up that nobody planned for.

The first is financial and immediate: the company stops paying for idle seats. That unused slice stops being a fixed cost and becomes available capacity for whoever has work to do — no new purchase, no new budget request.

The second is architectural, and it's worth more over the medium term. When cost has a name — "this agent cost this much last month" — the right question surfaces on its own: could it run on a lighter model and only scale when needed? or does it really need to run every night, or are Monday and Wednesday enough? Nobody asks that when the license is already paid for and expiring anyway. With shared credits and per-agent reporting, it becomes a review routine — and it's that routine, not the cutting, that separates who cancels the project from who keeps it running past 2027.

Three questions to bring to your next meeting

Three questions to bring to your next meeting
  1. If each agent runs without a cap, how many calls per month is your company paying for without knowing? Start by asking for a real consumption report, not a forecast. If 78% of IT leaders have already been surprised by consumption-based charges, the odds of a surprise sitting inside your bill are high.
  2. What model does your team use for each task: the most expensive for everything, or the cheapest one that meets the accuracy you need? If the answer is "always the same one," you're paying, on every routine task, for headroom you don't use.
  3. If an agent asks for authorization before changing a critical document, who approves it — and how fast? Human approval is what prevents the expensive, irreversible action. But if it becomes a bottleneck, the team goes back to doing it by hand and the agent loses its value. A single queue for pending decisions solves both sides.

Discover Skyller