According to Microsoft's 2026 Work Trend Index, published in May, 86% of people who use AI at work say they treat the text it produces as a draft, not a finished answer. Most AI use today is exactly what it does best: writing first drafts of emails, summaries, and reports.
The problem shows up when the question changes shape. Writing a text and recommending a decision don't call for the same level of caution from the reader — and research shows that, faced with an answer that sounds certain, many people stop checking it.
For decision-makers, that distinction separates a real gain from a quiet risk: the same AI that saves time on drafting can get expensive when it weighs in on something that should have been verified first.
The cost of trusting without checking
In September 2023, researchers from Harvard, working with the consulting firm Boston Consulting Group, tested 758 junior consultants on two families of tasks: one inside what GPT-4 did well, product innovation, and one outside it, a business problem that required judgment and information the tool didn't have.
On the task inside the AI's competence, people who used it performed 40% better than those who didn't. On the task outside it, the result flipped: the group using AI performed 23% worse than the group that worked with no tool at all.
The reason wasn't a lack of warning. Participants knew the tool could get that kind of task wrong. Even so, they accepted wrong answers because the explanation the AI offered sounded convincing — trusting it felt easier than working through the reasoning themselves. Familiarity with the tool didn't protect anyone from this pattern: the issue wasn't a lack of practice, it was the way the answer itself arrived, polished and confident.
The phenomenon has a name in decision-support research: automation bias, the tendency to accept a system's output instead of independently evaluating the information. An academic review that gathered studies on decision-support systems found a 26% higher risk of a wrong decision when the system offered incorrect advice. In some of the studies reviewed, up to 7% of answers people had already gotten right were flipped to wrong after seeing the system's suggestion.
In Brazil and Latin America, where generative AI adoption has moved fast across operations, finance, and support teams, this is exactly the blind spot: the tool became a habit before the company decided what to do when it gets something wrong with confidence. The missing question isn't "does the team use AI?" — it's "when the AI suggests something wrong, is there a step that catches it before it becomes a decision?"
The warning that doesn't change the habit

The most common response has been to ask for more care: a banner on the screen, a best-practices training session, the generic instruction to "always double-check". The BCG experiment itself shows the limit of that approach — participants were warned about the tool's limitation on that specific task, and they still accepted the wrong answer.
The academic review on automation bias reaches the same conclusion from a different angle: asking for generic vigilance doesn't work well. What reduces the problem is changing what the person actually sees — showing the reasoning behind the answer, indicating how confident the system is in the recommendation, and presenting information for the person to decide with, rather than a ready-made suggestion to accept. The same review also lists holding the decision-maker accountable for the outcome and reducing the visual prominence of the suggestion, so it doesn't read as the only option on the screen.
In other words, the problem is rarely inattention. It's that the answer's design gives the person nothing concrete to check. Without a source, a document, or a data point backing the recommendation, verifying it becomes extra work that the day-to-day routine doesn't leave room for — and the shortest path becomes accepting what already arrived pre-packaged.
What has to be in place
An AI environment that supports decisions, not just drafts text, needs verifiable mechanisms — not another banner on the screen.
Every answer grounded in an internal document points to the source. Whoever receives the recommendation can see which policy, table, or process it came from, and can open the original before acting instead of blindly trusting the AI's explanation.
Recommendation and drafting are treated as different things. Writing an email draft is one thing; suggesting that a contract be renewed or a collection notice be sent is another — and the second calls for a confirmation step the first doesn't.
Human approval before a sensitive action. Critical documents and actions that change something in the real world can require a two-step review and approval, with the person who drafts kept separate from the person who approves, configurable by area rather than applied to every answer.
Access according to each person's role. Who can approve a recommendation, and on which topic, follows the same permission rule as any other company system, instead of being left open to whoever gets there first.
An audit trail of what was approved. If a recommendation became an action, there's a record of who approved it, when, and based on which source — exactly the material that's usually missing when an error needs to be investigated later.
Reuse with permission, not loose copying. A response template that already went through review can be made available to other people in the same role or group, instead of each person rebuilding their own shortcut with no one aware of what it contains.
This is how Skyller was designed: corporate knowledge with a citable source and human approval before any sensitive action, as the default, not the exception.
Less review, faster decisions

The gain shows up first in speed. When the answer already arrives with its source attached, the decision-maker spends less time hunting for the document that backs that policy and more time judging whether it applies to the case at hand.
The second gain is reconstructing what happened. In an audit, a customer complaint, or an internal investigation, the difference between reconstructing a decision in minutes and not being able to reconstruct it at all is the record of who approved what.
The third is cultural: when approval has an owner and a defined step, the team stops treating the AI's answer as the final word. It becomes an input — a good input, but an input, like any report or spreadsheet that arrives ready for critical reading.
There's also a compounding effect: a response template that someone approved once, with the right source attached, stays available for the next similar case. Instead of every decision starting from zero, the team builds up a set of already-checked recommendations — and that's where the time saved stops being individual and starts showing up in the whole area's results.
Questions to ask before trusting the answer
Before the next recommendation AI brings to the table, it's worth asking:
- Does the answer point to the source behind it? If there's no citable document, data point, or policy, treat it as an opinion, not a fact, and check it before passing it along.
- Does someone approve it before the recommendation becomes an action? If the answer can trigger a message, a collection notice, or a contract change on its own, a step is missing between the suggestion and the real-world effect.
- Has a convincing explanation already been mistaken for a correct one? The study with 758 consultants shows the two aren't the same — and that trained professionals fall into that confusion too.
- Is there a record of who approved what? Without that trail, a bad decision only surfaces after its cost has already been paid.
- Does the tool separate a text draft from a recommended action? If both arrive with the same air of certainty, the reader loses exactly the signal that should call for more attention.






