In May 2024, researchers at Stanford University published the first preregistered study on the reliability of AI tools built for legal research. The result: the two leading platforms on the market, built by major legal publishers, get something wrong or invented — a citation that doesn't exist, or a claim that a document says something it doesn't say — between 17% and 33% of the time. A general-purpose model, with no tuning for the task at all, got legal questions wrong up to 82% of the time.

The reason this matters beyond the legal field is simple: none of these answers arrived labeled "I'm not sure about this." They all came in the same confident tone, well-written, with the appearance of complete research. The difference between the right answer and the wrong one wasn't visible in the text — it only became visible to whoever checked the source.

That's the problem any company faces when it uses AI to answer questions about internal policy, a contract, or a process, even outside the legal field: a plausible answer with no verifiable origin isn't information. It's an opinion dressed up as data.

For companies in Brazil and Latin America that put AI to work answering questions about internal rules — a reimbursement policy, a vacation deadline, a clause in a contract —, the risk is the same at a smaller scale: every confident, wrong answer becomes a decision made on a foundation that doesn't exist.

The risk of a confident answer

Outside the courts, the same kind of error shows up in smaller, far more frequent decisions: an outdated reimbursement policy cited as if it were current, an expense cap that changed two months ago and that AI keeps repeating from the old version, a contract clause interpreted out of context from the paragraph right after it. None of these cases make the news. All of them cost money and time the same way, just quietly, with no one writing a story about it.

In March 2026, a federal court in Oregon ruled on a case that shows this risk outside the lab. Over five months, attorneys for one side filed three separate briefs containing 15 citations to cases that didn't exist and 8 "quotations" from rulings that never said any such thing. The court fined the responsible attorney $15,500, struck the briefs, and dismissed that side's claims with prejudice — the harshest penalty, reserved, in the court's own words, for someone who showed no remorse whatsoever when confronted.

The case isn't isolated. Researchers now maintain a public registry just to track court rulings where AI-generated content with invented citations reached a judge — a sign the problem has stopped being a rare incident and become a recognized pattern, with hundreds of cataloged cases.

The definition Stanford's researchers used helps explain why this keeps happening: a serious error isn't only information invented from nothing. It's also a real citation, from a real document, used to support something that document doesn't actually say. That second form is more dangerous because it survives any surface-level check — the source exists, it's only the link between it and what was claimed that's false.

Why tone isn't enough

Why tone isn't enough

The most common reaction to hearing this kind of data is to look for a "better" tool, or to trust that the newest model will solve the problem on its own. Stanford's own study shows the limits of that bet: the tools tested were built specifically for legal research, with access to real legal databases — and even so, one of them got it wrong in 1 out of every 3 answers. Switching AI models without changing the structure around it solves very little, because the problem isn't how much the model "knows." It's that nothing in the structure forces the answer to point back to where it came from.

The other common reaction is to rely on generic human review — "someone always checks it later." But review with no defined process tends to happen only after something has already gone wrong, not before. It's the same pattern seen in the fabricated-citation court cases: most were only caught because the opposing side, or the judge, manually checked every reference — work the structure should have done before the document ever went out.

What's missing in both cases is the same thing: a mandatory link between an answer and its origin, and a defined approval step before critical content becomes official knowledge that AI uses afterward.

What has to be in place

Corporate knowledge with an owner, a current version, and a defined scope. Every document that feeds AI has a person responsible for it, a version marked as current, and a defined scope — does it apply company-wide, or only to one area? Without that, a policy revoked two years ago competes on equal footing with the current one.

Answers that can cite the source used. When an answer points to which document it came from, whoever receives it can check it in seconds — the same check that, in the case of the legal tools, only a trained specialist could realistically do by hand.

Two-step review and approval for critical documents. Before a sensitive document becomes knowledge that AI uses to answer questions, it goes through whoever wrote it and whoever approves it, with separation of duties and logged exceptions. This doesn't mean approving every single AI-generated answer; it means approving the material that feeds those answers.

An audit trail of who approved what, and when. In an audit, or when a question comes up about a specific answer, reconstructing the origin in minutes is the difference between resolving the case and being unable to prove anything.

This is how Skyller was designed: corporate knowledge with a source, two-step review for critical documents, and answers that can cite where they came from.

The payoff of deciding on traceable ground

The payoff of deciding on traceable ground

When an answer can be checked in seconds, decision speed changes in practice: whoever receives it doesn't have to choose between blind trust and stopping everything to investigate on their own. The check becomes part of the workflow, not an extra step no one has time for.

There's also an effect that only shows up over time: when every answer carries its source, it becomes visible which documents are outdated, contradict each other, or are never actually consulted. The company gains a clarity over its own knowledge base that no manual audit reaches — because every AI query is, in practice, a real-world usage test of that document.

The cost of not having this rarely shows up all at once. It builds up in small decisions no one audits, until one of them turns into a problem big enough for someone to ask where that information came from — and the answer is "I don't know, the AI said so."

There's also the most direct payoff: no business decision should rest on an answer no one can trace back to its origin. That's just as true for a contract as it is for a reimbursement policy — the size of the document changes, the principle doesn't.

A roadmap to get started

  1. Start with the documents that generate the most doubt or the most sensitive decisions. There's no need to organize the entire knowledge base at once — start with the few documents that cost the most when they're wrong.
  2. Name an owner for each one. Without a named owner, no one updates the version or notices when it goes stale.
  3. Set up two-step review only for what's actually critical. A low-risk operational document doesn't need the same process as a financial or legal policy.
  4. Require that answers cite the source used. If the tool doesn't allow that, it shouldn't be answering anything the company treats as official.
  5. Run a real test before trusting it fully. Take everyday questions and manually check the answers for a few weeks — the same care that was missing in the fabricated-citation court cases.

Discover Skyller