In February 2024, the Civil Resolution Tribunal of British Columbia, Canada, decided a case that became a global reference on how companies use AI assistants. A customer asked Air Canada's website assistant about the airline's bereavement fare and got the wrong answer. He bought the ticket based on that answer and, when he asked for the promised discount, the company refused.

Air Canada argued that the assistant operated as its own separate entity. The tribunal rejected the argument and ordered the company to pay the fare difference: in its view, it makes no difference whether information comes from a static page or a conversational assistant — whoever published it is responsible for it.

Around the same time, courts in several countries were facing a similar problem inside legal proceedings: lawyers citing rulings that never existed, produced by an AI assistant and presented as real case law. Both situations share the same root. The wrong answer arrived with the same confidence as the right one, and nobody checked it before it became an official communication or evidence on the record.

The Answer Always Arrives With the Same Confidence

The Air Canada case was not isolated. In the United States, a public tracker maintained by researcher Damien Charlotin has already documented more than 1,400 court decisions in which an AI-produced citation or piece of evidence contained a factual error — most of them references to cases that simply do not exist. According to the researcher, quoted by Scientific American, the pace of new cases has settled at around 350 to 400 per quarter.

The reason behind that persistence shows up in a study by Stanford RegLab and the Stanford Institute for Human-Centered AI on AI tools built for legal research. Specialized tools, from vendors that promised hallucination-free answers, were wrong between 17% and 34% of the time. General-purpose assistants, with no adjustment for the law, were wrong between 58% and 82% of the time on the legal questions tested.

The detail that matters most in that study is not the number — it is the pattern of the error. The wrong answer did not arrive with hesitation or a caveat. It came formatted as a real citation, complete with a case name, a year, and a quoted passage, exactly as a correct answer would look. Whoever reads it has no clue to be suspicious without going back to the original source.

Inside a company, the same pattern repeats at a smaller scale, but with the same mechanics. A question about an internal policy, a contract, or an HR procedure gets a fluent, specific answer in exactly the right tone — just without saying where it came from. Without that reference, accepting it and checking it take exactly the same amount of work.

Warnings and Good Intentions Are Not Enough

Warnings and Good Intentions Are Not Enough

The most common reaction to this risk is to ask for common sense: "double-check before using it," "don't take it literally," "always review it." The problem is that checking an answer takes, in practice, the same work the person tried to avoid by asking the assistant in the first place. Without an attached source, "checking" means redoing the research from scratch.

The Air Canada case shows exactly why a warning protects no one. The customer trusted the information because it came through the company's official channel — the same way he would trust any other page on the site. A disclaimer saying the assistant "may make mistakes" does not change what the person actually receives: an answer that looks exactly like any other piece of company information.

The same logic applies to anyone citing case law. No sanctioned lawyer has ever claimed not to know that an AI assistant can invent things — courts have already warned about it, in public rulings, more than once. Even so, the pattern repeats, because the warning does not fix the structural problem: nothing in the answer itself signals that it was invented.

Banning the tool does not solve it either, for the opposite reason: the task still has to get done, just now with no record of how it was done. The right question is not whether the assistant will get something wrong — it will, at the rate the studies themselves document. The question is what stands between the answer and the action that depends on it.

What Is at Stake for Decision-Makers

The precedent set by Air Canada is not limited to airlines. It establishes something simpler and far broader: what an AI assistant says on a company's behalf is, for practical purposes, what the company said — whether that is an HR policy, a commercial term, or an answer to a customer.

It should be obvious to Air Canada that it is responsible for all the information on its website, regardless of whether it comes from a static page or a chatbot.

Civil Resolution Tribunal of British Columbia, 2024

For lawyers, the consequence has already shown up in concrete rulings: fines, dismissed claims, suspension from practice. In one U.S. case, an attorney was sanctioned for citing nonexistent cases — and cited another nonexistent case in the very next sentence of the same filing, after already having been caught. The error was not the assistant's. It was the absence, somewhere in the workflow, of a point where someone checked the source before signing off.

Outside the courtroom, the effect is the same on a smaller scale: a business decision made on a contract clause the assistant "remembered" wrong, an external communication built on a number that appears in no company document. The cost does not always make headlines. But it lands on whoever signed off — not on the assistant that answered.

What Has to Be in Place

What Has to Be in Place

A governed AI environment answers this risk with verifiable mechanisms, not with warnings or trust.

Knowledge with an owner and a traceable source. The answer draws on internal documents that have an owner and a current version, and it can point to which document it came from, for anyone who wants to check before using it.

Human approval before a sensitive action. Critical documents may require two-step review and approval, with segregation of duties and audited exceptions, before becoming official knowledge or going out in an external communication.

Access based on each person's role. Anyone consulting legal, financial, or HR information sees only what their own role authorizes, which limits how many people can publish something without a check.

An audit trail. Every answer backed by internal knowledge is logged: who asked, what was used as the basis, and when it happened.

Reuse with permission. A response script already reviewed and approved by one person can be reused by others, under the same access controls, instead of each person redoing the same risky question alone.

This is how Skyller was designed: corporate knowledge with a stated source, human approval before a sensitive action, and an audit trail as the default, not as an extra setting.

From a Loose Answer to Data You Can Check

The most immediate gain of an environment like this is time. When an answer arrives with its source attached, checking it takes seconds, not a stalled investigation: the person reads the referenced document, confirms the citation is correct, and moves on.

There is also a gain when something has to be reconstructed later. On a platform like Skyller, every answer backed by internal knowledge keeps the source document and a record of who asked — turning a reconstruction of the facts into a lookup, not an investigation.

That changes the role of whoever reviews the answer. Instead of choosing between blind trust in the assistant and distrust of everything, the person examines one specific, concrete source — the same task they already know how to do, just without redoing the research from scratch.

Questions to Bring to Your Next Meeting

Before drafting a new policy on AI use, it is worth bringing these questions to the next meeting with IT and business leadership.

  1. When the assistant answers something about a contract or an internal policy, where did that information come from? If no one can point to the source document, the answer cannot be checked before it becomes a decision.
  2. Is there an approval step before an answer becomes an external communication or a sensitive action? Without that checkpoint, the first person to notice a mistake may be the customer — or a court.
  3. If a mistake reaches a customer or a legal case, can you reconstruct who asked what, and on what basis? Without an audit trail, the answer is always "we don't know."

Discover Skyller