On the morning of August 31, 2026, Outlook and Microsoft 365 (formerly Office 365) stopped sending and receiving email for a number of businesses that Microsoft never fully disclosed. Downdetector, the service that tracks outages in real time, logged more than 5,000 problem reports by early afternoon alone, according to TechCrunch. Hours earlier, Microsoft had already acknowledged the issue internally, tracking it under the code EX1464935.

The cause, according to the company itself, was a failure in an authentication component used by the Exchange Online infrastructure — the service behind Microsoft 365 inboxes. The problem did not stay confined to email: it also reached Teams, SharePoint, OneDrive and other services that share the same corporate login. Inbox access returned to normal for most users about 12 hours later, but some features only fully recovered three days after, on September 3.

For people who work in technology, it was one more logged incident. For the company that morning that could not send an invoice, confirm an order, or answer a client chasing a deadline, it was a lost day of work — and the question left behind was why no one knew what to do.

That question is what matters to whoever makes decisions in a company, not the technical detail of the failure. The business owner doesn't need to understand what an authentication component is. They need to know what the team does in the first hours of an outage, and that is exactly what tends to be missing.

When the email provider goes down

It wasn't an isolated episode. On May 13, 2026, a different failure knocked Outlook out specifically across South America: Downdetector received 711 problem reports for Outlook and more than 270 for Microsoft 365 within a few hours, according to Canaltech. Microsoft attributed the incident to a portion of its own network infrastructure in the region, which caused intermittent disruptions for users in Brazil and neighboring countries.

The two episodes share something in common: the cause was neither an attack nor a user error. It was a failure inside the provider, the kind that no contract eliminates entirely — it only lowers the odds and shortens the time to recovery. That holds for any cloud provider, not only Microsoft, and it also holds for a company that keeps email on a server of its own: the point isn't which path is safer, it's that neither one is failure-proof.

What makes these outages different from any other IT hiccup is what travels through email: client orders, invoices, signed contracts, billing notices and, often, the password-reset link for other systems the company relies on daily. When it stops, it isn't just one tool going dark. It's the channel through which the company talks to the people who pay the bills.

And email rarely fails loudly enough to announce itself. It stops working quietly, with no warning on the screen, while someone keeps typing a reply that will never go out — and the problem is only noticed when the client calls asking why nothing arrived. That gap between the real failure and the company noticing it is often the most expensive part of the whole incident.

Why nobody knew what to do

Why nobody knew what to do

When email goes down, the first problem is rarely technical — it's a decision problem. Nobody can say whether the failure is the internet provider, the cloud provider, or just one person's computer, and each hypothesis calls for a different fix. While the doubt lingers, nobody calls anybody, and the client on the other end goes unanswered.

A continuity strategy has to account for different failure scenarios.

Penso Tecnologia, 2026

The underlying problem is usually the lack of a plan. According to a survey by US insurer Nationwide with small and mid-size business owners, published in February 2025, 21% of them have no business continuity plan at all — even though nearly 90% regularly review other risk policies. The gap sits exactly where it costs the most: in the middle of a real emergency.

Relying only on the cloud provider to sort everything out is another common mistake. The provider takes care of its own infrastructure, but it has no idea who inside the company needs to be notified first, or which order is sitting there waiting for confirmation. That part is always on the company — plan or no plan.

It also doesn't help to wait for a problem to figure out who to call. Many small businesses only have an IT contact for when something has already broken, with nobody watching the network beforehand — and it's exactly that lack of ongoing monitoring that turns a few-hour outage into a full day of uncertainty.

What has to be in place

A well-run IT setup doesn't prevent every outage; it shortens the confusion and the downtime when one hits. That depends on concrete mechanisms, not luck:

An alternate channel, already agreed on. Phone, a business WhatsApp line, or a notice on the website — decided ahead of time, not improvised on the spot, to tell clients and vendors that email is down and their request wasn't lost.

Someone who confirms the cause in minutes, not hours. One single owner, in-house or outsourced, who can tell the difference between an internet outage, a cloud provider failure, or one specific computer, instead of the whole team standing around guessing.

Shared accounts and cc'd lists, not lone individual inboxes, for orders, billing and contracts — so information doesn't sit trapped in a single inbox that's down.

Monitoring that spots the outage before the team does. The sooner someone knows, the sooner the backup channel kicks in.

A logged ticket with cause and resolution for every incident, so the same problem doesn't catch the company off guard next time.

A simple inventory of what the company uses and what depends on what. Without it, nobody knows off the top of their head what else goes down along with email — and the list of who to notify first gets built at the worst possible moment.

This is how Skills IT works: constant network monitoring, a single point of contact to confirm the cause of a problem, and every ticket logged — so the next outage finds the company more ready than the last one.

The payoff of deciding this before the outage

The payoff of deciding this before the outage

The return on having this plan ready doesn't show up only on the day of the incident — it shows up in the time the company stops losing trying to figure out what to do. With a backup channel already agreed on, the client gets a heads-up within minutes, instead of concluding on their own that they've been ignored. With someone quickly confirming the cause, the team moves on to other work while email is down, instead of standing idle waiting for a verdict.

It's also a financial matter, even without an exact figure: every hour of indecision is an hour of idle staff, unconfirmed orders, and rework once everything is back up and the backlog needs sorting out. A company with a plan decides fast; a company without one decides later, under pressure, with the client on the other end of the line.

There's also a gain that only shows up the second time around. A company that has already logged how it resolved the last outage doesn't start from zero on the next one — it knows who to notify, who to call to confirm the cause, where to look first. The lesson from one incident becomes a shorter response time on the next, and that's what separates the company that treats every outage as a novelty from the one that treats it as an already-mapped routine.

A playbook for the day email goes down

Nobody writes a continuity plan in the middle of an outage. It needs to exist beforehand, even if it's simple:

  1. Pick the backup channel now. Phone, WhatsApp, or a website notice — agree on which one and where it's written down, so it isn't decided while the problem is happening.
  2. Decide who confirms the cause. One person, or one provider, who knows how to check whether it's the internet, the cloud provider, or a specific computer, instead of leaving it to "whoever figures it out first."
  3. Put critical processes on cc'd accounts. Orders, billing and contracts shouldn't depend on a single inbox.
  4. Write down the first three steps of the next outage. Who notifies the client, who confirms the cause, who logs what happened — in writing, today, not from memory later.
  5. Review the plan after every incident, even the small ones. That's what separates the company that learns from the one that repeats the same scare.