Operations · 10 min

Microsoft 365, ChatGPT, and AWS all went down this week. What your business should have ready

By Xenith Editorial

In the space of a few days: a Microsoft 365 and Azure outage that took Teams and SharePoint down for hours, a global ChatGPT outage, and an AWS disruption that cascaded into consumer services. None of it was your fault, and all of it stopped work anyway.

The realistic goal is not zero downtime. You cannot make someone else's cloud more reliable. The goal is that an outage costs you two mildly annoying hours instead of a lost day and an angry client.

What happened, as of 26 July 2026

IncidentCause and scale
Microsoft 365 and Azure, 23 JulyMicrosoft attributed it to a bug in its automated network maintenance system that removed IP routes from more devices than intended, per BleepingComputer. Teams, SharePoint and other services disrupted for several hours
ChatGPT, 25 JulyGlobal outage affecting the app and APIs, with elevated error rates reported before a fix was applied
AWS disruptionCascaded into dependent consumer platforms, including gaming services

The Microsoft cause is the instructive one. It was not an attack or a data centre fire. It was routine automation doing slightly more than intended. That is the ordinary way large systems fail, and it is why "our provider is enormous and professional" is not a continuity plan.

Work out what actually stops

Most teams have never listed this, which is why outages feel chaotic. Fifteen minutes with a sheet of paper:

If this is downWhat stopsCan you work around it?
EmailClient communication, quotes, approvalsPhone and WhatsApp, if you have the numbers offline
Chat and meetingsInternal coordination, client callsA second platform, agreed in advance
Cloud file storageEverything, if nothing is localLocal copies of active work
Your website or app hostingSales and credibilityStatus page and a way to tell people
Payment gatewayCollecting moneyA second gateway account, kept live
An AI tool in a workflowThat workflowThe manual process it replaced
Internet at the officeAll of itMobile hotspot, and a documented switch

Anything in that table with no workaround is a single point of failure you have accepted without deciding to.

The plan that fits a small business

1. A second channel everyone already has

When Teams or Slack dies, coordination should not die with it. Agree the fallback in advance — usually WhatsApp or phone — and make sure numbers are stored somewhere that does not depend on the thing that is down. A contact list living only in the broken system is a familiar and avoidable trap.

2. Local copies of active work

Not your whole archive. Just what is being worked on this week. If cloud storage goes read-only for four hours, someone should still be able to finish the deliverable due tomorrow.

3. A written outage page

One page, reachable when your systems are not, listing: provider support contacts and account numbers, where the spare hotspot is, who tells clients, and the first three steps. During an incident nobody researches calmly. Same principle as the incident page in our security baseline.

4. Know where the truth is

Bookmark your providers' status pages and subscribe to their alerts. Ten minutes of "is it us or them" is ten minutes wasted, and Downdetector-style reports confirm scale quickly.

5. Tell clients before they ask

A short message — we are affected by a provider outage, here is what it means for your deadline, here is when we will update you — converts a frustrating experience into a demonstration of competence. Silence does the opposite.

6. A second internet connection

The cheapest continuity purchase if your work stops without internet. Covered in our small office network guide, along with why the network gear belongs on a UPS.

The five-minute test: imagine your main email and file platform is unavailable for the next four hours. Can your team still work, and can you still reach every client? If not, that gap is your entire action list.

What is worth paying for, and what is not

Worth it: a second internet connection, backups you have actually restored, a payment gateway kept live as a spare, and local copies of active files.

Usually not, at your size: multi-cloud architecture, hot standby infrastructure, or enterprise continuity software. These solve problems you do not have yet, and the money is better spent on the boring items above. Our software audit method is the way to check you are not already paying for resilience you never configured.

Two things this week specifically should prompt

If you built an AI step into a client workflow, the ChatGPT outage is your warning. Any workflow with an AI dependency needs a documented manual fallback, or a client-facing promise you cannot keep. Write down what happens when the model is unavailable.

If you host client sites or apps, decide now who tells the client during an outage and how. Handling that badly damages a relationship more than the downtime itself, and it belongs in the scope document alongside your other assumptions — see the scope guide.

The honest summary

Outages at this scale are not preventable by you, they are not rare, and they will happen again. Providers publish uptime figures that sound reassuring and still lose an afternoon to an automation bug. Plan for the afternoon, keep the plan on one page, and the next one becomes an inconvenience rather than a crisis.

No company paid for placement in this article. Verify current prices and terms with each provider before buying.