Microsoft 365, ChatGPT, and AWS all went down this week. What your business should have ready
In the space of a few days: a Microsoft 365 and Azure outage that took Teams and SharePoint down for hours, a global ChatGPT outage, and an AWS disruption that cascaded into consumer services. None of it was your fault, and all of it stopped work anyway.
What happened, as of 26 July 2026
| Incident | Cause and scale |
|---|---|
| Microsoft 365 and Azure, 23 July | Microsoft attributed it to a bug in its automated network maintenance system that removed IP routes from more devices than intended, per BleepingComputer. Teams, SharePoint and other services disrupted for several hours |
| ChatGPT, 25 July | Global outage affecting the app and APIs, with elevated error rates reported before a fix was applied |
| AWS disruption | Cascaded into dependent consumer platforms, including gaming services |
The Microsoft cause is the instructive one. It was not an attack or a data centre fire. It was routine automation doing slightly more than intended. That is the ordinary way large systems fail, and it is why "our provider is enormous and professional" is not a continuity plan.
Work out what actually stops
Most teams have never listed this, which is why outages feel chaotic. Fifteen minutes with a sheet of paper:
| If this is down | What stops | Can you work around it? |
|---|---|---|
| Client communication, quotes, approvals | Phone and WhatsApp, if you have the numbers offline | |
| Chat and meetings | Internal coordination, client calls | A second platform, agreed in advance |
| Cloud file storage | Everything, if nothing is local | Local copies of active work |
| Your website or app hosting | Sales and credibility | Status page and a way to tell people |
| Payment gateway | Collecting money | A second gateway account, kept live |
| An AI tool in a workflow | That workflow | The manual process it replaced |
| Internet at the office | All of it | Mobile hotspot, and a documented switch |
Anything in that table with no workaround is a single point of failure you have accepted without deciding to.
The plan that fits a small business
1. A second channel everyone already has
When Teams or Slack dies, coordination should not die with it. Agree the fallback in advance — usually WhatsApp or phone — and make sure numbers are stored somewhere that does not depend on the thing that is down. A contact list living only in the broken system is a familiar and avoidable trap.
2. Local copies of active work
Not your whole archive. Just what is being worked on this week. If cloud storage goes read-only for four hours, someone should still be able to finish the deliverable due tomorrow.
3. A written outage page
One page, reachable when your systems are not, listing: provider support contacts and account numbers, where the spare hotspot is, who tells clients, and the first three steps. During an incident nobody researches calmly. Same principle as the incident page in our security baseline.
4. Know where the truth is
Bookmark your providers' status pages and subscribe to their alerts. Ten minutes of "is it us or them" is ten minutes wasted, and Downdetector-style reports confirm scale quickly.
5. Tell clients before they ask
A short message — we are affected by a provider outage, here is what it means for your deadline, here is when we will update you — converts a frustrating experience into a demonstration of competence. Silence does the opposite.
6. A second internet connection
The cheapest continuity purchase if your work stops without internet. Covered in our small office network guide, along with why the network gear belongs on a UPS.
What is worth paying for, and what is not
Worth it: a second internet connection, backups you have actually restored, a payment gateway kept live as a spare, and local copies of active files.
Usually not, at your size: multi-cloud architecture, hot standby infrastructure, or enterprise continuity software. These solve problems you do not have yet, and the money is better spent on the boring items above. Our software audit method is the way to check you are not already paying for resilience you never configured.
Two things this week specifically should prompt
If you built an AI step into a client workflow, the ChatGPT outage is your warning. Any workflow with an AI dependency needs a documented manual fallback, or a client-facing promise you cannot keep. Write down what happens when the model is unavailable.
If you host client sites or apps, decide now who tells the client during an outage and how. Handling that badly damages a relationship more than the downtime itself, and it belongs in the scope document alongside your other assumptions — see the scope guide.
The honest summary
Outages at this scale are not preventable by you, they are not rare, and they will happen again. Providers publish uptime figures that sound reassuring and still lose an afternoon to an automation bug. Plan for the afternoon, keep the plan on one page, and the next one becomes an inconvenience rather than a crisis.
No company paid for placement in this article. Verify current prices and terms with each provider before buying.