Microsoft Office 365 suffers global outage

Microsoft Office 365 suffers global outage

Thousands of Microsoft Office 365 users experienced issues due to a global outage.

Australia appeared to be one of the worst hit countries, with many users unable to send and receive emails or log into their Outlook accounts.

The incident

The issue was caused by a global outage within Microsoft Office 365, designated as incident EX103663.

Microsoft acknowledged the issue and worked towards resolution. During the outage, users were unable to connect to their mailbox from:

blueAPACHE monitored the situation and provided updates to clients throughout.

Resolution

Update: Microsoft confirmed the issue was resolved and services restored.

The root cause was identified as a code issue in a regularly scheduled deployment.

If you continue to experience any issue, please contact your account manager.

What the root cause tells you

That closing detail is the most instructive part of the incident, and it is usually the part that gets skipped.

The outage was not caused by a cyber attack, a data centre failure, a fibre cut or a natural disaster. It was caused by a routine deployment.

This is the dominant failure mode in large-scale cloud services. Hyperscale platforms are engineered against hardware failure comprehensively — redundant power, redundant network, redundant compute, geographic distribution. What they cannot engineer away is the fact that software has to be changed, and that change carries risk regardless of how carefully it is staged.

The practical implication for a customer is uncomfortable but worth being clear about: the redundancy you are paying for does not protect against this class of event. Geographic redundancy protects against a facility failing. It does not protect against a code change deployed everywhere, because the same faulty code reaches every region.

The shared responsibility question

Outages of this kind expose a gap that most organisations have not thought through until it happens.

Microsoft is responsible for the availability of the service. Your organisation remains responsible for continuing to operate while the service is unavailable. Those are different obligations, and the second one is yours whether or not you have planned for it.

A few questions that are easier to answer before an outage than during one:

  1. How does the business communicate when email is down? If the answer is Teams, note that Microsoft 365 service incidents can affect multiple services at once.
  2. Which processes depend on email as a system of record, rather than as a convenience? Approvals, notifications, customer commitments with response-time obligations.
  3. Who is authorised to invoke a fallback, and what is it?
  4. How do you tell customers? If your only channel to them is the one that is down, you have a problem that is not technical.
  5. What is the escalation path — to Microsoft, and to your managed service provider?

What a managed service provider adds during a vendor outage

An MSP cannot fix a Microsoft outage. Nobody outside Microsoft can. What it can do is remove the two things that make an outage worse than it needs to be.

Diagnosis. The first twenty minutes of any outage are spent establishing whether the problem is yours or the vendor's. Is it our mail flow rules? Our connectors? Our licences? A firewall change from last night? That time is entirely wasted if someone is already monitoring the vendor's service health and can say immediately that it is global.

Communication. During the incident above, blueAPACHE monitored the situation and issued updates. A single authoritative source of "this is what is happening, this is who is affected, this is what we know about restoration" prevents an internal IT team from answering the same question two hundred times.

Neither is glamorous. Both materially reduce the cost of the event.

Where email continuity fits

This is the case for an email continuity and archiving service, which is otherwise easy to file under compliance.

A continuity service maintains an independent path for sending and receiving mail when the primary platform is unavailable, and it earns its cost the first time that happens. blueAPACHE's partnership with Mimecast covers exactly this ground — email security, archiving and continuity — and the company has been nominated for Customer Excellence and Partner Innovation awards in that programme.

Reading a service level agreement honestly

A final note, because outages tend to prompt questions about SLAs.

SaaS availability commitments are typically expressed as monthly uptime percentages with financial service credits as the remedy. Those credits are calculated against the subscription fee, not against your business losses — and they are usually a small proportion even of that.

This is not a criticism of any particular vendor; it is how the entire category is contracted. The correct conclusion is that an SLA is a statement of intent and a commercial penalty, not a business continuity plan. If a service being unavailable for several hours would cause material harm to your organisation, the mitigation has to come from your own architecture and process, not from the contract.

Frequently asked questions

What happened in this incident? A global Microsoft Office 365 outage, designated incident EX103663, left users unable to connect to their mailbox from Outlook desktop, Outlook on the web, mobile devices and other mail protocols. Australia appeared to be one of the worst hit countries. Microsoft subsequently confirmed the issue was resolved and services restored.

What caused it? A code issue in a regularly scheduled deployment — not a cyber attack, a data centre failure, a fibre cut or a natural disaster.

Why does that root cause matter so much? Because it is the dominant failure mode in large-scale cloud services. Hyperscale platforms are engineered comprehensively against hardware failure through redundant power, network, compute and geographic distribution. What they cannot engineer away is that software has to be changed, and change carries risk however carefully it is staged.

Does geographic redundancy protect against this? No, and the page is blunt about it. Geographic redundancy protects against a facility failing. It does not protect against a code change deployed everywhere, because the same faulty code reaches every region.

Who is responsible for what during a vendor outage? Microsoft is responsible for the availability of the service. Your organisation remains responsible for continuing to operate while the service is unavailable. Those are different obligations, and the second is yours whether or not you planned for it.

What should an organisation decide before an outage rather than during one? How the business communicates when email is down — noting that a Microsoft 365 incident can affect multiple services at once, so "we'll use Teams" may not hold; which processes depend on email as a system of record rather than a convenience; who is authorised to invoke a fallback and what it is; how customers are told if the only channel to them is the one that is down; and what the escalation path is, both to Microsoft and to the managed service provider.

What can a managed service provider actually do? Not fix the outage — nobody outside Microsoft can. What it removes is the first twenty minutes spent establishing whether the fault is yours or the vendor's, and the burden of communication, by providing one authoritative account of what is happening, who is affected and what is known about restoration.

Where does email continuity fit? A continuity service maintains an independent path for sending and receiving mail when the primary platform is unavailable, and earns its cost the first time that happens. blueAPACHE's Mimecast partnership covers email security, archiving and continuity.

Does an SLA protect against this? Not in the way people assume. SaaS availability commitments are typically monthly uptime percentages with financial service credits as the remedy, calculated against the subscription fee rather than business losses — and usually a small proportion even of that. The page's conclusion is that an SLA is a statement of intent and a commercial penalty, not a business continuity plan.

Related

Contact

If you continue to experience issues following any service incident, contact your blueAPACHE account manager.

Knowledge Base

What happened during the Microsoft Office 365 outage described on blueAPACHE's blog?

Thousands of Microsoft Office 365 users experienced issues due to a global outage, with Australia appearing to be one of the worst hit countries. Many users had trouble sending and receiving emails and logging into their Outlook accounts.

What was the designated incident number for this Microsoft Office 365 outage?

The issue was caused by a global outage within Microsoft Office 365 designated as incident EX103663.

What services were affected by the outage?

Users may have been unable to connect to their mailbox from Outlook, Outlook web, mobile devices, or other protocols.

Did Microsoft acknowledge the outage, and was there an estimated fix time?

Microsoft acknowledged the issue and was working towards a resolution, but there was no clear indication of when it was likely to be fixed at the time of the initial report.

Was the Microsoft Office 365 outage eventually resolved?

Yes. An update on the blog confirmed that Microsoft resolved the issue and restored services.

What was the root cause of the outage according to Microsoft?

The root cause was identified as a code issue in a regularly scheduled deployment.

When was this blog post about the Microsoft Office 365 outage published?

The blog post was written by blueAPACHE and dated May 31, 2017, with a 1-minute read time.

What did blueAPACHE do while the outage was ongoing?

blueAPACHE monitored the situation and continued to provide updates to readers.

What should someone do if they continue experiencing issues related to this outage?

The blog advises affected users to contact their account manager if they continue to experience any issues.

Images on This Page