General
July 2026 brought significant outages across AWS, Microsoft Azure, Google Cloud, and Cloudflare. Learn what caused these incidents, why they matter to MSPs, and how to build a stronger strategy for monitoring, business continuity, and vendor risk management.

Cloud platforms are designed to improve availability through redundancy, but they are not immune to failure. July 2026 demonstrated that even the world's largest cloud providers can experience regional connectivity problems, power and cooling failures, and configuration-related disruptions.
For managed service providers (MSPs), these incidents are more than headlines. They show how quickly a vendor issue can become a client-facing outage when multiple organizations rely on the same cloud region, identity provider, storage platform, or network path.
The lesson is not to avoid cloud services. Instead, MSPs should treat vendor availability as an external dependency that requires planning, monitoring, testing, and recovery.
One of the month's most visible incidents affected Amazon Web Services (AWS). According to the AWS Health Dashboard, AWS experienced connectivity issues affecting access to the US-West-2 (Oregon) region on July 24. While the incident was relatively brief, it affected connectivity to multiple AWS services within the region, illustrating how a regional networking issue can impact multiple workloads simultaneously.
Microsoft Azure experienced a longer disruption on July 23. Microsoft's Preliminary Post Incident Review for the West US region reported connectivity failures, elevated latency, and difficulty accessing Azure resources and other Microsoft cloud services hosted in West US. The incident lasted from 14:44 UTC to 19:41 UTC. Microsoft determined that a maintenance-related process incorrectly removed more network routes than intended, affecting traffic entering and leaving the region. Workloads operating entirely within West US continued functioning normally.
Google Cloud experienced one of the month's most significant infrastructure failures. Google's completed incident report describes a power disturbance in the europe-west4-a zone that led to failures in backup power transfer and cooling infrastructure. The resulting outage lasted 14 hours and 55 minutes and affected Google Cloud VMware Engine, Bare Metal Solution, and Google Cloud NetApp Volumes. Google reported that the event affected 24 private clouds belonging to 20 VMware Engine customers, nine Bare Metal Solution customers, and six NetApp clusters that shut down automatically to protect equipment from high temperatures.
Google also reported an extended VMware Engine networking incident on July 14. According to the official Google Cloud incident record, a network configuration update interrupted communication between stretched cluster zones, resulting in inter-site communication failures, BGP session instability, and loss of connectivity to witness appliances. Google restored service after rolling back the network configuration.
Cloudflare also experienced a storage platform incident. The company's official incident history reports that some R2 customers in the ENAM region received HTTP 500 errors between July 20 and July 21. Although Cloudflare did not publish detailed customer impact figures or a public root cause, the incident demonstrates that storage services can also become a point of failure for hosted applications.
Although each incident had a different root cause, several common themes emerged.
First, failures occurred across different layers of cloud infrastructure. Some incidents involved networking, while others originated from power distribution, cooling systems, or configuration changes.
Second, regional outages can affect many services simultaneously because they often share underlying infrastructure.
Finally, healthy cloud resources do not always mean users can reach them. Connectivity problems, routing failures, or identity issues can leave applications unavailable even when servers continue running normally.
For MSPs, these patterns reinforce the importance of preparing for failures outside their direct control.
A single vendor outage can quickly impact dozens or hundreds of client environments.
Many MSPs standardize on a common cloud provider, backup platform, identity service, DNS provider, or storage platform. While this simplifies administration, it also creates shared dependencies. When one of those services experiences an outage, multiple clients may be affected at the same time.
The NIST Cybersecurity Framework (CSF) 2.0 provides a practical structure for managing this type of operational risk through its six functions: Govern, Identify, Protect, Detect, Respond, and Recover.
Applying these functions means identifying vendor dependencies before an outage occurs, monitoring for service degradation, coordinating incident response, and validating recovery after services return.
The NIST Contingency Planning Guide for Federal Information Systems further recommends performing business impact analyses, defining recovery priorities, documenting alternate operating procedures, and regularly testing contingency plans. Although originally written for federal information systems, these planning principles remain highly applicable to modern MSP environments.
One of the most valuable preparation activities is documenting vendor dependencies across the client portfolio.
The NIST Cybersecurity Supply Chain Risk Management Practices for Systems and Organizations recommends identifying, assessing, and managing risks associated with suppliers, products, and services throughout their lifecycle.
For an MSP, a dependency map should identify every critical third-party service supporting each client, including cloud platforms, regions, identity providers, DNS services, backup platforms, remote management tools, security solutions, internet providers, messaging platforms, storage services, and business-critical integrations.
This visibility allows teams to quickly determine which clients may be affected when a vendor reports an incident.
The ENISA Good Practices for Supply Chain Cybersecurity similarly recommends structured supplier identification, risk assessment, monitoring, contractual planning, and coordinated incident response. Together, these practices reduce uncertainty during major service disruptions.
Vendor status pages provide valuable information, but they should not be the only source of operational visibility.
Monitoring should verify the complete user experience, including application availability, authentication, DNS resolution, API connectivity, storage operations, and endpoint behavior from outside the affected cloud environment.
Platforms such as Level can complement cloud monitoring by providing visibility into endpoint health during broader service disruptions. Endpoint monitoring cannot determine whether a cloud provider is experiencing an outage, but it can help MSPs distinguish widespread infrastructure issues from endpoint-specific or network-specific problems.
The objective is to understand what clients are actually experiencing rather than relying solely on infrastructure status indicators.
Cloud resilience depends on deliberate architecture and operational planning.
The AWS Well-Architected Framework Reliability Pillar emphasizes recovery testing, automation, and designing workloads to recover from failures. Microsoft's Azure Well-Architected Framework reliability guidance focuses on resilient architectures and clearly defined availability objectives. Google's Cloud Architecture Framework for Reliability highlights observability, failure planning, and resilient system design.
Together, these guidance documents encourage organizations to assume failures will occur and to validate recovery processes before production incidents happen.
MSPs should verify whether critical workloads can tolerate the loss of a service, an availability zone, or an entire region. They should also confirm that backups, recovery documentation, monitoring systems, and administrative access remain available throughout a vendor outage rather than depending on the same affected platform.
Recovery plans are only valuable if they have been tested.
The CISA Tabletop Exercise Packages provide practical scenarios and templates that organizations can adapt to evaluate incident response procedures.
An MSP exercise might simulate:
These exercises should include engineers, service desk personnel, leadership, account managers, and client communication teams. Technical recovery is only one component of incident response. Communication, prioritization, escalation, and decision making are equally important.
The NIST Incident Response Recommendations and Considerations for Cybersecurity Risk Management integrates incident preparation, detection, response, and recovery into the broader Cybersecurity Framework lifecycle, making it applicable to operational outages as well as cybersecurity incidents.
Every significant outage should result in operational improvements.
After services have recovered, MSPs should:
The ENISA Technical Implementation Guidance on Cybersecurity Risk-Management Measures reinforces the importance of business continuity, supplier management, change control, rollback procedures, and validating that recovery measures function as intended.
July's outages originated from very different causes, including networking, maintenance processes, power infrastructure, cooling systems, and storage services. No single technical solution would have prevented every disruption.
Instead, industry guidance consistently points toward layered resilience through planning, monitoring, supplier management, recovery testing, and continuous improvement.
Organizations that prepare for vendor outages are generally better positioned to reduce business disruption, communicate effectively with clients, and restore services in a controlled manner.
For MSPs, resilience is not measured by preventing every outage. It is measured by how effectively the organization prepares for, responds to, and recovers from events that are outside its direct control.
At Level, we understand the modern challenges faced by IT professionals. That's why we've crafted a robust, browser-based Remote Monitoring and Management (RMM) platform that's as flexible as it is secure. Whether your team operates on Windows, Mac, or Linux, Level equips you with the tools to manage, monitor, and control your company's devices seamlessly from anywhere.
Ready to revolutionize how your IT team works? Experience the power of managing a thousand devices as effortlessly as one. Start with Level today—sign up for a free trial or book a demo to see Level in action.