What happened
CrowdStrike's Falcon platform is endpoint detection and response software installed at the kernel level on Windows systems. On 19 July 2024, a content update contained a logic error that caused affected systems to enter a boot loop.
The outage was simultaneous and global. Airlines, hospitals, banks, broadcasters and emergency services were affected within minutes. Because affected machines could not boot normally, remediation commonly required manual intervention.
Delta Air Lines suffered the most prolonged airline disruption, cancelling about 7,000 flights over five days, affecting 1.3 million passengers and reporting at least $500 million in losses and costs.
Incident at a glance
- Windows systems affected
- 8.5M
- Delta flights cancelled
- ~7,000
- Delta reported impact
- $500M+
- Passengers disrupted
- 1.3M
- Remediation
- Manual intervention on affected machines
- Global sectors affected
- Aviation, banking, healthcare, broadcasting
The root cause
The CrowdStrike incident illustrated a category of third-party risk that traditional escrow does not solve by itself: faulty update risk. The vendor remained operational, but one bad update caused simultaneous failure across millions of systems.
The question for every institution is whether it has a verified prior release, a tested rollback target and an executable recovery procedure when critical vendor software fails after an update.
What would Proof of Recovery have changed?
- A current Proof of Recovery for the prior release would provide a verified rollback target.
- A signed deployment runbook would provide a systematic recovery procedure rather than an improvised response.
- A current SBOM and verified baseline would help security teams compare affected and known-good states.
- Physical access might still be required, but the recovery decision and procedure could be made faster.
“The vendor designed to protect critical systems became the source of failure. Recovery still depended on whether each organisation had a tested way back”
The regulatory consequence
The incident intensified scrutiny of ICT change-management and recovery testing. EU DORA, the UK PRA operational-resilience framework and APRA CPS 230 each reinforce the need to test recovery from critical technology failures, including rollback and third-party disruption.