CRISISTABLETOP
Rows of servers representing a mission-critical SaaS platform
← Exercise library

COMMERCIAL LEGAL INCIDENT · OPERATIONAL RESILIENCE

Significant SaaS Platform Outage

A routine cleanup script deletes hundreds of live customer tenants. Monitoring treats the destructive calls as valid, support records disappear with the service, and recovery designed for one tenant cannot scale. Over two weeks, the provider must prioritize customers, choose between availability and integrity risk, satisfy divergent contracts, support regulated customers, assess disclosure, respond to leaked warnings, and rebuild trust.

2.5 HRSuggested duration12Participant roles11Scenario stages
Stock photography via Unsplash

EXERCISE PURPOSE

Examine the provider's ability to coordinate technical recovery with customer commitments, continuity, regulatory support, disclosure, insurance, communications, remedies, and governance during a significant SaaS outage.

An approved maintenance action deletes hundreds of production tenants from a mission-critical SaaS platform, causing a prolonged outage, risky restoration, contractual disputes, regulatory scrutiny, claims, and loss of customer trust.

DECISION PRESSURE

Questions the team must answer together

The facilitator releases facts in stages. Participants should identify the decision owner, the authority being used, the information still needed, the immediate action, and the next escalation point.

01

Who can stop a destructive change and declare the highest-severity incident?

02

How should customers be prioritized when recovery cannot meet every commitment?

03

When is a faster restoration too risky for customer data integrity?

04

Which contractual rights and commercial remedies can be promised during the event?

05

What should customers, regulators, insurers, directors, investors, and media be told as estimates change?

INSIDE THE EXERCISE

A scenario that changes as the response develops

The template is already populated with public injects, facilitator-only context, discussion prompts, private role messages, and a final hotwash.

  1. Quarter 2
    Module 1 · Resilience before the incident

    A mission-critical platform grows faster than its controls

    Northstar Cloud provides a multitenant workflow and records platform used by financial institutions, manufacturers, health organizations, and public companies. Revenue has doubled in two years. Customer contracts promise 99.95% availability, regional hosting, daily backups, incident notice, and recovery objectives. Operations still depend on a shared global control plane and several privileged maintenance tools.

  2. Monday · 07:30 UTC
    Module 1 · Detection and escalation

    Routine cleanup targets live customer tenants

    An infrastructure team runs an approved script intended to remove retired feature instances. The input file contains tenant identifiers rather than feature identifiers. The deletion API accepts either type, performs hard deletes, and does not display the object type or require a second production confirmation. Hundreds of live tenants begin disappearing.

  3. Monday · 08:20 UTC
    Module 1 · Detection and escalation

    Scope is known but customer support is impaired

    Engineering stops the script after 846 tenants are deleted across 712 customers. The support portal, approved-customer contact directory, and several incident runbooks are hosted in the same environment. Customers cannot authenticate to open tickets, and response teams cannot retrieve every contract or escalation contact.

  4. Monday · 11:00 UTC
    Module 1 · Detection and escalation

    Recovery will take days, not hours

    Backups are intact, but restoration requires rebuilding identity, tenant configuration, application data, attachments, integrations, encryption references, search indexes, and audit history in a strict sequence. The tested process handles one tenant at a time. Engineering estimates that some customers may be unavailable for more than a week.

  5. Monday · 16:00 UTC
    Module 1 · Detection and escalation

    Customers demand rights the company cannot yet deliver

    Customers request database copies, backup images, audit logs, proof of deletion scope, restoration sequencing, and permission to rebuild in competing services. Some invoke disaster-recovery assistance, audit, step-in, business-continuity, and termination provisions. Sales asks whether account teams can promise service credits and customized recovery dates.

  6. Tuesday · 09:00 UTC
    Module 2 · Recovery and scrutiny

    A fast restore could overwrite customer changes

    Engineers build a parallel restoration pipeline that may cut recovery time in half. Validation finds a race condition: integrations that reconnect early can write new records before older records finish restoring, creating duplicates or overwriting later changes. The team can pause for another day of testing or proceed with monitoring and customer-specific rollback.

  7. Wednesday
    Module 2 · Recovery and scrutiny

    A regulated customer reports missed obligations

    A financial-services customer says the outage prevented required approvals and regulatory reporting. A healthcare customer reports delayed care coordination but no confirmed patient harm. A public-company customer says its close process and disclosure controls depend on the platform. Regulators begin contacting customers about third-party resilience.

PARTICIPANTS

Bring the decision-makers who would own the real event

Assign people to functions, not titles alone. If one person owns several functions, keep the roles distinct during discussion so conflicts and handoffs remain visible.

Incident CommanderSite Reliability and Engineering LeadChief Information Security and Privacy OfficerBusiness Continuity LeadGeneral CounselCommercial Legal LeadCustomer Operations LeadRegulatory and Compliance LeadCommunications LeadFinance and Insurance LeadBoard and Audit Committee LiaisonExecutive Sponsor

REAL-WORLD INCIDENTS

Ground the exercise in events teams can recognize

Use these cases during planning or the prebrief. They are factual anchors, not scripts. Adapt the scenario to the organization’s technology, industry, geography, contracts, regulators, and risk profile.

LEGAL, REGULATORY & STANDARDS LENS

Authorities worth testing against the scenario

Applicability depends on the organization and facts. Use counsel and subject-matter owners to tailor deadlines, thresholds, privileges, preservation, reporting, and communications.

PRACTITIONER PERSPECTIVES

Law-firm and security-professional guidance

External perspectives help the design team challenge internal assumptions. They do not replace organization-specific legal or technical advice.

EXPECTED OUTPUT

Finish with an improvement plan, not a score.

  • A tested out-of-band incident and customer-contact model
  • Safe tenant-scale recovery and acceptance criteria
  • A contract and remedy decision matrix
  • Regulated-customer and disclosure escalation paths
  • Validated architecture, continuity, and governance improvements
Start with this exerciseRead the facilitator guide