Your incident response team has brought systems back online. Your communications team has shifted from crisis mode to recovery language. Your board wants to move on.
But if you're measuring breach recovery by operational uptime alone, you're leaving the door open for the same adversary to walk back in.
This checklist helps you distinguish between service restoration and actual security recovery. It covers three distinct phases: operational recovery, adversary eviction, and governance remediation. Each item requires documented evidence before you can mark it complete.
What This Checklist Covers
This checklist applies after a confirmed security incident where an adversary gained unauthorized access to your systems. It's designed for privacy officers and security leaders who need to ensure that "back online" doesn't become "compromised again in three months."
The checklist assumes you've already completed immediate containment and are now evaluating whether it's safe to resume normal operations. It supplements your incident response plan by addressing gaps that most organizations defer or skip entirely.
Prerequisites
Before starting this checklist, confirm you have:
- Documented timeline of the incident from initial access through containment.
- Named incident owner with authority to delay service restoration if security questions remain unresolved.
- Access to forensic findings including indicators of compromise, affected systems, and credential exposure.
- Executive sponsor who understands that recovery completion is an evidence-based decision, not a calendar deadline.
Checklist Items
Operational Recovery
1. Restore systems from known-good sources
Rebuild compromised systems from verified clean backups or gold images taken before the incident window. Do not simply patch or clean infected systems in place.
Good looks like: Documented restoration date for each affected system, source backup timestamp, and verification that the backup predates the earliest evidence of compromise.
2. Re-enable user access with fresh credentials
Force password resets for all users with access to affected systems. Revoke and reissue API keys, service account credentials, and access tokens that existed during the incident window.
Good looks like: Credential reset completion report showing 100% coverage of in-scope accounts, with no exceptions for convenience or business pressure.
3. Validate application functionality in a monitored environment
Test restored applications under enhanced logging before returning them to production. Confirm that business processes work as expected and that monitoring coverage captures relevant activity.
Good looks like: Test plan with pass/fail criteria, completed test results, and confirmation that logging captures authentication, privilege escalation, and data access events.
Adversary Eviction
4. Document how the attacker gained initial access
Identify the specific vulnerability, credential, or configuration that allowed the adversary's first foothold. If you cannot determine this with confidence, document what you do know and what remains uncertain.
Good looks like: Forensic report section titled "Initial Access Vector" with specific technical detail (e.g., "unpatched VPN appliance CVE-2024-XXXX" or "compromised contractor credential"), not vague language like "phishing" or "social engineering."
5. Search for and remove persistence mechanisms
Hunt for scheduled tasks, registry modifications, dormant accounts, compromised service principals, tampered monitoring agents, and cloud access tokens that could allow the adversary to return.
Good looks like: Documented search methodology, tools used, systems scanned, findings list (even if zero findings), and remediation evidence for each discovery.
6. Verify identity system integrity
Confirm that privileged groups, federation trusts, and identity provider configurations have not been modified. Review recent changes to service principals, application registrations, and conditional access policies.
Good looks like: Identity audit report covering privileged group membership, federated trust relationships, and application permissions, with comparison to pre-incident baseline or documented expected state.
7. Confirm monitoring coverage and alert function
Validate that security monitoring tools were not disabled or tampered with during the incident. Test that alerts fire as expected for behaviors consistent with the attack pattern.
Good looks like: Monitoring validation test showing that detection rules trigger correctly for the techniques the adversary used, with alert delivery confirmed to the security team.
Governance Remediation
8. Identify the control or governance failure that enabled the incident
Determine which control gap, unmanaged exception, delayed patch, weak segmentation, or accepted risk made the attack possible or allowed it to progress undetected.
Good looks like: Root cause statement in the incident report that names a specific control failure (e.g., "VPN appliances excluded from patch management scope due to vendor support concern") with accountable owner identified.
9. Assign remediation ownership with executive oversight
Designate a named individual responsible for closing the identified control gap, with a deadline and reporting line to an executive who can remove obstacles.
Good looks like: Remediation plan entry showing owner name, target completion date, executive sponsor, and monthly status reporting requirement until closure.
10. Update risk register to reflect residual uncertainty
If you cannot achieve full confidence in adversary removal or control remediation, document what remains uncertain and what compensating controls are in place.
Good looks like: Risk register entry describing the specific uncertainty (e.g., "OT segment rebuilt but cannot verify absence of firmware-level persistence"), compensating controls deployed, and acceptance signed by accountable executive.
Common Mistakes
Declaring recovery when systems are online but questions remain unanswered. If you cannot confirm adversary removal, you haven't recovered. You've resumed operations in a potentially compromised state.
Treating credential resets as optional or phased. Attackers often establish access through multiple credentials. Partial resets leave paths open.
Skipping governance remediation because the technical fix is complete. If the control gap that enabled the breach remains unaddressed, you're preserving the conditions for the next incident.
Confusing forensic uncertainty with completion. You won't always have perfect visibility. Document what you don't know and manage it as residual risk, but don't pretend uncertainty doesn't exist.
Allowing business pressure to override evidence requirements. The board question "Are we back up?" should be followed by "What evidence do we have that it's safe to be back up?" If you can't answer the second question, the first one is premature.
Next Steps
Once you've completed this checklist with documented evidence for each item, schedule a recovery closure review with your incident owner, executive sponsor, and legal counsel. Present the evidence, acknowledge any residual uncertainty, and obtain written approval to declare the incident closed.
If items remain incomplete, document why and establish a timeline for completion. Do not declare recovery complete simply because operational systems are available.
For GDPR-regulated organizations, remember that the 72-hour notification to your supervisory authority under Article 33 is separate from recovery completion. Your breach register should reflect the actual closure date based on this checklist, not the date systems came back online.
Recovery isn't finished when the crisis ends. It's finished when you can demonstrate that the adversary is gone and the weakness that let them in has been addressed.



