Production Incident Post-Mortem Narrative Verification Checklist
Rigorous quality assurance checklist for long-form incident root cause analyses and technical post-mortem reports.
Use this template after critical outages or production bugs to review the technical post-mortem report before broad distribution. It ensures root-cause depth, timeline fidelity, telemetry evidence, and systemic remediation tracking.
Role: Staff Site Reliability Engineer & Incident Review Lead specializing in distributed systems post-mortems and resilience engineering.
Context
- Incident Identifier: {{incident_id}}
- Core Affected Services: {{affected_services}}
- Severity Classification: {{severity_tier}}
- Total Service Degradation Time: {{outage_duration}}
- Impact Footprint: {{customer_impact_scope}}
- Action Item Owners: {{remediation_owners}}
Task
Produce a comprehensive, forensic review checklist to validate that the long-form post-mortem report for {{incident_id}} provides a blameless, factually precise, and deeply technical analysis of systemic failure points and sustainable remediations.
Method
- Review the scope of {{affected_services}} alongside the incident duration of {{outage_duration}} to define timeline fidelity benchmarks.
- Construct verification checks for the chronological timeline, ensuring log timestamps, telemetry spans, and engineer actions correlate without gaps.
- Create inspection criteria for the technical causal chain, validating that root mechanisms (e.g., race conditions, thread pool exhaustion, cascading timeouts) are proven rather than assumed.
- Design checklist items to verify that {{customer_impact_scope}} is represented with exact percentiles, error rates, and degraded API endpoints.
- Draft verification gates for preventative action items assigned to {{remediation_owners}}, ensuring work is scoped with concrete delivery milestones.
- Formulate checks for detection and monitoring telemetry, identifying why automated alerts failed or were delayed during the {{severity_tier}} incident.
- Detail human factors and operational tooling checklist items assessing runaway runbooks, blast radius containment, and communication cadence.
Constraints
- Checklist MUST enforce blameless language conventions, eliminating finger-pointing while maintaining accountability.
- Review criteria MUST require log excerpts, metrics dashboards, or distributed trace links as mandatory attachments.
- Action items must be validated for architectural prevention rather than purely surface-level patch fixes.
- Checklist items must be grouped chronologically and by investigative domain.
Output format
- Incident Metadata Header (ID, Severity, Scope)
- Phase 1: Incident Timeline & Telemetry Completeness Checklist (5 items)
- Phase 2: Technical Causal Mechanism & Failure Domain Checklist (6 items)
- Phase 3: Impact Analysis & Data Integrity Checklist (4 items)
- Phase 4: Detection, Alerting & Observability Gap Checklist (4 items)
- Phase 5: Remediation Action Item & SRE Ownership Checklist (5 items with owner validation fields)
Self-review
- Confirm that {{incident_id}}, {{severity_tier}}, and all other variables are integrated across the checklist stages.
- Check that each item specifies an explicit verification method (e.g., Trace verification, Code diff inspection).
- Ensure the checklist prevents superficial 'human error' classifications and demands systemic remediation.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.