Long-form
AuraScore 81/100

Production Incident Post-Mortem Narrative Verification Checklist

Rigorous quality assurance checklist for long-form incident root cause analyses and technical post-mortem reports.

Use this template after critical outages or production bugs to review the technical post-mortem report before broad distribution. It ensures root-cause depth, timeline fidelity, telemetry evidence, and systemic remediation tracking.

Template

Role: Staff Site Reliability Engineer & Incident Review Lead specializing in distributed systems post-mortems and resilience engineering.

Context

  • Incident Identifier: {{incident_id}}
  • Core Affected Services: {{affected_services}}
  • Severity Classification: {{severity_tier}}
  • Total Service Degradation Time: {{outage_duration}}
  • Impact Footprint: {{customer_impact_scope}}
  • Action Item Owners: {{remediation_owners}}

Task

Produce a comprehensive, forensic review checklist to validate that the long-form post-mortem report for {{incident_id}} provides a blameless, factually precise, and deeply technical analysis of systemic failure points and sustainable remediations.

Method

  1. Review the scope of {{affected_services}} alongside the incident duration of {{outage_duration}} to define timeline fidelity benchmarks.
  2. Construct verification checks for the chronological timeline, ensuring log timestamps, telemetry spans, and engineer actions correlate without gaps.
  3. Create inspection criteria for the technical causal chain, validating that root mechanisms (e.g., race conditions, thread pool exhaustion, cascading timeouts) are proven rather than assumed.
  4. Design checklist items to verify that {{customer_impact_scope}} is represented with exact percentiles, error rates, and degraded API endpoints.
  5. Draft verification gates for preventative action items assigned to {{remediation_owners}}, ensuring work is scoped with concrete delivery milestones.
  6. Formulate checks for detection and monitoring telemetry, identifying why automated alerts failed or were delayed during the {{severity_tier}} incident.
  7. Detail human factors and operational tooling checklist items assessing runaway runbooks, blast radius containment, and communication cadence.

Constraints

  • Checklist MUST enforce blameless language conventions, eliminating finger-pointing while maintaining accountability.
  • Review criteria MUST require log excerpts, metrics dashboards, or distributed trace links as mandatory attachments.
  • Action items must be validated for architectural prevention rather than purely surface-level patch fixes.
  • Checklist items must be grouped chronologically and by investigative domain.

Output format

  1. Incident Metadata Header (ID, Severity, Scope)
  2. Phase 1: Incident Timeline & Telemetry Completeness Checklist (5 items)
  3. Phase 2: Technical Causal Mechanism & Failure Domain Checklist (6 items)
  4. Phase 3: Impact Analysis & Data Integrity Checklist (4 items)
  5. Phase 4: Detection, Alerting & Observability Gap Checklist (4 items)
  6. Phase 5: Remediation Action Item & SRE Ownership Checklist (5 items with owner validation fields)

Self-review

  • Confirm that {{incident_id}}, {{severity_tier}}, and all other variables are integrated across the checklist stages.
  • Check that each item specifies an explicit verification method (e.g., Trace verification, Code diff inspection).
  • Ensure the checklist prevents superficial 'human error' classifications and demands systemic remediation.
AuraScore breakdown
81/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering12/12 · Strong

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness3/5 · Adequate

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

writing-content
writing-long-form
software-engineering-debugging
sre
post-mortem
incident-management