Post-Mortem Remediation Plan for Critical Production Outages
Draft an executive-level operational incident post-mortem and remediation email following a major distributed system outage.
Use this template when an engineering incident commander must translate deep technical failure telemetry into a concise operational post-mortem for executive stakeholders. It outlines root causes, SLA breaches, and concrete infrastructure hardening milestones.
Role: Principal Reliability Operations Director with 15+ years managing tier-1 distributed infrastructure and incident response.
Context
- Incident Tracking Code: {{incident_identifier}}
- Core Impacted Systems: {{impacted_services}}
- Customer-Facing Outage Duration: {{downtime_duration}}
- Underlying Technical Fault: {{root_cause_summary}}
- Estimated SLA Penalty Exposure: {{sla_breach_cost}}
- Immediate Engineering Fixes: {{mitigation_actions}}
Task
Synthesize raw telemetry, architectural failure points, and mitigation data into a decisive executive post-mortem email that reassures leadership, explains system breakdown mechanics simply, and locks in operational hardening commitments.
Method
- Translate {{root_cause_summary}} into a clear dependency-failure chain without obscuring architectural mechanics.
- Calculate customer impact metrics contrasting {{downtime_duration}} against contractual availability baselines.
- Frame the total business impact incorporating {{sla_breach_cost}} alongside engineering recovery expenditure.
- Map each failure mode in {{impacted_services}} to its corresponding preventive patch in {{mitigation_actions}}.
- Establish a time-bound operational remediation matrix with named engineering DRI assignments.
- Structure recurring executive update intervals until all corrective items achieve production verification.
- Calibrate tone to convey engineering ownership, blameless clarity, and technical authority.
Constraints
- MUST adhere to an executive email format with clean markdown section dividers.
- MUST articulate technical root causes without using speculative or unverified hypotheses.
- MUST NOT shift blame to external vendors or individual software engineers.
- Keep the total email body between 350 and 500 words.
Output format
- Subject line formatted as: [INCIDENT REPORT] {{incident_identifier}} - Executive Summary & Remediation Plan
- Executive Summary (1 short paragraph)
- Outage Anatomy & Root Cause (bulleted breakdown)
- Financial & SLA Exposure Assessment (table or bulleted list)
- Remediation Roadmap & Milestone Deadlines (chronological action plan)
- Operational Cadence (next communication checkpoints)
Self-review
- Are all variable inputs ({{incident_identifier}}, {{impacted_services}}, {{downtime_duration}}, {{root_cause_summary}}, {{sla_breach_cost}}, {{mitigation_actions}}) logically integrated?
- Is the tone blameless, authoritative, and direct?
- Does the action plan provide measurable commitments rather than vague intentions?
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.