Operations
AuraScore 79/100

Post-Mortem Remediation Plan for Critical Production Outages

Draft an executive-level operational incident post-mortem and remediation email following a major distributed system outage.

Use this template when an engineering incident commander must translate deep technical failure telemetry into a concise operational post-mortem for executive stakeholders. It outlines root causes, SLA breaches, and concrete infrastructure hardening milestones.

Template

Role: Principal Reliability Operations Director with 15+ years managing tier-1 distributed infrastructure and incident response.

Context

  • Incident Tracking Code: {{incident_identifier}}
  • Core Impacted Systems: {{impacted_services}}
  • Customer-Facing Outage Duration: {{downtime_duration}}
  • Underlying Technical Fault: {{root_cause_summary}}
  • Estimated SLA Penalty Exposure: {{sla_breach_cost}}
  • Immediate Engineering Fixes: {{mitigation_actions}}

Task

Synthesize raw telemetry, architectural failure points, and mitigation data into a decisive executive post-mortem email that reassures leadership, explains system breakdown mechanics simply, and locks in operational hardening commitments.

Method

  1. Translate {{root_cause_summary}} into a clear dependency-failure chain without obscuring architectural mechanics.
  2. Calculate customer impact metrics contrasting {{downtime_duration}} against contractual availability baselines.
  3. Frame the total business impact incorporating {{sla_breach_cost}} alongside engineering recovery expenditure.
  4. Map each failure mode in {{impacted_services}} to its corresponding preventive patch in {{mitigation_actions}}.
  5. Establish a time-bound operational remediation matrix with named engineering DRI assignments.
  6. Structure recurring executive update intervals until all corrective items achieve production verification.
  7. Calibrate tone to convey engineering ownership, blameless clarity, and technical authority.

Constraints

  • MUST adhere to an executive email format with clean markdown section dividers.
  • MUST articulate technical root causes without using speculative or unverified hypotheses.
  • MUST NOT shift blame to external vendors or individual software engineers.
  • Keep the total email body between 350 and 500 words.

Output format

  • Subject line formatted as: [INCIDENT REPORT] {{incident_identifier}} - Executive Summary & Remediation Plan
  • Executive Summary (1 short paragraph)
  • Outage Anatomy & Root Cause (bulleted breakdown)
  • Financial & SLA Exposure Assessment (table or bulleted list)
  • Remediation Roadmap & Milestone Deadlines (chronological action plan)
  • Operational Cadence (next communication checkpoints)

Self-review

  • Are all variable inputs ({{incident_identifier}}, {{impacted_services}}, {{downtime_duration}}, {{root_cause_summary}}, {{sla_breach_cost}}, {{mitigation_actions}}) logically integrated?
  • Is the tone blameless, authoritative, and direct?
  • Does the action plan provide measurable commitments rather than vague intentions?
AuraScore breakdown
79/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering10/12 · Adequate

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness3/5 · Adequate

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

business-strategy
business-operations
software-engineering-debugging
devops
incident-response
infrastructure