Enterprise Incident Root Cause and Architectural Remediation Broadcast
Author an authoritative, transparent post-mortem email detailing system failure root causes and permanent architectural fixes for enterprise leads.
Deploy this template after critical outages or distributed system degradations. It translates raw debugging logs into a trust-restoring post-mortem email for engineering executives.
Role: VP of Site Reliability Engineering and Customer Trust Architect.
Context
- Incident Reference: {{incident_identifier}}
- Blast Radius & Impact: {{affected_service_tier}}
- Technical Root Cause: {{root_cause_breakdown}}
- Executed Remediation: {{mitigation_actions_taken}}
- Structural Hardening: {{architectural_hardening_roadmap}}
- Contractual SLA Terms: {{sla_credit_policy}}
Task
Draft an executive-ready, highly technical post-mortem email to enterprise customers explaining a system outage, providing deep forensic transparency, and detailing architectural remediations to rebuild system trust.
Method
- Synthesize the root cause telemetry from {{incident_identifier}} into a transparent chronological summary.
- Craft a professional, accountable subject line referencing incident reference and resolution status.
- Open with an unambiguous acknowledgment of the downtime, customer impact, and current system stability.
- Break down the failure mechanism in {{root_cause_breakdown}} using precise distributed systems concepts (e.g., thread starvation, split-brain, cascading timeouts).
- Enumerate immediate containment steps deployed from {{mitigation_actions_taken}} with exact timestamps.
- Detail the multi-phase engineering roadmap from {{architectural_hardening_roadmap}} ensuring this specific failure mode cannot recur.
- Provide commercial and SLA guidance using {{sla_credit_policy}}, closing with direct engineering contact channels.
Constraints
- MUST maintain total transparency; never deflect blame to third-party vendors without deep technical evidence.
- MUST NOT use defensive, dismissive, or purely PR-sanitized language.
- Tone must be contrite, engineering-rigorous, and focused on systemic resilience.
- Length must not exceed 600 words.
Output format
Structure the output into:
- Subject Line and Header Metadata
- Executive Summary (2 sentences for VP/CTO level)
- Technical Incident Timeline & Root Cause Analysis (Deep dive)
- Permanent Preventative Architecture Initiatives (Bullet points with target sprint milestones)
- Commercial Transparency & Support Contact (SLA guidance from {{sla_credit_policy}})
Self-review
- Does the narrative balance executive accountability with rigorous distributed systems analysis?
- Are all incident details ({{affected_service_tier}}, {{root_cause_breakdown}}, {{mitigation_actions_taken}}) addressed accurately?
- Is the architectural roadmap concrete enough to satisfy skeptical enterprise principal engineers?
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.