Distributed Systems Major Outage Root Cause Email Brief
Architect an enterprise post-incident transparency email brief for system outages.
Use this template when formulating post-mortem email communications following a critical production incident. It equips technical leaders to deliver blame-free, technically thorough incident debriefs to enterprise clients.
Role: Principal SRE Communications Strategist and Cloud Infrastructure Product Marketer.
Context
- Incident identifier: {{service_incident_name}}
- System downtime footprint: {{outage_duration_and_impact}}
- Root architectural defect: {{root_cause_architecture_summary}}
- Fix deployment and hash: {{remediation_commit_details}}
- Affected stakeholder tier: {{enterprise_customer_segment}}
- Contractual SLA remedy: {{sla_credit_policy}}
Task
Create an enterprise post-incident transparency email brief that translates the distributed systems failure in {{service_incident_name}} into an authoritative, engineering-grade post-mortem broadcast for {{enterprise_customer_segment}} leadership.
Method
- Analyze the system architecture breakdown and synthesize {{root_cause_architecture_summary}} into clear causal chains.
- Structure the timeline detailing initial latency spike, failover partition isolation, and full recovery according to {{outage_duration_and_impact}}.
- Detail the exact patch, configuration rollout, or kernel fix referenced in {{remediation_commit_details}}.
- Formulate executive reassurance messaging that outlines long-term architectural redundancy improvements without obfuscating fault.
- Draft precise protocol instructions for claiming commercial remedies under {{sla_credit_policy}}.
- Define audience segmentation rules separating VP/CTO high-level summaries from Staff SRE deep-dive appendices.
- Establish customer success enablement guidelines for handling incoming support tickets from affected engineering teams.
Constraints
- MUST avoid passive corporate apologies; maintain rigorous engineering transparency with precise fault isolation terminology.
- MUST NOT omit technical metrics such as P99 latency, MTTR, packet drop rates, or circuit breaker states.
- Tone must balance high accountability with enterprise system reliability rigor.
- All sections must directly reference the provided context variables.
Output format
- Incident Impact & Blast Radius Summary (Max 200 words)
- Technical Deep-Dive Email Narrative (Subject line options, timeline architecture, code/config fix breakdown)
- SLA & Preventive Action Matrix (Remediation roadmap and commercial resolution protocol)
- Field Enablement & FAQ Directive (Guidance for technical account managers)
Self-review
- Ensure all technical terms align with distributed computing realities.
- Check that the SLA recovery steps in {{sla_credit_policy}} are actionable and clear.
- Confirm exact character length constraints are satisfied.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.