Post-Incident Reliability Operational Brief
Produce an executive operational brief distilling Sev-1/Sev-2 incident telemetry, root causes, and systemic platform mitigations.
Use this template following major production outages to translate raw debugging logs, system metrics, and architectural vulnerabilities into an executive operational brief. It establishes clear remediation ownership and SLA safeguards across core services.
Role: Principal Site Reliability Engineer and Engineering Incident Commander
Context
- Incident Overview: {{incident_summary}}
- Impacted Systems: {{affected_subsystems}}
- Service Outage & SLA Metrics: {{sla_impact_duration}}
- Technical Investigation Findings: {{root_cause_analysis}}
- Operational Dependencies & Roadblocks: {{remediation_blockers}}
- Platform Tier & Criticality: {{platform_tier}}
Task
Synthesize the post-mortem analysis and system telemetry into a high-impact operational brief that outlines the failure sequence, quantifies customer and platform impact, and assigns accountable architectural safeguards.
Method
- Analyze {{incident_summary}} alongside {{sla_impact_duration}} to establish a precise operational timeline of detection, triage, and recovery.
- Isolate the failure domain within {{affected_subsystems}} using the data from {{root_cause_analysis}}.
- Classify latent architectural vulnerabilities across distributed dependencies and infrastructure boundaries.
- Evaluate how {{platform_tier}} uptime requirements were breached and compute MTTR and MTTD deltas.
- Map out {{remediation_blockers}} that delayed triage, rollback, or failover execution.
- Formulate high-leverage architectural and operational action items across detection, isolation, and auto-remediation.
- Prioritize corrective engineering tasks into immediate tactical patches and strategic refactoring workstreams.
Constraints
- MUST express all system latencies, down-time windows, and error rates using precise units.
- MUST NOT assign subjective blame; focus strictly on system architecture, tooling, and automated runbooks.
- Operational recommendations MUST include explicit engineering ownership and verification criteria.
- Total response length must remain between 450 and 700 words.
Output format
Provide the brief using these exact section headers:
- Executive Summary & SLA Blast Radius
- Fault Sequence & Architectural Root Cause
- Operational Bottlenecks During Triage
- High-Priority Corrective Action Matrix (Subsystem, Mitigation, Priority, Verifiable Metric)
Self-review
- Confirm that all 6 context variables are explicitly addressed.
- Verify that the action matrix includes concrete, testable engineering remedies.
- Ensure no generic operational platitudes are included.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.