Grid SCADA Database Incident Post-Mortem Briefing
Compose a developer-focused post-mortem email explaining a high-severity SCADA cluster replication lag or failover incident.
Use this template following an unscheduled database failover or storage bottleneck in energy grid SCADA systems. It produces an executive-ready post-mortem email for engineering leads and control room managers.
Role: Lead Industrial Data Architect specializing in mission-critical SCADA database clusters.
Context
- Grid operating entity: {{grid_operator_name}}
- Database cluster ID: {{cluster_identifier}}
- Root cause trigger: {{incident_trigger_cause}}
- Telemetry ingestion downtime: {{telemetry_loss_duration}}
- Maximum observed replication lag: {{replication_lag_seconds}}
- Primary corrective remediation: {{remediation_action}}
Task
Generate an incident post-mortem email addressed to grid software engineering leadership and operational stakeholders, detailing the root cause, recovery timeline, replication dynamics, and permanent corrective actions for the SCADA database event.
Method
- Formulate an informative subject line incorporating cluster severity and incident reference code.
- State the timeline of the disruption impacting {{cluster_identifier}} at {{grid_operator_name}}.
- Describe the mechanism by which {{incident_trigger_cause}} degraded write throughput and caused {{replication_lag_seconds}} seconds of replica drift.
- Clarify the telemetry exposure window of {{telemetry_loss_duration}} and how backpressure buffering handled incoming substation signals.
- Detail the immediate containment steps taken by database on-call engineers to stabilize primary-replica quorum.
- Elaborate on {{remediation_action}} as the permanent preventative measure.
- Provide an itemized action matrix assigning technical tasks, Jira epics, and target resolution dates.
Constraints
- MUST maintain an objective, blame-free blameless post-mortem tone.
- MUST NOT expose raw credential strings, internal IP ranges, or restricted grid operational topology.
- Keep total email length under 450 words while preserving technical rigor.
- MUST explicitly differentiate between database downtime and end-user grid control impairment.
Output format
An email formatted with:
- Email Header: Subject, Intended Audience, Incident Reference Number, Severity Classification
- Incident Overview (3 bullet points covering duration, cluster, and impact)
- Chronological Recovery Timeline (3-5 milestone timestamps)
- Root Cause & Data Integrity Analysis (concise technical paragraph)
- Corrective Action Items (markdown table with Action, Owner, Due Date)
Self-review
- Confirm that the replication lag metric of {{replication_lag_seconds}} seconds is placed in technical context.
- Ensure the relationship between {{incident_trigger_cause}} and {{telemetry_loss_duration}} is clearly articulated.
- Validate that all action items are assigned to operational roles rather than specific named individuals.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.