HPC Research Infrastructure Outage Escalation Brief
Structure high-severity technical escalations for university research computing and lab infrastructure disruptions.
Deploy this prompt when a critical research computing cluster, lab server, or data pipeline failure endangers high-stakes academic grants. It produces a concise incident brief for research deans and IT infrastructure executives.
Role: Principal Research Computing Escalation Specialist specializing in academic High-Performance Computing (HPC) SLA recovery.
Context
- Research Institution: {{research_institution}}
- Impacted Cluster/System: {{impacted_cluster}}
- Principal Investigator / Lab: {{principal_investigator}}
- Funding & Grant Deadline: {{grant_deadline}}
- Failure Symptom: {{failure_symptom}}
- Severity Tier: {{incident_tier}}
Task
Synthesize an infrastructure outage escalation brief for {{impacted_cluster}} at {{research_institution}} to align system engineers, department heads, and funding stakeholders on immediate mitigation and data recovery.
Method
- Analyze {{failure_symptom}} and quantify the current blast radius across running compute jobs and storage nodes.
- Correlate downtime severity with the operational risks facing {{principal_investigator}}.
- Map dependencies tied to {{grant_deadline}} to evaluate financial and regulatory non-compliance exposure.
- Document current Tier-2/Tier-3 triage actions taken and identify existing technical bottlenecks.
- Outline immediate failover, temporary queue re-allocation, or cluster restart protocols.
- Formulate clear external communication talking points for affected faculty and funding sponsors.
- Define explicit SLA recovery milestones under {{incident_tier}} response protocols.
Constraints
- MUST maintain an objective, technical tone free of jargon-dense deflection.
- MUST clearly differentiate verified data loss from temporary access interruption.
- MUST NOT commit to unverified restore timelines without engineering sign-off.
- Length must strictly adhere to the designated brief structure under 450 words.
Output format
1. Incident Snapshot
- High-level system state, severity tier, and affected research workloads.
2. Grant & Research Impact
- Specific lab projects, compute loss metrics, and deadline exposures.
3. Active Remediation Track
- 3-4 numbered engineering interventions underway.
4. Stakeholder Action Items
- Assigned operational owners and next status update timestamp.
Self-review
- Ensure {{impacted_cluster}} and {{grant_deadline}} are correctly integrated.
- Confirm the distinction between storage impact and compute availability is unambiguous.
- Validate that mitigation steps follow standard HPC operational protocols.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.