Infrastructure Telemetry and Chaos Engineering Reliability Dispatch
Creates an SRE newsletter report detailing cluster telemetry, chaos experiment outcomes, SLO trends, and platform hardening steps.
Use this template when publishing a periodic reliability and chaos engineering newsletter report. It transforms raw observability data, SLO burn rates, and resilience experiment outcomes into prioritized operational intelligence.
Role: Principal Site Reliability Engineer and Systems Performance Lead overseeing multi-region platform health, resilience patterns, and chaos engineering practices.
Context
- Infrastructure Footprint: {{infrastructure_footprint}}
- Observability Stack: {{observability_stack}}
- Service SLO Breaches & Burn: {{service_slo_breaches}}
- Chaos Engineering Experiments: {{chaos_experiments}}
- Cost & Latency Trade-offs: {{cost_latency_tradeoffs}}
- Critical Remediation Actions: {{critical_remediation_actions}}
Task
Compile a comprehensive reliability newsletter report analyzing infrastructure telemetry, resilience testing outcomes, and SLO compliance across {{infrastructure_footprint}} to align platform operators and application engineers on systemic reliability priorities.
Method
- Parse the metrics and telemetry gathered via {{observability_stack}} to evaluate multi-region platform stability.
- Quantify error budget consumption and threshold violations detailed in {{service_slo_breaches}}.
- Evaluate the hypothesis, fault injection method, and recovery blast radius for each item in {{chaos_experiments}}.
- Analyze compute, network, and storage cost implications against throughput requirements listed in {{cost_latency_tradeoffs}}.
- Categorize and prioritize the platform hardening tasks outlined in {{critical_remediation_actions}} by engineering effort vs. risk reduction.
- Formulate proactive resilience directives for application teams (e.g., circuit breaker tuning, deadline propagation, retry jitter).
- Produce a consolidated reliability score for critical platform components based on telemetry evidence.
Constraints
- MUST use precise quantitative metrics (percentages, percentiles like p95/p99/p99.9, error budgets) throughout.
- MUST NOT present chaos experiments without explicitly noting whether the system self-healed within targeted RTO/RPO limits.
- Recommendations MUST directly reference the monitoring and telemetry capabilities of {{observability_stack}}.
- Action items must clearly assign responsibility between Platform/Infra teams and Domain Product squads.
Output format
Generate a detailed reliability report formatted as follows:
- Global Reliability & SLO Health Dashboard (Summary table of services, SLO targets, actuals, error budget remaining)
- Chaos & Resilience Testing Digest (Case study breakdown of {{chaos_experiments}}: Hypothesis, Fault Injected, Observed Behavior, Gap Identified)
- Telemetry & Latency Profiling (In-depth analysis of anomalies surfaced by {{observability_stack}})
- Efficiency vs. Resilience Trade-off Ledger (Analysis of {{cost_latency_tradeoffs}})
- Mandated Platform Actions & Remediation Matrix (Prioritized list based on {{critical_remediation_actions}})
Self-review
- Confirm all 6 input variables are accurately reflected in their respective sections.
- Ensure chaos testing observations clearly distinguish between graceful degradation and hard service failure.
- Verify that every recommended remediation includes a clear justification based on the cited SLO breaches.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.