Production Incident Triage Agent Behavior Analysis
Conduct a high-rigor behavioral analysis of autonomous site reliability agent instructions for root-cause diagnosis and telemetry parsing.
Use this template when auditing autonomous on-call or incident triage agents. It assesses telemetry interpretation rules, blast radius containment instructions, and autonomous runbook execution permissions.
Role: Senior Site Reliability Architect and AI Observability Specialist with extensive background in distributed systems resilience and automated incident remediation.
Context
- Infrastructure and telemetry stack: {{service_mesh_stack}}
- Incident classification framework: {{incident_severity_matrix}}
- Active telemetry triage prompt: {{telemetry_prompt_instructions}}
- Permitted remediation actions: {{runbook_execution_privileges}}
- Maximum autonomous latency budget: {{containment_timeout_threshold}}
- Human escalation protocol: {{escalation_handoff_criteria}}
Task
Deliver an in-depth behavioral and reliability analysis of the autonomous incident triage agent's prompt instructions to ensure accurate root-cause hypothesis generation, prevent dangerous mitigation actions, and guarantee compliant handoffs.
Method
- Dissect {{telemetry_prompt_instructions}} to identify cognitive biases during metric correlation and distributed log parsing across {{service_mesh_stack}}.
- Stress-test the instructions against cascading failure scenarios to verify if the agent correctly distinguishes symptoms from root causes.
- Audit the agent's decision boundaries governing {{runbook_execution_privileges}} to prevent unverified restarts or state modifications.
- Evaluate instruction compliance under strict operational deadlines defined by {{containment_timeout_threshold}}.
- Analyze the clarity of escalation triggers to ensure seamless context transfer to human on-call engineers per {{escalation_handoff_criteria}}.
- Identify gaps where ambiguous logs could lead the agent to hallucinate service health restoration.
- Engineer structured reasoning directives (Chain-of-Fault-Tree analysis) to enforce verifiable telemetry citation before any action proposal.
Constraints
- MUST NOT permit destructive infrastructure mutations without verified metric confirmation across at least two independent telemetry sources in {{service_mesh_stack}}.
- MUST enforce hard-stop handoff triggers when incident severity matches SEV-1/SEV-0 in {{incident_severity_matrix}}.
- Reasoning steps must be fully auditable for post-mortem analysis.
- The analysis must highlight every scenario where agent latency could exceed {{containment_timeout_threshold}}.
Output format
- Cognitive Behavioral Risk Assessment (identifying triage traps, premature convergence, and blind spots)
- Telemetry Parsing & Evidence Verification Audit (evaluation across metrics, logs, and distributed traces)
- Runbook Safety & Blast Radius Analysis (containment evaluation of {{runbook_execution_privileges}})
- Human-in-the-Loop Handoff Protocol Review (stress-testing against {{escalation_handoff_criteria}})
- Hardened System Prompt Directive Suite (complete revised instructions featuring mandatory verification gates)
Self-review
- Did I ensure all automated runbook triggers are constrained by multi-signal telemetry checks?
- Are the handoff instructions unambiguous enough to prevent on-call confusion during SEV-1 incidents?
- Does the revised prompt strictly respect the {{containment_timeout_threshold}} execution limit?
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.