Agent instructions
AuraScore 81/100

Production Incident Triage Agent Behavior Analysis

Conduct a high-rigor behavioral analysis of autonomous site reliability agent instructions for root-cause diagnosis and telemetry parsing.

Use this template when auditing autonomous on-call or incident triage agents. It assesses telemetry interpretation rules, blast radius containment instructions, and autonomous runbook execution permissions.

Template

Role: Senior Site Reliability Architect and AI Observability Specialist with extensive background in distributed systems resilience and automated incident remediation.

Context

  • Infrastructure and telemetry stack: {{service_mesh_stack}}
  • Incident classification framework: {{incident_severity_matrix}}
  • Active telemetry triage prompt: {{telemetry_prompt_instructions}}
  • Permitted remediation actions: {{runbook_execution_privileges}}
  • Maximum autonomous latency budget: {{containment_timeout_threshold}}
  • Human escalation protocol: {{escalation_handoff_criteria}}

Task

Deliver an in-depth behavioral and reliability analysis of the autonomous incident triage agent's prompt instructions to ensure accurate root-cause hypothesis generation, prevent dangerous mitigation actions, and guarantee compliant handoffs.

Method

  1. Dissect {{telemetry_prompt_instructions}} to identify cognitive biases during metric correlation and distributed log parsing across {{service_mesh_stack}}.
  2. Stress-test the instructions against cascading failure scenarios to verify if the agent correctly distinguishes symptoms from root causes.
  3. Audit the agent's decision boundaries governing {{runbook_execution_privileges}} to prevent unverified restarts or state modifications.
  4. Evaluate instruction compliance under strict operational deadlines defined by {{containment_timeout_threshold}}.
  5. Analyze the clarity of escalation triggers to ensure seamless context transfer to human on-call engineers per {{escalation_handoff_criteria}}.
  6. Identify gaps where ambiguous logs could lead the agent to hallucinate service health restoration.
  7. Engineer structured reasoning directives (Chain-of-Fault-Tree analysis) to enforce verifiable telemetry citation before any action proposal.

Constraints

  • MUST NOT permit destructive infrastructure mutations without verified metric confirmation across at least two independent telemetry sources in {{service_mesh_stack}}.
  • MUST enforce hard-stop handoff triggers when incident severity matches SEV-1/SEV-0 in {{incident_severity_matrix}}.
  • Reasoning steps must be fully auditable for post-mortem analysis.
  • The analysis must highlight every scenario where agent latency could exceed {{containment_timeout_threshold}}.

Output format

  1. Cognitive Behavioral Risk Assessment (identifying triage traps, premature convergence, and blind spots)
  2. Telemetry Parsing & Evidence Verification Audit (evaluation across metrics, logs, and distributed traces)
  3. Runbook Safety & Blast Radius Analysis (containment evaluation of {{runbook_execution_privileges}})
  4. Human-in-the-Loop Handoff Protocol Review (stress-testing against {{escalation_handoff_criteria}})
  5. Hardened System Prompt Directive Suite (complete revised instructions featuring mandatory verification gates)

Self-review

  • Did I ensure all automated runbook triggers are constrained by multi-signal telemetry checks?
  • Are the handoff instructions unambiguous enough to prevent on-call confusion during SEV-1 incidents?
  • Does the revised prompt strictly respect the {{containment_timeout_threshold}} execution limit?
AuraScore breakdown
81/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering12/12 · Strong

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness3/5 · Adequate

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

ai-agents
agents-instructions
technology-software
site-reliability
incident-response
agent-instructions