Evaluation
AuraScore 83/100

Humanitarian Crisis Multi-Agent Triage Performance Brief

Evaluate autonomous multi-agent disaster response networks against humanitarian principles, latency constraints, and critical resource allocation metrics.

Deploy this template when nonprofit operational leads must assess autonomous multi-agent swarms operating in low-connectivity crisis environments. It establishes a multi-dimensional evaluation of triage accuracy, bandwidth resilience, and ethical aid distribution.

Template

Role: Lead Humanitarian AI Evaluation Specialist with expertise in field triage systems and NGO mission integrity.

Context

  • Humanitarian Initiative: {{ngo_initiative}}
  • Operational Environment: {{disaster_context}}
  • Multi-Agent Swarm Topology: {{agent_swarm_topology}}
  • Field Telemetry Logs: {{telemetry_logs}}
  • Humanitarian Mandate & Core Principles: {{ethical_triage_mandate}}
  • In-Country Field Feedback: {{partner_field_feedback}}

Task

Produce an authoritative triage performance evaluation brief analyzing multi-agent decision telemetry from {{ngo_initiative}} to ensure resource prioritization strictly complies with {{ethical_triage_mandate}} amidst degraded field conditions.

Method

  1. Deconstruct {{agent_swarm_topology}} to map inter-agent handoffs between information gathering, severity scoring, and logistical dispatch agents.
  2. Ingest {{telemetry_logs}} to calculate mean decision latency, consensus divergence, and packet-loss failure rates during high-stress simulations.
  3. Benchmark automated supply and medical triage recommendations against baseline standards defined in {{ethical_triage_mandate}}.
  4. Cross-examine agent supply allocations against qualitative ground truth reports provided in {{partner_field_feedback}} to identify geographic or demographic neglect.
  5. Audit fail-safe behavior when autonomous communication links degrade in the specified {{disaster_context}}.
  6. Compute the hallucination frequency of critical field metadata (e.g., road accessibility, casualty figures, shelter capacities).
  7. Formulate operational protocols for humanitarian leads to override agent swarm allocations during localized field anomalies.

Constraints

  • The evaluation MUST benchmark all agent behaviors against the Do No Harm framework in {{ethical_triage_mandate}}.
  • Critical triage errors MUST be classified using standard disaster triage triage-tag categories (Immediate, Delayed, Minimal, Expectant).
  • The brief MUST NOT approve any autonomous dispatch logic lacking an offline graceful degradation state.
  • Analysis must distinguish between network-induced latency and agent inference bottlenecks.

Output format

  • Operational Readiness Summary (Under 200 words with explicit Go/No-Go field deployment rating)
  • Inter-Agent Swarm Telemetry Breakdown (Metrics: Latency, Divergence Rate, Handoff Integrity, Bandwidth Overhead)
  • Ethical Triage & Bias Diagnostic (Detailed analysis contrasting agent outputs with {{ethical_triage_mandate}})
  • Field Discrepancy Analysis (Comparative table matching agent logs against {{partner_field_feedback}})
  • Fallback & Human Override Directives (Numbered list of mandatory operational constraints)

Self-review

  • Did I isolate how {{agent_swarm_topology}} responds to severe network degradation?
  • Are field discrepancies substantiated by qualitative reports from {{partner_field_feedback}}?
  • Does the Go/No-Go verdict explicitly reference compliance with {{ethical_triage_mandate}}?
AuraScore breakdown
83/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering12/12 · Strong

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness5/5 · Strong

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

ai-agents
agents-evaluation
public-sector-nonprofit
humanitarian
disaster response
multi-agent