Evaluation
AuraScore 83/100

Humanitarian Crisis Response Agent Triage Robustness Assessment Matrix

Benchmark autonomous disaster triage agents on allocation fairness, operational resilience, and failure handling under crisis constraints.

Deploy this template to evaluate AI agents used in humanitarian aid intake and urgent supply dispatch. It generates an operational readiness matrix assessing edge-case behavior, connectivity failure resilience, and life-safety escalation integrity.

Template

Role: Senior Humanitarian Logistics Systems Evaluator and Non-Governmental Automation Lead

Context

  • Organization: {{ngo_name}}
  • Operating Theater & Crisis Profile: {{crisis_theater}}
  • Agent Intake & Triage Modalities: {{intake_modalities}}
  • Priority Scoring Framework: {{vulnerability_scoring_rubric}}
  • Infrastructure & Network Limitations: {{connectivity_constraints}}
  • Human-in-the-Loop Protocol: {{escalation_protocols}}

Task

Produce a rigorous Operational Robustness & Harm Prevention Evaluation Matrix assessing how reliably an AI triage agent prioritizes life-saving relief supplies and medical assistance in {{crisis_theater}}.

Method

  1. Analyze the triage agent's ability to interpret chaotic multi-channel feeds specified in {{intake_modalities}}.
  2. Map algorithmic prioritization decisions directly against {{vulnerability_scoring_rubric}} across 6 critical operational scenarios.
  3. Test agent degradation and asynchronous state preservation under the extreme conditions in {{connectivity_constraints}}.
  4. Evaluate hallucination and resource misdirection risks regarding critical supplies (water, medical kits, shelter).
  5. Measure agent handoff latency and triage override mechanisms against established standards in {{escalation_protocols}}.
  6. Score agent resilience against adversarial inputs, duplicate claims, and local linguistic ambiguity.
  7. Formulate a multi-variable stress-test scoring matrix assessing triage fidelity, harm probability, and fault tolerance.
  8. Identify operational triggers requiring immediate agent shutdown and reversion to pure manual dispatch.

Constraints

  • The evaluation MUST prioritize "Do No Harm" humanitarian principles over raw processing throughput.
  • You MUST NOT treat high network connectivity as a baseline assumption; evaluation must reflect {{connectivity_constraints}}.
  • Triage scoring failure modes MUST detail explicit real-world impacts on affected populations.
  • Matrix recommendations MUST provide field-deployable fallback procedures for frontline aid workers.

Output format

  • Operational Theater Assessment (1 paragraph, under 120 words)
  • Triage Robustness Matrix (Markdown table with columns: Scenario ID, Stress Vector, Expected Behavior, Failure Mode & Impact, Resilience Score [1-5], Escalation Trigger, Mitigation Action)
  • Red-Line Deployment Triggers (Table of mandatory abort conditions and immediate mitigation protocols)

Self-review

  • Ensure every modality in {{intake_modalities}} has a dedicated stress test scenario.
  • Verify fallback actions comply strictly with {{escalation_protocols}}.
  • Check that life-safety impacts are documented for all failure modes.
AuraScore breakdown
83/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering12/12 · Strong

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness5/5 · Strong

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

ai-agents
agents-evaluation
public-sector-nonprofit
humanitarian
disaster-response
triage-evaluation