Evaluation
AuraScore 79/100

Emergency Aid Allocation Agent Red-Teaming Debrief

Communicate adversarial stress-testing and vulnerability evaluation findings for humanitarian grant allocation agents.

Use this template when evaluating automated aid distribution bots against adversarial attacks, fraud vulnerabilities, and policy bypass exploits. It produces an urgent oversight debrief email for humanitarian donor committees.

Template

Role: Humanitarian AI Red-Teaming Director and Automated Governance Assurance Lead.

Context

  • Relief Program: {{grant_program_name}}
  • Red-Teaming Attack Vectors: {{adversarial_test_vectors}}
  • Allocation Variance Observed: {{funding_allocation_variance}}
  • Statutory Compliance Framework: {{compliance_standard}}
  • Emergency Override Mechanism: {{override_governance_protocol}}
  • Remediation Cutoff Date: {{remediation_deadline}}

Task

Draft an adversarial evaluation debrief email to the Donor Oversight Board and Chief Risk Officer detailing critical vulnerabilities uncovered during red-teaming of autonomous aid allocation systems.

Method

  1. Correlate targeted stress tests from {{adversarial_test_vectors}} against standard aid disbursement rules.
  2. Quantify financial and humanitarian exposure resulting from {{funding_allocation_variance}}.
  3. Assess compliance gaps relative to the {{compliance_standard}} regulatory framework.
  4. Evaluate system susceptibility to prompt injection, synthetic identity fraud, and sybil attacks.
  5. Audit the fail-safe reliability of the {{override_governance_protocol}} under simulated infrastructure degradation.
  6. Determine whether identified security gaps pose imminent threat to frontline aid delivery.
  7. Define time-critical mitigation tasks required prior to {{remediation_deadline}}.

Constraints

  • MUST write in an urgent, risk-focused, audit-grade communication style.
  • MUST NOT omit technical exploit vectors or downplay financial leakage risks.
  • MUST explicitly reference compliance mandates under {{compliance_standard}}.
  • Structure as an unclassified executive oversight debrief email.

Output format

  • Subject: [RISK ALERT] Red-Teaming Evaluation Findings for {{grant_program_name}}.
  • Threat Level Classification: Critical / High / Moderate with brief justification.
  • Exploit Vector Summary: Detailed review of successful vulnerabilities from {{adversarial_test_vectors}}.
  • Fiscal & Humanitarian Impact: Exposure calculation based on {{funding_allocation_variance}}.
  • Fail-Safe & Override Audit: Efficacy review of {{override_governance_protocol}}.
  • Mandatory Remediations: Required technical fixes due before {{remediation_deadline}}.
  • Word count: 450-600 words.

Self-review

  • Are the adversarial vulnerabilities explained with actionable clarity for non-technical board members?
  • Is the financial risk from {{funding_allocation_variance}} directly contextualized?
  • Is the cutoff date {{remediation_deadline}} prominently positioned?
AuraScore breakdown
79/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering10/12 · Adequate

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness3/5 · Adequate

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

ai-agents
agents-evaluation
public-sector-nonprofit
red-teaming
humanitarian-aid
adversarial-evaluation