Evaluation
AuraScore 83/100

Public Benefits Intake Agent Fairness Evaluation Brief

Evaluate automated constituent benefits triage agents for algorithmic bias, safety guardrails, and compliance risks.

Use this template when auditing autonomous public assistance intake workflows for systemic disparities or safety violations. It generates a comprehensive executive evaluation email tailored for nonprofit directors and public agency leadership.

Template

Role: Senior Algorithmic Justice Auditor and AI Evaluation Director for humanitarian operations.

Context

  • Target Public Entity: {{agency_name}}
  • Evaluated Intake System: {{intake_agent_name}}
  • Target Vulnerable Demographic: {{target_beneficiary_group}}
  • Evaluation Dataset Size: {{audit_sample_size}}
  • Acceptable Disparity Threshold: {{disparity_tolerance_rate}}
  • Critical Failure Incidents: {{observed_failure_modes}}

Task

Synthesize empirical evaluation results into an authoritative executive findings email for agency leadership, assessing whether the autonomous intake agent meets non-discrimination, safety, and reliability standards for public rollout.

Method

  1. Review the performance benchmarks of {{intake_agent_name}} across {{audit_sample_size}} test interactions.
  2. Calculate disparity ratios and false rejection rates specific to {{target_beneficiary_group}}.
  3. Benchmark observed variance against the mandatory {{disparity_tolerance_rate}} benchmark.
  4. Analyze the severity and root cause of {{observed_failure_modes}} in high-stakes public assistance routing.
  5. Evaluate automated guardrail interventions against statutory compliance and human-in-the-loop escalation criteria.
  6. Determine immediate operational go/no-go recommendation for {{agency_name}}.
  7. Outline mandatory pre-deployment remediation actions and recurring audit schedules.

Constraints

  • MUST structure output as a professional, high-stakes formal executive email.
  • MUST NOT use generic AI safety platitudes; cite concrete risk exposure for vulnerable claimants.
  • MUST present metric-driven findings alongside qualitative governance risks.
  • Tone must remain objective, rigorous, and legally cautious.

Output format

  • Subject line: Clear, risk-weighted formal notification.
  • Executive Summary: 3-4 sentence verdict on agent viability.
  • Evaluation Metrics Table: Key performance, parity metrics, and observed variances.
  • Root Cause Analysis of {{observed_failure_modes}}.
  • Formal Deployment Determination: Immediate release, conditional hold, or rejection.
  • Corrective Action Roadmap: 3-5 prioritized remediation requirements.
  • Word count: 450-650 words.

Self-review

  • Does the analysis explicitly address protection for {{target_beneficiary_group}}?
  • Are all {{observed_failure_modes}} mapped to clear remediation tasks?
  • Is the go/no-go determination unequivocally stated?
AuraScore breakdown
83/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering12/12 · Strong

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency7/10 · Adequate

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness3/5 · Adequate

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

ai-agents
agents-evaluation
public-sector-nonprofit
algorithmic-fairness
agent-audit
public-assistance