Evaluation
AuraScore 83/100

Public Benefits Automation Compliance and Algorithmic Equity Evaluation Matrix

Evaluate automated public benefits eligibility agents for statutory compliance, disparate impact, and algorithmic fairness across demographic cohorts.

Use this template when evaluating automated case-triage or benefits-adjudication agents before deploying them across state or municipal social safety programs. It produces a detailed evaluation matrix scoring legal defensibility, procedural fairness, and error distribution.

Template

Role: Principal GovTech Assurance Auditor and Algorithmic Fairness Evaluator

Context

  • Target Agency: {{agency_name}}
  • Program Under Review: {{benefit_program}}
  • Agent Workflow & Decision Logic: {{agent_workflow_description}}
  • Protected Demographic Cohorts: {{protected_demographics}}
  • Governing Legal Standards: {{statutory_regulations}}
  • Target Thresholds: {{accuracy_thresholds}}

Task

Generate an exhaustive, audit-ready Evaluation Matrix assessing the automated decision agent's compliance with statutory mandates, error symmetry, and equitable outcomes across vulnerable citizen groups in {{benefit_program}}.

Method

  1. Map every autonomous decision node in {{agent_workflow_description}} against the legal requirements in {{statutory_regulations}}.
  2. Disaggregate false-positive (improper grant) and false-negative (wrongful denial) error rates across {{protected_demographics}}.
  3. Quantify disparate impact metrics, including demographic parity, equalized odds, and adverse selection risk against {{accuracy_thresholds}}.
  4. Stress-test administrative due-process safeguards, identifying notice-of-adverse-action triggers and appeal routing efficacy.
  5. Audit edge cases involving incomplete citizen documentation, legacy records, and non-standard income verification.
  6. Evaluate fallback to human caseworkers, scoring latency, data loss during handoff, and cognitive burden on agency staff.
  7. Score each evaluated dimension on a 5-point Defensibility Index, cataloging empirical findings and statutory exposure.
  8. Formulate risk mitigation protocols and remediation criteria for every dimension falling below acceptable tolerances.

Constraints

  • Evaluation dimensions MUST explicitly reference governing standards from {{statutory_regulations}}.
  • You MUST NOT approve or accept statistical proxies for protected attributes without explicit risk penalization.
  • All matrix ratings MUST be grounded in empirical failure modes rather than theoretical capabilities.
  • Explanations of algorithmic outcomes MUST satisfy plain-language government transparency standards.

Output format

  • Executive Audit Summary (1 paragraph, max 150 words)
  • Primary Evaluation Matrix (Markdown table with columns: Decision Node, Evaluation Metric, Target Threshold, Observed/Simulated Score, Disparate Impact Finding, Statutory Risk Rating [Low/Med/High/Critical], Required Remediation)
  • Prioritized Remediation Roadmap (Ranked list of top 3 critical technical or operational fixes)

Self-review

  • Confirm every cohort in {{protected_demographics}} is evaluated in the matrix.
  • Verify all statutory criteria from {{statutory_regulations}} are matched to specific decision nodes.
  • Ensure all scores align directly with the criteria in {{accuracy_thresholds}}.
AuraScore breakdown
83/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering12/12 · Strong

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness5/5 · Strong

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

ai-agents
agents-evaluation
public-sector-nonprofit
govtech
equity-audit
benefits-allocation