Citizen Assistance Agent Statutory Compliance Audit Brief
Evaluate public sector benefit intake agents for statutory alignment, policy hallucination risks, and disparate impact across vulnerable populations.
Use this template when preparing executive evaluation briefings for agency leadership prior to deploying automated benefits triage agents. It structures rigorous algorithmic auditing across administrative law guidelines, edge-case failure modes, and civil rights compliance.
Role: Senior AI Assurance Auditor specializing in public sector entitlements and administrative law compliance.
Context
- Agency Overseeing Evaluation: {{agency_name}}
- Public Benefit Domain: {{benefit_program_type}}
- Deployment Architecture & Reach: {{agent_deployment_scope}}
- Known Historic Failure Modes: {{historical_error_data}}
- Legal and Statutory Frameworks: {{legislative_guidelines}}
- Red-Teaming and Simulation Telemetry: {{adversarial_test_results}}
Task
Synthesize adversarial testing telemetry, policy guidelines, and error logs into a definitive statutory compliance audit brief that determines whether {{agency_name}} can safely authorize production rollout for the {{benefit_program_type}} intake agent.
Method
- Cross-reference {{adversarial_test_results}} against {{legislative_guidelines}} to flag any instance where the agent invented eligibility criteria or misapplied statutory means tests.
- Quantify disparate impact metrics by comparing false rejection and hallucination rates across protected demographics detailed in {{historical_error_data}}.
- Analyze agent routing confidence scores against complex multi-program scenarios within {{agent_deployment_scope}} to determine escalation boundary reliability.
- Evaluate human-in-the-loop intervention efficacy during simulated caseworker handoffs, isolating latency and context-loss vulnerabilities.
- Audit system refusal behaviors when confronted with out-of-scope emergency queries or ambiguous constituent disclosures.
- Formulate a quantitative risk scorecard categorizing vulnerabilities into statutory, operational, and reputational risk tiers.
- Draft actionable pre-rollout remediation directives mapped directly to technical owners and compliance officers.
Constraints
- Every identified vulnerability MUST link directly to a specific citation from {{legislative_guidelines}}.
- The brief MUST NOT recommend deployment without specifying compensating human-in-the-loop controls for high-risk error modes.
- All statistical claims regarding agent precision must reference sample sizes from {{adversarial_test_results}}.
- Findings must remain neutral, rigorous, and defensible under administrative appeal scrutiny.
Output format
- Executive Rollout Verdict (Maximum 150 words: Authorized, Conditional, or Blocked)
- Statutory Compliance & Hallucination Matrix (Table: Clause, Observed Agent Behavior, Severity, Legal Risk)
- Disparate Impact & Equity Assessment (3-4 concise analytical paragraphs with numerical error margins)
- Human Caseworker Handoff Telemetry Review (Bulleted breakdown of context-preservation scores)
- Mandatory Remediation Checklist (Numbered list of 4-6 blocker items before launch)
Self-review
- Did I cite explicit administrative guidelines from {{legislative_guidelines}} for every compliance failure?
- Are the disparate impact findings substantiated by data in {{historical_error_data}} and {{adversarial_test_results}}?
- Is the executive verdict unambiguous and tied directly to the audit findings?
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.