Evaluation
AuraScore 81/100

Municipal Civic Assistant Trust and Accessibility Performance Benchmark Matrix

Benchmark public-facing municipal AI agents across accessibility compliance, multilingual fidelity, data sovereignty, and service resolution.

Use this prompt to systematically evaluate local government chatbots and automated civic service agents before public release. It delivers a comprehensive evaluation matrix covering civic trust, WCAG accessibility, dialectal parity, and PII protection.

Template

Role: Chief Digital Ethics Officer for Municipal Public Services

Context

  • Municipality: {{municipality_name}}
  • Covered Public Services: {{service_catalog}}
  • Linguistic & Dialectal Cohorts: {{linguistic_cohorts}}
  • Digital Accessibility Standards: {{accessibility_standards}}
  • Privacy & Data Sovereignty Mandates: {{data_sovereignty_mandates}}
  • Current Baseline Error Rate: {{baseline_error_rate}}

Task

Develop a comprehensive Civic Trust and Accessibility Evaluation Matrix to benchmark the municipal AI agent against public sector accessibility, privacy, and service delivery standards across all items in {{service_catalog}}.

Method

  1. Establish benchmark criteria across four civic pillars: Accessibility, Multilingual Fidelity, Privacy Preservation, and Resolution Accuracy.
  2. Cross-reference automated responses with technical accessibility guidelines in {{accessibility_standards}} (including screen reader parseability and cognitive clarity).
  3. Test conversational accuracy and cultural nuance across every group identified in {{linguistic_cohorts}}.
  4. Measure PII handling, data retention, and zero-trust boundary enforcement against {{data_sovereignty_mandates}}.
  5. Audit the agent's ability to navigate complex municipal procedures (permitting, tax relief, utility assistance) compared to {{baseline_error_rate}}.
  6. Evaluate "I don't know" boundaries, deceptive certainty detection, and seamless transfer to municipal ombudsman or desk staff.
  7. Populate a multi-tier Evaluation Matrix scoring each service domain with quantitative performance indicators.
  8. Produce a governance determination outlining conditional launch prerequisites for municipal leadership.

Constraints

  • The evaluation MUST grade each linguistic cohort in {{linguistic_cohorts}} independently without aggregating scores.
  • The agent MUST NOT retain or log identifiable civic inquiries containing PII under {{data_sovereignty_mandates}}.
  • Accessibility benchmarks MUST include specific pass/fail criteria for assistive tech interoperability.
  • Matrix findings MUST distinguish between content errors and systemic model safety failures.

Output format

  • Municipal Readiness Appraisal (1 paragraph, max 140 words)
  • Civic Agent Performance Matrix (Markdown table with columns: Service Category, Evaluated Dimension, Standard Referenced, Target Metric, Assessed Score, Accessibility/Equity Gap, Governance Action)
  • Launch Readiness Determination (Structured checklist with Go / No-Go sign-off gates)

Self-review

  • Confirm every service in {{service_catalog}} is represented in the matrix.
  • Check that each mandate in {{data_sovereignty_mandates}} is mapped to a validation check.
  • Verify that performance is explicitly compared against {{baseline_error_rate}}.
AuraScore breakdown
81/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering12/12 · Strong

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness3/5 · Adequate

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

ai-agents
agents-evaluation
public-sector-nonprofit
smart-cities
accessibility
civic-tech