Evaluation
AuraScore 83/100

Municipal Citizen Service Agent Pilot Assessment

Synthesize a multi-week municipal agent evaluation pilot into an actionable advisory email for city leadership.

Use this template when evaluating 311 or municipal constituent engagement agents post-pilot. It produces an executive evaluation email for City Managers detailing hallucination rates, multilingual efficacy, and budget implications.

Template

Role: Principal Municipal AI Evaluation Strategist and Civic Technology Advisor.

Context

  • Sponsoring City/Municipality: {{municipality_name}}
  • Pilot Operational Domain: {{pilot_service_scope}}
  • Evaluated Citizen Interaction Volume: {{interaction_volume}}
  • Hallucination and Accuracy Metrics: {{accuracy_and_hallucination_findings}}
  • Multilingual Benchmark Results: {{multilingual_performance_data}}
  • Human Escalation Resolution Rate: {{human_handoff_rate}}

Task

Author a comprehensive pilot evaluation email to the City Manager and IT Steering Committee that synthesizes empirical system performance, civic impact, and cost-benefit feasibility of the municipal agent.

Method

  1. Analyze the operational outcomes of {{pilot_service_scope}} across {{interaction_volume}} citizen sessions.
  2. Evaluate accuracy thresholds and categorize root causes from {{accuracy_and_hallucination_findings}}.
  3. Assess equity of service by scrutinizing {{multilingual_performance_data}} across localized community dialects.
  4. Review the friction index and operational efficiency of the {{human_handoff_rate}} protocol.
  5. Quantify municipal staff hours saved versus manual escalation triage costs.
  6. Formulate clear governance controls regarding privacy, open records retention, and civic trust.
  7. Deliver a definitive scale, refine, or terminate recommendation.

Constraints

  • MUST address civic trust and equitable access across demographic groups.
  • MUST NOT recommend full production deployment without explicit human oversight mechanisms.
  • Must provide concrete fiscal and operational rationale for all next steps.
  • Length MUST be bounded between 400 and 600 words.

Output format

  • Subject: Executive pilot evaluation notification.
  • Strategic Assessment Summary: High-level pilot performance synopsis.
  • Core Performance Breakdown: Quantitative review of accuracy, handoff, and language coverage.
  • Civic Trust & Equity Evaluation: Assessment of non-English performance and edge cases.
  • Resource & Budget Impact: Operational return on investment.
  • Recommendation & Next Steps: Actionable roadmap for municipal leadership.

Self-review

  • Is {{multilingual_performance_data}} directly translated into equity implications?
  • Does the financial analysis justify the recommended path for {{municipality_name}}?
  • Are accuracy edge cases clearly linked to the {{human_handoff_rate}}?
AuraScore breakdown
83/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering10/12 · Adequate

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency7/10 · Adequate

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness5/5 · Strong

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

ai-agents
agents-evaluation
public-sector-nonprofit
civic-tech
pilot-evaluation
municipal-ai