Evaluation
AuraScore 81/100

Autonomous Returns Agent Policy Adherence Evaluation Brief

Evaluate autonomous customer service return and refund agent decisions against retailer policy and fraud thresholds.

Use this template when auditing autonomous customer service agents handling return authorizations, appeasements, and chargeback disputes. It assesses policy compliance, edge-case handling, and financial leakage across digital touchpoints.

Template

Role: Principal Conversational AI Quality and Risk Auditor with fifteen years evaluating automated retail customer service systems.

Context

  • Brand and operating unit: {{retail_brand_name}}
  • Operational workflow scope: {{agent_workflow_scope}}
  • Benchmark dispute dataset: {{historical_dispute_logs}}
  • Anti-fraud policy tolerance: {{fraud_tolerance_threshold}}
  • Customer touchpoints evaluated: {{channel_touchpoints}}
  • Analysis evaluation timeframe: {{evaluation_timeframe}}

Task

Generate an executive evaluation brief assessing the automated agent's policy compliance, edge-case adjudication quality, and financial leakage risk across historical return and refund interactions to determine readiness for wider rollout.

Method

  1. Ingest {{historical_dispute_logs}} and segment interactions by return reason codes, customer tier, and transaction size within {{evaluation_timeframe}}.
  2. Evaluate the automated decision path in {{agent_workflow_scope}} against official merchandise return rules and {{fraud_tolerance_threshold}}.
  3. Measure deterministic policy adherence versus autonomous hallucinated concession rates across {{channel_touchpoints}}.
  4. Identify false-positive fraud blocks that create customer friction versus false-negative approvals leading to inventory loss.
  5. Benchmark agent sentiment handling, escalation trigger reliability, and human handoff latency on disputed claims.
  6. Quantify net financial variance between autonomous settlement values and baseline human agent settlements for {{retail_brand_name}}.
  7. Formulate a risk-tiered remediation backlog for unhandled edge cases, customer appeasement caps, and policy ambiguities.

Constraints

  • All evaluations MUST classify leakage into preventable merchant errors, customer fraud, or system edge-case failures.
  • The brief MUST NOT recommend full automation expansion for product categories with an automated policy violation rate above 1.5%.
  • Analysis MUST explicitly separate physical inventory return exceptions from non-return appeasement grants.
  • Maintain focus strictly on agent decision quality without diagnosing underlying conversational UI latency.

Output format

  • Executive Summary: Exactly 150 words summarizing agent compliance score and net financial risk.
  • Policy Adherence Matrix: Table comparing 5 key return workflows, compliance pass rate, and primary failure modes.
  • Fraud and Leakage Breakdown: 3 detailed paragraphs analyzing concession overruns and false positives.
  • Remediation Action Plan: 4 numbered tactical directives for engineering and policy teams.

Self-review

  • Confirm all context variables appear naturally in the analytical framing.
  • Verify that both false-positive customer friction and false-negative inventory loss are quantified.
  • Check that the output format adheres strictly to the section names and word count constraints.
AuraScore breakdown
81/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering10/12 · Adequate

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness5/5 · Strong

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

ai-agents
agents-evaluation
retail-consumer-goods
retail
customer-service
returns