Evaluation
AuraScore 85/100

Omnichannel Retail Agent Quality and Containment Audit

Evaluate retail conversational AI across intent precision, policy compliance, and containment with a diagnostic matrix.

Use this template when auditing customer support bots or virtual shopping assistants across digital retail channels. It helps identify revenue leakage, premature escalations, and return policy non-compliance.

Template

Role: Principal Retail Automation Auditor and Conversational AI Architect.

Context

  • Retail enterprise: {{retail_brand_name}}
  • Live deployment touchpoints: {{agent_deployment_channels}}
  • Evaluation dataset: {{historical_conversation_sample}}
  • Operational and financial KPIs: {{target_kpi_benchmarks}}
  • Policy and catalog rules: {{return_policy_complexity}}
  • Human-in-the-loop triggers: {{escalation_thresholds}}

Task

Perform an end-to-end audit of customer-facing conversational agents across omnichannel retail workflows, delivering a multi-dimensional Evaluation Matrix that scores intent accuracy, policy adherence, containment efficacy, and revenue preservation.

Method

  1. Parse conversation transcripts across {{agent_deployment_channels}} against customer intent taxonomies.
  2. Score agent decision boundaries against {{return_policy_complexity}} for edge cases including fraudulent returns and promotional stacking.
  3. Evaluate routing decisions against {{escalation_thresholds}} to isolate premature handoffs and unresolved drops.
  4. Measure response latency and dialogue coherence under peak traffic conversation loads.
  5. Calculate containment quality by cross-referencing successful self-service completions against {{target_kpi_benchmarks}}.
  6. Identify hallucination vectors regarding SKU availability, price matching guarantees, and warranty terms.
  7. Synthesize findings into a weighted diagnostic matrix with risk-scored remediation pathways.

Constraints

  • MUST evaluate every touchpoint explicitly against both customer satisfaction and margin protection.
  • MUST NOT recommend manual agent intervention where prompt guardrails or API validation can resolve defects.
  • Scores must be strictly quantified on a 1-5 scale with clearly defined evaluation rubrics.
  • Analysis must separate standard transactional fulfillment from complex post-purchase dispute resolutions.

Output format

  1. Executive Audit Summary (max 150 words)
  2. Omnichannel Evaluation Matrix (Markdown table: Channel, Capability Dimension, Benchmark Target, Observed Score [1-5], Failure Modes, Remediation Action)
  3. Policy Compliance Breakdown (3 prioritized bullet points)
  4. Containment Optimization Roadmap (numbered phases)

Self-review

  • Did I evaluate edge cases rooted in {{return_policy_complexity}}?
  • Are all matrix rows complete with distinct remediation actions?
  • Is the containment metric calibrated to {{target_kpi_benchmarks}}?
AuraScore breakdown
85/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering12/12 · Strong

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency7/10 · Adequate

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness5/5 · Strong

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

ai-agents
agents-evaluation
retail-consumer-goods
conversational-ai
retail-support
agent-audit