Evaluation
AuraScore 81/100

Multi-Hop Research Synthesis Agent Stress Audit

Deliver a comprehensive evaluation email assessing complex multi-document synthesis and contradiction-handling capabilities.

Use this template when auditing multi-hop research agents that synthesize conflicting academic or scientific literature. It structures rigorous evaluation metrics, counterfactual robustness, and gatekeeper recommendations into an executive briefing email.

Template

Role: Senior Cognitive Evaluation Scientist & AI Safety Auditor

Context

  • Pipeline Version: {{agent_pipeline_version}}
  • Lead Evaluator: {{primary_evaluator_name}}
  • Target Domain: {{corpus_domain}}
  • Permissible Error Margin: {{hallucination_tolerance_rate}}
  • Adversarial Dataset: {{counterfactual_challenge_set}}
  • Addressee: {{engineering_director}}

Task

Compose an audit evaluation email to {{engineering_director}} detailing the stress-test results of {{agent_pipeline_version}} across multi-document synthesis tasks in {{corpus_domain}}, determining whether the system satisfies {{hallucination_tolerance_rate}} under adversarial conditions.

Method

  1. Define synthesis accuracy across multi-source dependency graphs.
  2. Quantify source attribution precision and cross-document citation grounding.
  3. Audit agent handling of contradictory evidence within {{corpus_domain}}.
  4. Measure degradation curves when exposed to {{counterfactual_challenge_set}}.
  5. Isolate intermediate reasoning drift in chains longer than five hops.
  6. Compare empirical failure rates against {{hallucination_tolerance_rate}}.
  7. Assess latency and context token overhead during complex reconciliation cycles.
  8. Generate definitive release gating criteria authored by {{primary_evaluator_name}}.

Constraints

  • MUST cite quantitative attribution error rates alongside qualitative failure examples.
  • MUST NOT approve release if counterfactual vulnerability exceeds {{hallucination_tolerance_rate}}.
  • Maintain an authoritative, audit-ready tone suited for executive leadership.
  • Restrict the email body to 500-750 words.
  • Structure findings into precisely 4 numbered analytical sections.

Output format

Subject: [AUDIT REPORT] {{agent_pipeline_version}} Multi-Hop Synthesis Evaluation

  • Formal Greeting to {{engineering_director}}
    1. Executive Summary & Gating Decision
    1. Source Attribution & Contradiction Resolution Metrics
    1. Counterfactual Stress-Test Results (via {{counterfactual_challenge_set}})
    1. Required Algorithmic Hardening Measures
  • Formal Sign-off by {{primary_evaluator_name}}

Self-review

  • Confirm attribution error rates are explicitly benchmarked against {{hallucination_tolerance_rate}}.
  • Validate that all multi-hop reasoning edge cases reflect {{corpus_domain}} dynamics.
  • Ensure each numbered section meets the specified depth criteria.
AuraScore breakdown
81/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering8/12 · Adequate

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency7/10 · Adequate

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness5/5 · Strong

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

ai-agents
agents-evaluation
complex-reasoning-analysis-math
research-synthesis
multi-hop-reasoning
audit