Multi-Hop Research Synthesis Agent Stress Audit
Deliver a comprehensive evaluation email assessing complex multi-document synthesis and contradiction-handling capabilities.
Use this template when auditing multi-hop research agents that synthesize conflicting academic or scientific literature. It structures rigorous evaluation metrics, counterfactual robustness, and gatekeeper recommendations into an executive briefing email.
Role: Senior Cognitive Evaluation Scientist & AI Safety Auditor
Context
- Pipeline Version: {{agent_pipeline_version}}
- Lead Evaluator: {{primary_evaluator_name}}
- Target Domain: {{corpus_domain}}
- Permissible Error Margin: {{hallucination_tolerance_rate}}
- Adversarial Dataset: {{counterfactual_challenge_set}}
- Addressee: {{engineering_director}}
Task
Compose an audit evaluation email to {{engineering_director}} detailing the stress-test results of {{agent_pipeline_version}} across multi-document synthesis tasks in {{corpus_domain}}, determining whether the system satisfies {{hallucination_tolerance_rate}} under adversarial conditions.
Method
- Define synthesis accuracy across multi-source dependency graphs.
- Quantify source attribution precision and cross-document citation grounding.
- Audit agent handling of contradictory evidence within {{corpus_domain}}.
- Measure degradation curves when exposed to {{counterfactual_challenge_set}}.
- Isolate intermediate reasoning drift in chains longer than five hops.
- Compare empirical failure rates against {{hallucination_tolerance_rate}}.
- Assess latency and context token overhead during complex reconciliation cycles.
- Generate definitive release gating criteria authored by {{primary_evaluator_name}}.
Constraints
- MUST cite quantitative attribution error rates alongside qualitative failure examples.
- MUST NOT approve release if counterfactual vulnerability exceeds {{hallucination_tolerance_rate}}.
- Maintain an authoritative, audit-ready tone suited for executive leadership.
- Restrict the email body to 500-750 words.
- Structure findings into precisely 4 numbered analytical sections.
Output format
Subject: [AUDIT REPORT] {{agent_pipeline_version}} Multi-Hop Synthesis Evaluation
- Formal Greeting to {{engineering_director}}
-
- Executive Summary & Gating Decision
-
- Source Attribution & Contradiction Resolution Metrics
-
- Counterfactual Stress-Test Results (via {{counterfactual_challenge_set}})
-
- Required Algorithmic Hardening Measures
- Formal Sign-off by {{primary_evaluator_name}}
Self-review
- Confirm attribution error rates are explicitly benchmarked against {{hallucination_tolerance_rate}}.
- Validate that all multi-hop reasoning edge cases reflect {{corpus_domain}} dynamics.
- Ensure each numbered section meets the specified depth criteria.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.