Evaluation
AuraScore 81/100

Multi-Hop Scientific Research Synthesis Agent Quality Gate

Systematically verify cross-disciplinary academic synthesis agents on citation fidelity, contradiction handling, and multi-hop epistemic claims.

Use this template when validating complex literature review and research synthesis agents that ingest dozens of conflicting empirical papers and generate structured systematic reviews. It ensures grounded claims, resolved contradictions, and zero citation hallucinations.

Template

Role: Lead Research Synthesis Auditor specializing in bibliometrics, epistemic conflict resolution, and evidence-based meta-analysis validation.

Context

  • Target research domain: {{target_research_domain}}
  • Contradiction resolution protocol: {{contradiction_resolution_protocol}}
  • Source qualification tiering: {{source_tier_criteria}}
  • Synthesis depth threshold: {{synthesis_depth_threshold}}
  • Claim attribution metric: {{claim_attribution_metric}}
  • Ingestion and retrieval pipeline: {{retrieval_agent_pipeline}}

Task

Produce an operational evaluation checklist that measures the evidence ground truth, cross-study synthesis coherence, and bibliographic integrity of an agent conducting multi-hop research analysis in {{target_research_domain}}.

Method

  1. Map out validation gates for raw source parsing from {{retrieval_agent_pipeline}} against {{source_tier_criteria}}.
  2. Define checks to detect source attribution degradation across multi-hop reasoning leaps.
  3. Construct evaluation criteria for statistical interpretation (e.g., p-values, sample sizes, confidence intervals, effect sizes).
  4. Specify deterministic tests for how the agent identifies and handles conflicting study findings via {{contradiction_resolution_protocol}}.
  5. Design rigorous checks to uncover extrapolation bias, over-generalization, and unstated causal claims.
  6. Formulate precise metrics based on {{claim_attribution_metric}} to catch hallucinated references and misattributed quotes.
  7. Establish validation checks for synthesis depth up to {{synthesis_depth_threshold}} across diverse empirical methodologies.
  8. Build summary integrity checks to guarantee non-omission of negative results or dissenting literature.

Constraints

  • MUST verify that every synthesized claim contains a verifiable pointer to an accepted source document.
  • MUST NOT accept vague summaries; each criterion must enforce explicit quantification of empirical findings.
  • Checklist items MUST evaluate both micro-level claim accuracy and macro-level thematic synthesis.
  • Must enforce strict disqualification for fabricated DOIs, authors, or statistical outputs.

Output format

A comprehensive quality assurance checklist divided into:

  • Part I: Source Provenance and Ingestion Integrity (4-5 checklist items)
  • Part II: Statistical Extraction & Multi-Hop Attribution (5-6 checklist items)
  • Part III: Contradiction Reconciliation & Epistemic Balance (4-5 checklist items)
  • Part IV: Structural Coherence & Over-Claim Prevention (4-5 checklist items)
  • Quality threshold summary specifying mandatory blockers vs. warnings.

Self-review

  • Verify that the checklist tests multi-source synthesis, not just single-document extraction.
  • Ensure all variables, including {{contradiction_resolution_protocol}}, are directly applied.
  • Check that citation fidelity checks explicitly prevent subtle author-year mismatches.
AuraScore breakdown
81/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering12/12 · Strong

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness3/5 · Adequate

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

ai-agents
agents-evaluation
complex-reasoning-analysis-math
research-synthesis
literature-review
epistemic-validation