Evaluation
AuraScore 81/100

Multi-Source Research Synthesis Agent Quality Assessment Matrix

Assess literature review and synthesis agents on citation fidelity, bias detection, and cross-document reconciliation.

Use this evaluation framework to audit research synthesis AI agents working across high-volume academic and technical literature. It scores epistemic consistency, claim attribution accuracy, and contradiction resolution.

Template

Role: Senior Research Methodologist and Meta-Synthesis Evaluation Director.

Context

  • Synthesis agent pipeline: {{synthesis_pipeline_name}}
  • Corpus domain and scope: {{research_corpus_domain}}
  • Source document batch: {{source_document_batch}}
  • Benchmark ground truth synthesis: {{expert_ground_truth_synthesis}}
  • Minimum citation precision requirement: {{citation_precision_threshold}}
  • Primary conflict resolution policy: {{conflict_resolution_policy}}

Task

Produce an exhaustive multi-dimensional evaluation matrix that assesses how accurately {{synthesis_pipeline_name}} extracts, aggregates, and resolves competing claims from {{source_document_batch}} within {{research_corpus_domain}} against {{expert_ground_truth_synthesis}}.

Method

  1. Index and categorize all empirical claims, effect sizes, and qualitative assertions inside {{source_document_batch}}.
  2. Cross-reference agent-generated citations against exact source paragraphs to verify reference validity.
  3. Measure claim extraction precision against {{citation_precision_threshold}} to flag fabricated claims.
  4. Identify contradictory assertions across disparate sources within {{research_corpus_domain}}.
  5. Audit agent synthesis behavior against {{conflict_resolution_policy}} to verify balanced reporting of conflicting data.
  6. Evaluate thematic coverage and structural completeness against {{expert_ground_truth_synthesis}}.
  7. Score synthesis outputs across claim grounding, nuance retention, source weight balancing, and omission rate.
  8. Construct a systematic scoring matrix summarizing performance across all synthesis dimensions.

Constraints

  • MUST explicitly distinguish between hallucinated citations and misplaced in-line references.
  • MUST NOT score a synthesis as acceptable if contradictory source findings are silently omitted.
  • Every synthesis claim must be traced to verifiable document identifiers.
  • Score variance across evaluation dimensions must be statistically substantiated.

Output format

  • Synthesis Audit Matrix (Markdown table with columns: Thematic Dimension, Extraction Accuracy %, Citation Grounding %, Conflict Reconciliation Score, Omission Penalty, Overall Rating)
  • Claim Attribution Breakdown (Table detailing Source ID, Extracted Claim, Ground Truth Status, Attribution Verdict)
  • Discrepancy Analysis (Max 300 words analyzing systemic synthesis errors)
  • Quality Scorecard Summary (4-bullet metric summary)

Self-review

  • Check that every cited claim in the matrix links to an item in {{source_document_batch}}.
  • Validate that attribution verdicts align strictly with {{citation_precision_threshold}}.
  • Confirm conflict reconciliation scoring accurately reflects {{conflict_resolution_policy}}.
AuraScore breakdown
81/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering12/12 · Strong

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness3/5 · Adequate

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

ai-agents
agents-evaluation
complex-reasoning-analysis-math
research-synthesis
literature-review
citation-analysis