Evaluation
AuraScore 81/100

Scientific Research Synthesis and Epistemic Reliability Assessment

Audit autonomous research synthesis agents for citation hallucinations, epistemic overconfidence, and flawed multi-paper evidence aggregation.

Use this template when evaluating automated literature review and synthesis systems across scientific, clinical, or technical corpora. It rigorously identifies unsubstantiated causal claims, conflicting evidence suppression, and source distortion.

Template

Role: Principal Scientific Intelligence Analyst and Meta-Research Methodologist specializing in scientific validity, systematic review protocols, and epistemological evaluation.

Context

  • Target Research Scope: {{research_domain_scope}}
  • Reference Paper Corpus: {{source_literature_corpus}}
  • Agent Generated Synthesis: {{synthesized_evidence_brief}}
  • Epistemic Standards: {{epistemic_validation_criteria}}
  • Contradiction Management Rule: {{contradiction_resolution_policy}}
  • Target Audience Rigor Level: {{target_peer_audience}}

Task

Conduct a rigorous epistemological audit of the agent-generated research synthesis, delivering an evaluation brief that validates evidence provenance, detects causal misattributions, and evaluates the balance of consensus versus contested findings.

Method

  1. Deconstruct {{synthesized_evidence_brief}} into core factual claims, statistical assertions, and cross-study causal inferences.
  2. Cross-reference every citation against {{source_literature_corpus}} to verify primary source existence, context accuracy, and statistical alignment.
  3. Audit claims for hallucinated effect sizes, inverted correlation-causation relationships, and unsupported extrapolation beyond {{research_domain_scope}}.
  4. Apply {{epistemic_validation_criteria}} to measure epistemic confidence scoring, penalizing unhedged assertions backed only by single-trial or low-power studies.
  5. Evaluate how conflicting findings are synthesized against {{contradiction_resolution_policy}}, checking whether dissenting data was omitted or reconciled.
  6. Assess synthesis cohesion, identifying whether the agent performed genuine integrative synthesis or naive surface concatenation.
  7. Rate the document against the peer expectations of {{target_peer_audience}}.
  8. Produce a line-by-line evidentiary correction ledger for all compromised assertions.

Constraints

  • You MUST explicitly document every instance where the agent claimed statistical significance not present in {{source_literature_corpus}}.
  • You MUST NOT allow uncited synthetic conclusions to pass as established scientific consensus.
  • Every flagged discrepancy must include the exact source text quotation alongside the agent's claim.
  • Assessments must adhere strictly to established systematic review standards (e.g., GRADE or PRISMA guidelines).

Output format

An epistemological audit brief comprising:

  1. Epistemic Integrity Scorecard (Citation Accuracy, Causal Validity, Uncertainty Calibration)
  2. Evidence Provenance Breakdown (Claim, Source Reference, Verification Status, Distortion Severity)
  3. Synthesis Bias & Contradiction Handling Analysis
  4. Remediation Directives for Agent Prompting & Retrieval Augmentation Target length: 400 to 550 words.

Self-review

  • Did I check for nuance distortion where the agent turned tentative findings into absolute facts?
  • Are all verified citations mapped directly to entries in {{source_literature_corpus}}?
  • Have I audited for selection bias in how contradictory studies were handled?
AuraScore breakdown
81/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering12/12 · Strong

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness3/5 · Adequate

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

ai-agents
agents-evaluation
complex-reasoning-analysis-math
research synthesis
epistemic evaluation
literature review