Testing
AuraScore 79/100

Chain-of-Thought Reasoning Benchmark Failure Analysis

Report complex logic drift, fallacy trends, and multi-step reasoning failures to AI evaluation stakeholders.

Use this template when synthetic evaluation suites identify breakdowns in deductive, inductive, or multi-hop logic within AI models. It crafts a comprehensive evaluation email detailing error patterns and test suite updates.

Template

Role: Staff AI Reasoning Evaluation Scientist leading multi-step logic benchmarking and epistemic testing pipelines.

Context

  • Target model checkpoint: {{model_checkpoint}}
  • Complex reasoning evaluation harness: {{reasoning_benchmark_suite}}
  • Categorized logical failure patterns: {{fallacy_taxonomy}}
  • Observed performance degradation rate: {{drift_rate_percentage}}
  • Inference testing compute envelope: {{compute_budget_limit}}
  • Primary stakeholder recipient: {{primary_stakeholder}}

Task

Compose an evaluation debrief email to {{primary_stakeholder}} analyzing multi-step deductive reasoning failures in {{model_checkpoint}}, diagnosing systematic logic fallacies, and prescribing evaluation harness refinements.

Method

  1. Aggregate reasoning trace failures recorded across {{reasoning_benchmark_suite}} for {{model_checkpoint}}.
  2. Classify individual step breakdowns against {{fallacy_taxonomy}} (e.g., premise hallucination, circular deduction, quantifier shifts).
  3. Correlate {{drift_rate_percentage}} against reasoning depth to identify the exact step count where logical coherence collapses.
  4. Evaluate token-level compute usage against {{compute_budget_limit}} to separate compute exhaustion from fundamental epistemic failures.
  5. Isolate one representative, end-to-end failure trace demonstrating intermediate reasoning drift.
  6. Formulate new regression test patterns designed to penalize invalid multi-hop leaps.
  7. Package the statistical and qualitative findings into an authoritative, high-signal debrief email.

Constraints

  • MUST categorize all failure instances using explicit terminology from {{fallacy_taxonomy}}.
  • MUST distinguish between epistemic knowledge voids and invalid structural inference steps.
  • MUST NOT recommend model deployment without establishing verified gate criteria.
  • Format all benchmark figures and error rates in structured, scannable markdown blocks.

Output format

Subject line: EVALUATION DEBRIEF: Multi-Step Reasoning Failures in {{model_checkpoint}}

  1. Executive Summary & Benchmark Scorecard: Overview of {{reasoning_benchmark_suite}} outcomes and {{drift_rate_percentage}}.
  2. Fallacy Taxonomy Breakdown: Distribution of observed reasoning errors.
  3. Dissected Failure Case Study: Multi-step prompt, erroneous reasoning path, and step-level root cause.
  4. Compute & Test Suite Recommendations: Actionable next steps within {{compute_budget_limit}} for {{primary_stakeholder}}.

Self-review

  • Is the breakdown between invalid inference and premise distortion sharply defined?
  • Does the email provide concrete reasoning traces rather than high-level generalities?
  • Are all variable placeholders populated consistently with testing context?
AuraScore breakdown
79/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering10/12 · Adequate

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness3/5 · Adequate

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

developers
developers-testing
complex-reasoning-analysis-math
ai-benchmarking
reasoning-evaluation
complex-analysis