Chain-of-Thought Reasoning Benchmark Failure Analysis
Report complex logic drift, fallacy trends, and multi-step reasoning failures to AI evaluation stakeholders.
Use this template when synthetic evaluation suites identify breakdowns in deductive, inductive, or multi-hop logic within AI models. It crafts a comprehensive evaluation email detailing error patterns and test suite updates.
Role: Staff AI Reasoning Evaluation Scientist leading multi-step logic benchmarking and epistemic testing pipelines.
Context
- Target model checkpoint: {{model_checkpoint}}
- Complex reasoning evaluation harness: {{reasoning_benchmark_suite}}
- Categorized logical failure patterns: {{fallacy_taxonomy}}
- Observed performance degradation rate: {{drift_rate_percentage}}
- Inference testing compute envelope: {{compute_budget_limit}}
- Primary stakeholder recipient: {{primary_stakeholder}}
Task
Compose an evaluation debrief email to {{primary_stakeholder}} analyzing multi-step deductive reasoning failures in {{model_checkpoint}}, diagnosing systematic logic fallacies, and prescribing evaluation harness refinements.
Method
- Aggregate reasoning trace failures recorded across {{reasoning_benchmark_suite}} for {{model_checkpoint}}.
- Classify individual step breakdowns against {{fallacy_taxonomy}} (e.g., premise hallucination, circular deduction, quantifier shifts).
- Correlate {{drift_rate_percentage}} against reasoning depth to identify the exact step count where logical coherence collapses.
- Evaluate token-level compute usage against {{compute_budget_limit}} to separate compute exhaustion from fundamental epistemic failures.
- Isolate one representative, end-to-end failure trace demonstrating intermediate reasoning drift.
- Formulate new regression test patterns designed to penalize invalid multi-hop leaps.
- Package the statistical and qualitative findings into an authoritative, high-signal debrief email.
Constraints
- MUST categorize all failure instances using explicit terminology from {{fallacy_taxonomy}}.
- MUST distinguish between epistemic knowledge voids and invalid structural inference steps.
- MUST NOT recommend model deployment without establishing verified gate criteria.
- Format all benchmark figures and error rates in structured, scannable markdown blocks.
Output format
Subject line: EVALUATION DEBRIEF: Multi-Step Reasoning Failures in {{model_checkpoint}}
- Executive Summary & Benchmark Scorecard: Overview of {{reasoning_benchmark_suite}} outcomes and {{drift_rate_percentage}}.
- Fallacy Taxonomy Breakdown: Distribution of observed reasoning errors.
- Dissected Failure Case Study: Multi-step prompt, erroneous reasoning path, and step-level root cause.
- Compute & Test Suite Recommendations: Actionable next steps within {{compute_budget_limit}} for {{primary_stakeholder}}.
Self-review
- Is the breakdown between invalid inference and premise distortion sharply defined?
- Does the email provide concrete reasoning traces rather than high-level generalities?
- Are all variable placeholders populated consistently with testing context?
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.