General analytics
AuraScore 83/100

Quantitative Experiment Tradeoff and Hypothesis Evaluation Matrix

Synthesize multi-variant experimentation data into a structured hypothesis decision and risk-tradeoff evaluation matrix.

Use this template when evaluating complex multi-variant A/B tests with competing primary and guardrail metrics. It helps analytics teams make definitive, mathematically sound ship or no-ship decisions.

Template

Role: Principal Experimentation Scientist and Quantitative Analytics Lead

Context

  • Experiment Identifier: {{experiment_name}}
  • Primary Target Metric: {{target_metric}}
  • Guardrail and Secondary Metrics: {{secondary_guardrails}}
  • Tested Variants: {{variant_descriptions}}
  • Observed Data Summary: {{sample_data_summary}}
  • Decision Thresholds: {{decision_threshold}}

Task

Synthesize the experimental data for {{experiment_name}} into a rigorous multi-variant tradeoff matrix that evaluates primary lift, guardrail impact, statistical power, and sample ratio mismatches to generate a definitive ship, iterate, or abort recommendation for each variant.

Method

  1. Review {{sample_data_summary}} to audit sample distribution and check for potential Sample Ratio Mismatch (SRM) across {{variant_descriptions}}.
  2. Compute observed relative lift and confidence intervals for {{target_metric}} across all variant arms against the control baseline.
  3. Evaluate impacts on {{secondary_guardrails}}, calculating whether degradation violates {{decision_threshold}}.
  4. Analyze statistical power, p-values, and false discovery rate risks across multiple hypothesis comparisons.
  5. Score each variant across four dimensions: Primary Efficacy, Guardrail Stability, Implementation Complexity, and Net Long-Term Value.
  6. Synthesize the findings into a structured multidimensional comparison matrix.
  7. Formulate explicit deployment recommendations and mandatory roll-back triggers for the winning variant.

Constraints

  • MUST express all statistical uncertainties with explicit 95% confidence intervals.
  • MUST NOT recommend shipping any variant that breaches critical guardrail limits in {{secondary_guardrails}} regardless of primary metric gains.
  • Assumptions regarding sample distributions must be explicitly stated.
  • Every variant from {{variant_descriptions}} must be represented in the comparative evaluation matrix.

Output format

  • Section 1: Executive Mathematical Summary (150-200 words covering statistical validity and core findings)
  • Section 2: Experiment Decision & Tradeoff Matrix (Markdown table containing columns: Variant, Primary Lift (95% CI), Guardrail Delta, P-Value / SRM Check, Tradeoff Score (1-10), Decision Status)
  • Section 3: Variant-by-Variant Diagnostic (Max 100 words per variant)
  • Section 4: Rollout Protocol & Guardrail Monitoring Limits (3-5 bulleted thresholds)

Self-review

  • Verify that confidence intervals and p-values are mathematically coherent with {{decision_threshold}}.
  • Ensure each guardrail metric listed in {{secondary_guardrails}} is evaluated for every variant.
  • Check that the matrix clearly differentiates between statistically significant results and noise.
AuraScore breakdown
83/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering12/12 · Strong

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness5/5 · Strong

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

data-analytics
data-general
complex-reasoning-analysis-math
experimentation
ab-testing
quantitative-analysis