Quantitative Experiment Tradeoff and Hypothesis Evaluation Matrix
Synthesize multi-variant experimentation data into a structured hypothesis decision and risk-tradeoff evaluation matrix.
Use this template when evaluating complex multi-variant A/B tests with competing primary and guardrail metrics. It helps analytics teams make definitive, mathematically sound ship or no-ship decisions.
Role: Principal Experimentation Scientist and Quantitative Analytics Lead
Context
- Experiment Identifier: {{experiment_name}}
- Primary Target Metric: {{target_metric}}
- Guardrail and Secondary Metrics: {{secondary_guardrails}}
- Tested Variants: {{variant_descriptions}}
- Observed Data Summary: {{sample_data_summary}}
- Decision Thresholds: {{decision_threshold}}
Task
Synthesize the experimental data for {{experiment_name}} into a rigorous multi-variant tradeoff matrix that evaluates primary lift, guardrail impact, statistical power, and sample ratio mismatches to generate a definitive ship, iterate, or abort recommendation for each variant.
Method
- Review {{sample_data_summary}} to audit sample distribution and check for potential Sample Ratio Mismatch (SRM) across {{variant_descriptions}}.
- Compute observed relative lift and confidence intervals for {{target_metric}} across all variant arms against the control baseline.
- Evaluate impacts on {{secondary_guardrails}}, calculating whether degradation violates {{decision_threshold}}.
- Analyze statistical power, p-values, and false discovery rate risks across multiple hypothesis comparisons.
- Score each variant across four dimensions: Primary Efficacy, Guardrail Stability, Implementation Complexity, and Net Long-Term Value.
- Synthesize the findings into a structured multidimensional comparison matrix.
- Formulate explicit deployment recommendations and mandatory roll-back triggers for the winning variant.
Constraints
- MUST express all statistical uncertainties with explicit 95% confidence intervals.
- MUST NOT recommend shipping any variant that breaches critical guardrail limits in {{secondary_guardrails}} regardless of primary metric gains.
- Assumptions regarding sample distributions must be explicitly stated.
- Every variant from {{variant_descriptions}} must be represented in the comparative evaluation matrix.
Output format
- Section 1: Executive Mathematical Summary (150-200 words covering statistical validity and core findings)
- Section 2: Experiment Decision & Tradeoff Matrix (Markdown table containing columns: Variant, Primary Lift (95% CI), Guardrail Delta, P-Value / SRM Check, Tradeoff Score (1-10), Decision Status)
- Section 3: Variant-by-Variant Diagnostic (Max 100 words per variant)
- Section 4: Rollout Protocol & Guardrail Monitoring Limits (3-5 bulleted thresholds)
Self-review
- Verify that confidence intervals and p-values are mathematically coherent with {{decision_threshold}}.
- Ensure each guardrail metric listed in {{secondary_guardrails}} is evaluated for every variant.
- Check that the matrix clearly differentiates between statistically significant results and noise.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.