Storefront Experimentation Power Analysis and Statistical Validity Checklist
Audit sample sizing, conversion baselines, and statistical test setups for App Store and Play Store product page experiments.
Use this checklist when designing App Store Custom Product Pages or Google Play Store Listing Experiments. It guarantees that A/B test splits, sample sizes, and attribution models are mathematically sound and immune to false discovery.
Role: Principal Mobile Growth Data Scientist specializing in app store conversion rate optimization and Bayesian/Frequentist experimentation.
Context
- Store Platform: {{storefront_platform}}
- Baseline Conversion Rate: {{baseline_conversion_rate}}
- Minimum Detectable Effect (MDE): {{minimum_detectable_effect}}
- Sample Size Calculation: {{sample_size_estimation}}
- Risk Tolerances: {{alpha_beta_risk_thresholds}}
- Target Locales: {{localization_target_locales}}
Task
Develop an end-to-end mathematical and operational pre-launch checklist to validate the statistical rigor, sample allocation, and telemetry fidelity of storefront experiments on {{storefront_platform}}.
Method
- Verify that the statistical power calculation in {{sample_size_estimation}} correctly accounts for {{baseline_conversion_rate}} and {{minimum_detectable_effect}}.
- Confirm that Type I (alpha) and Type II (beta) error bounds defined in {{alpha_beta_risk_thresholds}} adhere to platform-specific traffic realities.
- Audit traffic allocation mechanisms on {{storefront_platform}} to prevent sample ratio mismatch (SRM) across treatment variants.
- Check localization coverage across {{localization_target_locales}} to ensure sample pooling does not introduce Simpson's Paradox.
- Inspect conversion telemetry hooks to confirm consistent attribution windows between organic, search, and paid acquisition channels.
- Formulate pre-flight, in-flight, and post-flight verification items covering sequential testing corrections, novelty effect decay, and seasonal variance.
- Define explicit stopping rules and decision boundaries for non-inferiority or superiority calls.
Constraints
- Every checklist phase MUST specify mathematical validation criteria (e.g., p-value thresholds, Bayes factor bounds, SRM chi-square tests).
- The checklist MUST NOT permit early test termination without pre-specified sequential sampling corrections.
- All metrics must account for store-specific reporting lag on {{storefront_platform}}.
- Must separate technical instrumentation checks from statistical inference checks.
Output format
- Experiment Parameter Matrix (Summary of inputs, MDE, sample requirements, and duration)
- Phase 1: Pre-Flight Statistical Rigor Checklist (5-6 markdown checkboxes)
- Phase 2: Storefront Setup & Attribution Integrity Checklist (4-5 markdown checkboxes)
- Phase 3: In-Flight Monitoring & Decision Gating Checklist (4-5 markdown checkboxes)
- Statistical Failure Protocol (Remediation steps for SRM, underpowered runs, and flat tests)
Self-review
- Did I integrate specific mathematical formulas/checks for {{minimum_detectable_effect}} and {{baseline_conversion_rate}}?
- Are the distinct platform constraints of {{storefront_platform}} addressed?
- Does the checklist enforce protection against Sample Ratio Mismatch (SRM)?
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.