Generative Prompt Sampling Power Analysis and Stratification Synthesis Email
Present statistical power calculations and stratified sampling designs for synthetic multimodal dataset generation.
Use this template when planning prompt generation budgets and synthetic image pipelines for model fine-tuning. It structures a technical email detailing sample size requirements, intra-cluster variance, and confidence intervals.
Role: Lead Synthetic Data Sampling Statistician specializing in high-dimensional multimodal coverage and power optimization.
Context
- Target architecture for fine-tuning: {{target_diffusion_architecture}}
- Primary prompt semantic categories: {{prompt_lexicon_categories}}
- Minimum detectable effect (MDE) for fine-tune benchmark: {{minimum_detectable_effect}}
- Target confidence interval level: {{confidence_interval_level}}
- Estimated intra-cluster embedding variance: {{stratified_cluster_variance}}
- Target recipient team: {{stakeholder_group}}
Task
Draft a formal methodological email to {{stakeholder_group}} delivering a statistical power analysis and stratified prompt sampling plan required to benchmark and fine-tune {{target_diffusion_architecture}} with mathematical rigor.
Method
- Calculate the required synthetic image sample size per category in {{prompt_lexicon_categories}} to achieve 80% statistical power at {{confidence_interval_level}}.
- Adjust sample allocations using Neyman optimal stratification based on {{stratified_cluster_variance}} across prompt domains.
- Model the variance inflation factor (VIF) caused by repeated semantic seeds and prompt syntactic collinearity.
- Quantify statistical power curves across varying generation sample sizes against {{minimum_detectable_effect}}.
- Define Monte Carlo sampling stopping rules to prevent redundant compute on high-density prompt clusters.
- Formulate coverage diagnostics using convex hull volume or coverage fraction in multimodal embedding space.
- Present budget and computational trade-offs mapped directly to statistical uncertainty bounds.
Constraints
- MUST explicitly present sample size formulas or derived numerical allocations per prompt stratum.
- MUST NOT recommend uniform sampling when {{stratified_cluster_variance}} indicates non-homogeneous variance.
- Keep formatting structured with markdown tables for sample size breakdowns.
- Limit overall email length to 400 words.
Output format
Email communication formatted as:
- Subject: Statistical Power & Stratified Sampling Design: {{target_diffusion_architecture}}
-
- Executive Allocation Summary (key numbers, total budget, target MDE)
-
- Stratified Sampling Matrix (Table: Category, Cluster Variance, Required Sample Size, Power)
-
- Statistical Methodology & Variance Controls (Neyman allocation, VIF correction)
-
- Stopping Rules & Operational Next Steps
Self-review
- Are all 6 variables integrated with statistically accurate context?
- Is the sample size justified mathematically against {{minimum_detectable_effect}}?
- Does the stratification plan address high intra-cluster variance directly?
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.