Statistics
AuraScore 77/100

Generative Prompt Sampling Power Analysis and Stratification Synthesis Email

Present statistical power calculations and stratified sampling designs for synthetic multimodal dataset generation.

Use this template when planning prompt generation budgets and synthetic image pipelines for model fine-tuning. It structures a technical email detailing sample size requirements, intra-cluster variance, and confidence intervals.

Template

Role: Lead Synthetic Data Sampling Statistician specializing in high-dimensional multimodal coverage and power optimization.

Context

  • Target architecture for fine-tuning: {{target_diffusion_architecture}}
  • Primary prompt semantic categories: {{prompt_lexicon_categories}}
  • Minimum detectable effect (MDE) for fine-tune benchmark: {{minimum_detectable_effect}}
  • Target confidence interval level: {{confidence_interval_level}}
  • Estimated intra-cluster embedding variance: {{stratified_cluster_variance}}
  • Target recipient team: {{stakeholder_group}}

Task

Draft a formal methodological email to {{stakeholder_group}} delivering a statistical power analysis and stratified prompt sampling plan required to benchmark and fine-tune {{target_diffusion_architecture}} with mathematical rigor.

Method

  1. Calculate the required synthetic image sample size per category in {{prompt_lexicon_categories}} to achieve 80% statistical power at {{confidence_interval_level}}.
  2. Adjust sample allocations using Neyman optimal stratification based on {{stratified_cluster_variance}} across prompt domains.
  3. Model the variance inflation factor (VIF) caused by repeated semantic seeds and prompt syntactic collinearity.
  4. Quantify statistical power curves across varying generation sample sizes against {{minimum_detectable_effect}}.
  5. Define Monte Carlo sampling stopping rules to prevent redundant compute on high-density prompt clusters.
  6. Formulate coverage diagnostics using convex hull volume or coverage fraction in multimodal embedding space.
  7. Present budget and computational trade-offs mapped directly to statistical uncertainty bounds.

Constraints

  • MUST explicitly present sample size formulas or derived numerical allocations per prompt stratum.
  • MUST NOT recommend uniform sampling when {{stratified_cluster_variance}} indicates non-homogeneous variance.
  • Keep formatting structured with markdown tables for sample size breakdowns.
  • Limit overall email length to 400 words.

Output format

Email communication formatted as:

  • Subject: Statistical Power & Stratified Sampling Design: {{target_diffusion_architecture}}
    1. Executive Allocation Summary (key numbers, total budget, target MDE)
    1. Stratified Sampling Matrix (Table: Category, Cluster Variance, Required Sample Size, Power)
    1. Statistical Methodology & Variance Controls (Neyman allocation, VIF correction)
    1. Stopping Rules & Operational Next Steps

Self-review

  • Are all 6 variables integrated with statistically accurate context?
  • Is the sample size justified mathematically against {{minimum_detectable_effect}}?
  • Does the stratification plan address high intra-cluster variance directly?
AuraScore breakdown
77/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering8/12 · Adequate

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness3/5 · Adequate

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

data-analytics
data-statistics
image-multimodal-prompting
statistics
power-analysis
synthetic-data