Data cleaning
AuraScore 79/100

Longitudinal Study Harmonization and Missing Data Imputation Brief

Formulate a systematic data cleaning and multiple imputation strategy for multi-cohort longitudinal research synthesis.

Use this template when synthesizing disparate research cohorts or longitudinal studies with non-random missingness and conflicting variable scales. It generates a rigorous harmonization and imputation protocol.

Template

Role: Principal Research Synthesis Methodologist and Biostatistical Modeler

Context

  • Study Cohorts: {{cohort_study_sources}}
  • Longitudinal Horizons: {{longitudinal_timepoints}}
  • Missingness Hypothesis: {{missingness_mechanism_hypothesis}}
  • Harmonization Criteria: {{covariate_harmonization_rules}}
  • Primary Targets: {{primary_effect_metrics}}
  • Sensitivity Boundaries: {{sensitivity_analysis_targets}}

Task

Develop a comprehensive data harmonization and missingness mitigation brief that establishes clear operational protocols to clean, align, and impute longitudinal observations across {{cohort_study_sources}} without introducing synthetic bias into {{primary_effect_metrics}}.

Method

  1. Map disparate schema definitions across {{cohort_study_sources}} against {{covariate_harmonization_rules}} to identify structural misalignments.
  2. Evaluate longitudinal dropout and panel attrition patterns across {{longitudinal_timepoints}} to test {{missingness_mechanism_hypothesis}} (MCAR, MAR, MNAR).
  3. Define deterministic unit standardizations and semantic cross-walks for non-conforming continuous and categorical covariates.
  4. Design a multiple imputation framework (e.g., MICE, Full Information Maximum Likelihood) tailored to the observed missingness structure.
  5. Establish auxiliary variable injection rules to preserve covariance structures without causing collinearity.
  6. Formulate longitudinal consistency constraints to prevent biologically or temporally impossible imputed values.
  7. Outline the sensitivity analysis framework targeting {{sensitivity_analysis_targets}} to benchmark model stability.

Constraints

  • MUST distinguish clearly between handling strategies for MAR versus suspected MNAR data.
  • MUST NOT allow unconstrained stochastic imputation across deterministic longitudinal bounds.
  • Imputation convergence parameters and pooling rules MUST adhere to Rubin's rules.
  • Deliver the response in an executive research brief format under 750 words.

Output format

Format the deliverable with the following exact headers:

  1. Cross-Cohort Harmonization Specification
  2. Missingness Mechanism & Attrition Assessment
  3. Imputation Protocol & Longitudinal Constraints
  4. Sensitivity Benchmarking & Validation Plan

Self-review

  • Ensure all cohorts in {{cohort_study_sources}} are addressed in the harmonization mapping.
  • Confirm that the imputation mechanism explicitly respects the temporal order in {{longitudinal_timepoints}}.
  • Validate that boundary conditions prevent invalid imputed values.
AuraScore breakdown
79/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering10/12 · Adequate

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness3/5 · Adequate

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

data-analytics
data-cleaning
complex-reasoning-analysis-math
data cleaning
research synthesis
imputation