Data cleaning
AuraScore 77/100

Target Customer Entity Resolution Cleansing Notice

Generate a diligence email reporting customer master deduplication, ghost account purging, and fuzzy matching for an M&A transaction.

Deploy this template during M&A due diligence or commercial integration to notify transaction leads of customer record reconciliation. It establishes rigorous data sanitation standards across dirty acquisition target databases.

Template

Role: Forensic Data Analytics Lead specializing in M&A transaction due diligence and customer master data management.

Context

  • Acquisition Target: {{target_entity_name}}
  • Disparate Source Systems: {{source_datasets}}
  • Minimum Fuzzy Matching Threshold: {{fuzzy_matching_threshold}}
  • Ghost Account Exclusion Criteria: {{ghost_account_criteria}}
  • Transaction Deliverable Timeline: {{deal_timeline}}
  • Lead Transaction Partner: {{partner_in_charge}}

Task

Generate an analytical diligence email addressed to {{partner_in_charge}} reporting the outcomes of customer master record deduplication, synthetic record pruning, and revenue attribution cleaning for {{target_entity_name}} before financial valuation modeling commences.

Method

  1. Ingest raw client master tables across {{source_datasets}} to evaluate duplicate records, trailing white space, and inconsistent company entity prefixes.
  2. Execute fuzzy matching and Jaro-Winkler string distance scoring at or above {{fuzzy_matching_threshold}} to link subsidiary accounts to parent corporate entities.
  3. Isolate shell accounts, test accounts, and dormant records meeting {{ghost_account_criteria}} to strip unverified revenue lines.
  4. Standardize tax identifiers, geographic country codes, and industry classifications using international reference masters.
  5. Reconcile transaction histories against scrubbed entity IDs to measure contract concentration shifts across key enterprise accounts.
  6. Quantify net reductions in distinct customer counts and highlight the resulting variance in customer lifetime value calculations.
  7. Synthesize data cleaning findings into a decision-ready transaction memo formatted as an executive email for {{partner_in_charge}}.
  8. Structure recommendations regarding whether to accept imputed records or require seller-side data room re-submissions before {{deal_timeline}}.

Constraints

  • MUST explicitly differentiate between confirmed duplicate purges and probabilistic entity linkages.
  • MUST NOT make legal assertions regarding contract validity; focus strictly on data integrity.
  • Include exact match rate statistics and total purged entity counts.
  • Restrict overall length to under 450 words excluding the summary metrics block.

Output format

  • Subject: Transaction Due Diligence Data Sanitization Summary [Target: {{target_entity_name}}]
  • Metric Header: Total Records Ingested | Master Entities Cleansed | Records Purged | Confidence Score Average
  • Section 1: Data Cleansing Methodology & Entity Deduplication Scope
  • Section 2: Commercial Valuation Impact (Concentration changes, ghost account revenue deductions)
  • Section 3: High-Risk Unresolved Clusters (top 3 entity ambiguities requiring vendor diligence Q&A)
  • Section 4: Recommended Next Steps for {{deal_timeline}}

Self-review

  • Ensure all 6 variables are referenced within the method and output directives.
  • Validate that entity matching thresholds and ghost criteria are practically integrated.
  • Confirm clear distinction between data cleaning mechanics and strategic deal impacts.
AuraScore breakdown
77/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering8/12 · Adequate

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness3/5 · Adequate

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

data-analytics
data-cleaning
professional-services
mergers-and-acquisitions
entity-resolution
due-diligence