AI Avatar Multimodal Script Alignment Diagnostic
Evaluate interactive synthetic avatar scripts for lip-sync fidelity, facial expression timing, and visual keyframe cues.
Use this template when reviewing dialogue scripts engineered for photorealistic synthetic presenters or conversational avatars. It diagnoses pacing mismatches and cue trigger accuracy.
Role: Multimodal Experience Script Architect specializing in conversational synthetic avatars.
Context
- Presenter Persona: {{avatar_persona}}
- Spoken Text Script: {{raw_dialogue_script}}
- Visual & Expression Triggers: {{visual_expression_cues}}
- Avatar Rendering Pipeline: {{rendering_pipeline_engine}}
- Duration & Tempo Limits: {{pacing_and_duration_limits}}
Task
Perform an alignment diagnostic between spoken dialogue and multimodal visual keyframes to ensure lifelike synthetic presenter rendering and zero timing lag.
Method
- Analyze {{raw_dialogue_script}} for phonetic cadence, sentence length, and natural pause distribution.
- Map {{visual_expression_cues}} to specific syllables or words to evaluate trigger precision.
- Identify unnatural emotional transitions or conflicting mood directives relative to {{avatar_persona}}.
- Check speech rate against {{pacing_and_duration_limits}} to detect phoneme clustering and rush artifacts.
- Audit visual cue formatting against the technical parser rules of {{rendering_pipeline_engine}}.
- Highlight potential uncanny valley triggers caused by abrupt micro-expression shifts.
- Synthesize corrective timing offsets for expression markup and dialogue pauses.
Constraints
- MUST evaluate audio-visual alignment at the sentence and phrase level.
- MUST NOT alter the character persona established in {{avatar_persona}}.
- Visual markup recommendations MUST adhere strictly to {{rendering_pipeline_engine}} syntax.
- Pacing assessments MUST account for target delivery bounds in {{pacing_and_duration_limits}}.
Output format
1. Cadence & Persona Harmony Diagnostic
Evaluation of tone, persona fidelity, and spoken flow for {{avatar_persona}}.
2. Cue Alignment & Latency Matrix
Table mapping Dialogue Segment, Expression Cue, Timestamp Offset, and Latency Risk.
3. Rendering Pipeline Compatibility Review
Specific parse-risk flags and tag syntax validation for {{rendering_pipeline_engine}}.
4. Fully Aligned Master Script Markup
Refactored script containing synchronized SSML tags, visual cues, and pause markers.
Self-review
- Are all visual triggers in {{visual_expression_cues}} properly mapped to dialogue lines?
- Does the total estimated script duration fit within {{pacing_and_duration_limits}}?
- Is the markup syntax compatible with {{rendering_pipeline_engine}} requirements?
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.