Enterprise Image Model Fidelity and Prompt Adherence Audit
Compile an exhaustive technical benchmark report auditing prompt adherence and visual artifact rates across image models.
Use this template to assess image generation models before rolling them into automated creative pipelines. It delivers a structured comparative evaluation report identifying prompt drift, coherence flaws, and stylistic adherence.
Role: Staff Multimodal AI Researcher and Quality Benchmarking Lead
Context
- Enterprise Client: {{enterprise_client}}
- Evaluated Models: {{evaluated_models}}
- Core Subject Domains: {{core_subject_domains}}
- Prompt Complexity Tiers: {{prompt_complexity_tiers}}
- Fidelity Evaluation Criteria: {{fidelity_criteria}}
- Artifact Risk Threshold: {{artifact_risk_threshold}}
Task
Generate an in-depth benchmark audit report evaluating prompt following accuracy, spatial relation coherence, and artifact prevalence across {{evaluated_models}} for {{enterprise_client}} within {{core_subject_domains}}.
Method
- Define prompt test suites stratified across {{prompt_complexity_tiers}} (simple, multi-subject, relational, and contextual).
- Establish scoring rubrics for text alignment based on {{fidelity_criteria}}, measuring text-to-image semantic translation.
- Measure multi-subject spatial positioning accuracy, counting attribute bleeding and relational placement failures.
- Audit anatomical and geometric integrity against {{artifact_risk_threshold}} across repeated seed batches.
- Evaluate text rendering capabilities and typographic coherence across models when explicit lettering is requested.
- Compare prompt latency, token degradation thresholds, and cross-attention weight behavior across {{evaluated_models}}.
- Synthesize findings into actionable deployment recommendations for commercial asset pipelines.
Constraints
- MUST provide quantitative benchmark scoring rubrics (scale of 1-10) with explicit qualitative criteria for each level.
- MUST evaluate specific failure modes including color contamination, limb distortion, and background hallucination.
- MUST NOT recommend models that exceed the specified {{artifact_risk_threshold}} without mitigation steps.
- Analysis MUST compare model performance directly across identical prompt strings in {{core_subject_domains}}.
Output format
Deliver the audit as a formal technical evaluation report with the following structure:
Executive Summary & Risk Assessment
Benchmark Methodology & Prompt Stratification
Comparative Model Scorecards (Detailed Analysis of {{evaluated_models}})
Subject Domain Analysis (Coverage of {{core_subject_domains}})
Artifact Prevalence & Failure Mode Matrix
Production Deployment Roadmap
Total report length should span 1,400 to 2,000 words.
Self-review
- Verify that comparative scoring tables assess all models listed in {{evaluated_models}} equitably.
- Ensure failure modes are explicitly detailed with concrete prompt examples demonstrating breakdown.
- Check that deployment recommendations address the constraints defined in {{artifact_risk_threshold}}.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.