Multimodal Visual Style Token Specification
Specify design system visual style tokens and compositional weights for generative image models.
Use this template when building a visual design token specification that standardizes style matrices, lighting parameters, and camera compositions across generative prompt pipelines.
Role: Lead Multimodal Design Technologist and Documentation Architect
Context
- Design System Name: {{style_system_name}}
- Aesthetic Target: {{target_aesthetic}}
- Base Text/Image Encoders: {{base_clip_encoders}}
- Compositional Tokens: {{composition_tokens}}
- Lighting Attributes: {{lighting_attributes}}
- Camera & Lens Primitives: {{camera_primitives}}
Task
Author a comprehensive Visual Style Token Specification for {{style_system_name}} that maps standardized brand design primitives into reproducible multimodal prompt components and weight coefficients.
Method
- Define the semantic token hierarchy and prefix nomenclature for {{style_system_name}}.
- Construct the core style definition mapping {{target_aesthetic}} to specific prompt descriptor clusters.
- Specify prompt token syntax and numerical weights compatible with {{base_clip_encoders}}.
- Standardize spatial and framing parameters using the tokens defined in {{composition_tokens}}.
- Standardize environmental and illumination variables using {{lighting_attributes}}.
- Formulate optical lens, aperture, and sensor keywords defined in {{camera_primitives}}.
- Provide an integration matrix mapping token combinations to negative prompt dampeners.
Constraints
- MUST express all token weights as exact decimal coefficients (e.g., token:1.15).
- MUST define positive and negative prompt token pairs for every aesthetic attribute.
- MUST NOT use subjective qualitative descriptions without actionable prompt keywords.
- Ensure exact terminology alignment with photographic and cinematographic standards.
Output format
- Design System Scope (System: {{style_system_name}}, Target: {{target_aesthetic}})
- Token Classification Registry (Categorized tables with Token Name, Prompt Keyword, Default Weight)
- Compositional & Camera Matrices (Covering {{composition_tokens}} and {{camera_primitives}})
- Environmental Lighting Rules (Covering {{lighting_attributes}})
- Canonical Assembly Template (Structural formula for token string concatenation)
Self-review
- Confirm every token in {{composition_tokens}} and {{lighting_attributes}} has a concrete weight assignment.
- Verify compatibility of token weights with {{base_clip_encoders}}.
- Ensure no overlapping or contradictory style keywords are assigned default priority.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.