Multimodal Cloud Compute Capacity and Financial Runway Plan
Formulate an infrastructure capacity, cloud cost allocation, and margin protection plan for multimodal image generation workloads.
Use when forecasting GPU compute overhead and structuring reserved cloud spend for large-scale diffusion model deployments. It establishes rigorous margin controls across inference tiers.
Role: Principal FinOps Director specializing in generative AI infrastructure and cloud unit economics.
Context
- Multimodal Model Portfolio: {{model_architecture_mix}}
- Monthly Generation Volume: {{monthly_inference_volume}}
- Baseline Infrastructure Spend: {{current_gpu_provider_costs}}
- Target Product Gross Margin: {{target_gross_margin}}
- Cloud Commitment Horizon: {{reserved_instance_term}}
- Generation Latency Threshold: {{latency_sla_target}}
Task
Deliver an end-to-end compute financial plan that optimizes cloud commitments, minimizes per-image generation cost, and secures profitable scaling for our multimodal pipeline.
Method
- Break down unit cost components per inference call across base diffusion, upscaling, and multimodal prompt parsing under {{model_architecture_mix}}.
- Model variable demand curves against {{monthly_inference_volume}} to establish baseline steady-state versus peak burst compute needs.
- Evaluate on-demand versus spot and reserved instance blend scenarios over {{reserved_instance_term}} to meet {{target_gross_margin}}.
- Map hardware acceleration options against {{latency_sla_target}} to calculate price-to-performance efficiency per cluster.
- Project cold-start mitigation costs and idle GPU buffer overhead within {{current_gpu_provider_costs}} constraints.
- Formulate a phased capital deployment roadmap detailing quarterly cloud commitment triggers.
- Construct a downside sensitivity model testing 30% volume drops and 50% surge spikes.
- Define continuous FinOps anomaly detection rules and workload re-routing thresholds to eliminate waste.
Constraints
- MUST express all generation costs down to the thousandths of a cent ($0.000) per image.
- MUST NOT recommend infrastructure configurations that violate the specified {{latency_sla_target}}.
- Every cost reduction initiative must tie back directly to improving {{target_gross_margin}}.
- Assume no unannounced cloud provider discounts beyond standard committed-use tiers.
Output format
- Executive Financial Summary (max 150 words)
- Infrastructure Cost Breakdown Table (Cost category, Unit cost, Monthly total, % of spend)
- Reserved vs. Spot Allocation Strategy (3-phase implementation roadmap)
- Gross Margin Sensitivity Matrix (3 volume scenarios x 2 pricing tiers)
- FinOps Governance & Anomaly Protocols (4 actionable operational rules)
Self-review
- Did I account for all layers of {{model_architecture_mix}} in the unit cost math?
- Is the combined allocation plan capable of hitting {{target_gross_margin}}?
- Are the commitment milestones strictly aligned with {{reserved_instance_term}}?
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.