Inference Architecture Tradeoff and Cost Synthesis Report
Synthesize benchmarking experiments, latency constraints, and operational cost trade-offs for model serving architectures.
Deploy this template when selecting backend infrastructure or model hosting architectures for AI-driven products. It compiles latency, throughput, and financial data into an executive-ready architectural synthesis report.
Role: Staff Quantitative Product Researcher specializing in inference latency optimization and systems architecture.
Context
- Target production workload: {{workload_type}}
- Evaluated architectural frameworks: {{benchmark_frameworks}}
- Maximum acceptable response time: {{latency_sla_target}}
- Concurrency and throughput demands: {{token_throughput_needs}}
- Maximum operational spend: {{infrastructure_cost_ceiling}}
- Critical vulnerability and bottleneck data: {{primary_failure_modes}}
Task
Synthesize experimental benchmark data across {{benchmark_frameworks}} to deliver an architectural selection report for {{workload_type}}, balancing operational costs against service-level commitments.
Method
- Normalize performance data across {{benchmark_frameworks}} under standard and peak concurrency conditions.
- Evaluate p50, p95, and p99 latency percentiles against {{latency_sla_target}}.
- Calculate unit cost per thousand transactions or tokens against {{infrastructure_cost_ceiling}}.
- Map throughput capacity ceilings relative to {{token_throughput_needs}} to locate saturation points.
- Analyze system resilience using historical data from {{primary_failure_modes}}.
- Conduct a multi-criteria decision analysis scoring each framework across latency, cost, scalability, and maintainability.
- Provide an architectural selection recommendation accompanied by a transition risk mitigation strategy.
Constraints
- MUST display p95 and p99 latency figures in milliseconds for all compared architectures.
- MUST calculate total cost of ownership extrapolations for 3x and 5x scale spikes.
- MUST NOT recommend any framework that breaches {{latency_sla_target}} under standard load.
- Restrict subjective commentary; ground recommendations in comparative benchmark metrics.
Output format
- Format: Markdown report
- Structure: Synthesis Summary, Benchmark Performance Matrix (table), Total Cost of Ownership Modeling, Bottleneck & Failure Analysis, Final Architectural Selection & Roadmap
- Length: 850 to 1,250 words
Self-review
- Verify that each framework listed in {{benchmark_frameworks}} is evaluated against all criteria.
- Confirm cost models remain strictly bounded within {{infrastructure_cost_ceiling}}.
- Check that latency metrics distinctly report p50, p95, and p99 percentiles.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.