AI Algorithm Benchmarking and Evidence Synthesis Action Plan
Structure an end-to-end literature review and comparative benchmarking plan for specialized machine learning algorithms.
Use this plan when evaluating research papers and technical whitepapers to select optimization techniques for resource-constrained AI deployments. It delivers a verifiable synthesis matrix and validation framework for research teams.
Role: Staff Applied AI Research Scientist leading model optimization, inference efficiency, and state-of-the-art literature evaluation.
Context
- Algorithmic Focus: {{algorithmic_domain}}
- Target Hardware Platform: {{hardware_target_constraints}}
- Key Performance Indicators: {{primary_evaluation_metrics}}
- Literature Sources: {{repository_indexing_sources}}
- Core Validity Risks: {{threats_to_validity_focus}}
- Target System: {{deployment_environment}}
Task
Design an evidence extraction, benchmarking verification, and literature review project plan to systematically evaluate published research in {{algorithmic_domain}} and identify candidate techniques meeting {{hardware_target_constraints}}.
Method
- Define targeted research questions assessing theoretical convergence, empirical acceleration, and implementation complexity.
- Formulate database search operators, pre-print filtration criteria, and citation-chaining protocols across {{repository_indexing_sources}}.
- Establish reproducible data extraction schemas focusing on dataset lineage, hardware baselines, and reported trade-offs.
- Design a cross-paper metric normalization process to standardize heterogeneous evaluation results against {{primary_evaluation_metrics}}.
- Structure an audit mechanism to interrogate {{threats_to_validity_focus}} including compute-parity verification and test-set contamination.
- Create a synthesis matrix grouping surveyed techniques by memory footprint, inference latency, implementation difficulty, and accuracy trade-offs.
- Detail an empirical replication triage plan to select the top candidate techniques for internal sandbox validation on {{deployment_environment}}.
- Produce a phased resource and milestone execution roadmap covering systematic review through final candidate selection.
Constraints
- MUST define explicit quantitative normalization formulas to compare benchmarks across divergent hardware setups.
- MUST NOT consider pre-prints lacking open-source weights or reproducible artifact repositories for final round shortlisting.
- All evaluation rubrics must explicitly balance raw model capability against {{hardware_target_constraints}}.
- Every stage must include designated artifact checkpoints (e.g., deduplicated bibtex, annotated extraction sheets).
Output format
Provide the review and benchmarking plan with the following five explicit sections:
- Systematic Review Scope & Research Inquiries (formal problem statements and metric boundaries)
- Source Ingestion & Screening Protocol (search strings, citation snowballing steps, and inclusion filters)
- Metric Extraction & Normalization Schema (field-by-field extraction schema and standardization formulas)
- Validity Assessment & Empirical Triaging Strategy (rubric for evaluating paper credibility and contamination risks)
- Execution Roadmap & Candidate Selection Milestone Matrix (structured phases with deliverables and transition gates)
Self-review
- Confirm that the data extraction schema captures all variables needed to compute {{primary_evaluation_metrics}}.
- Check that the validity audit specifically targets {{threats_to_validity_focus}}.
- Ensure the protocol establishes strict criteria for distinguishing peer-reviewed findings from unverified pre-prints.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.