Academic Research Data Storage Engine Comparison
Evaluate candidate database architectures for large-scale multi-institutional research data repositories.
Use this template when planning infrastructure for scientific research data platforms requiring complex metadata governance and high-volume ingest. It generates a multi-dimensional matrix comparing storage engines across compliance, performance, and operational constraints.
Role: Principal Research Data Architect specializing in distributed scientific data management.
Context
- Research Organization: {{research_institution}}
- Primary Data Format: {{primary_data_modality}}
- Ingestion Throughput Target: {{ingestion_throughput_target}}
- Governance Standard: {{metadata_governance_standard}}
- Compliance Boundary: {{compliance_boundary}}
- Budget Tier: {{budget_tier}}
Task
Generate a structured decision matrix comparing at least three database storage engines to identify the optimal platform for scientific data sharing and long-term research archival.
Method
- Review the ingest requirements defined by {{ingestion_throughput_target}} against standard database write patterns.
- Analyze how {{primary_data_modality}} structures dictate schema flexibility, indexing strategies, and relational versus non-relational storage needs.
- Identify three distinct database engine architectures appropriate for {{research_institution}} and its resource tier {{budget_tier}}.
- Map each candidate engine against compliance requirements enforced by {{compliance_boundary}}.
- Evaluate support for metadata cataloging and FAIR data principles aligned with {{metadata_governance_standard}}.
- Score each candidate across query performance, schema evolution, access controls, and maintenance overhead.
- Synthesize findings into a final comparative matrix accompanied by an architectural recommendation summary.
Constraints
- MUST evaluate at least three distinct database engine paradigms (e.g., distributed SQL, document/graph, columnar object storage).
- MUST explicitly score each engine against {{compliance_boundary}} and {{metadata_governance_standard}}.
- Do not include vendor sales assertions without explicit technical justification.
- Scoring MUST use a consistent 1-5 scale with defined scoring criteria.
Output format
- Executive Context (2-3 sentences summarizing the deployment target).
- Candidate Engine Overview (bulleted list defining the 3 engines evaluated).
- Comparative Evaluation Matrix (Markdown table with columns: Candidate Engine, Schema Fit, {{compliance_boundary}} Compliance, Ingestion Capacity, Governance Fit, Total Score [out of 25]).
- Architectural Recommendation (1 paragraph justifying the top-scoring engine).
Self-review
- Verify all 6 context variables are directly addressed in the evaluation criteria.
- Confirm table columns strictly match the specified layout.
- Check that scoring totals mathematically sum up correctly.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.