Legal Discovery Hybrid Search Database Selection Matrix
Compare database backends for legal discovery, semantic retrieval, and policy archive indexing.
Use this template when selecting database technologies for legal discovery, regulatory record archives, or compliance document intelligence. It creates a technical trade-off matrix balancing vector recall, relational metadata filtering, and strict audit logging.
Role: Senior Legal Engineering Database Specialist focused on regulatory compliance and search infrastructure.
Context
- Practice Area: {{legal_practice_area}}
- Corpus Size: {{total_document_volume}}
- Query Latency SLA: {{retrieval_latency_sla}}
- Residency Requirements: {{jurisdictional_residency_rules}}
- Search Workload Mix: {{search_modality_mix}}
- Base Infrastructure: {{existing_infrastructure}}
Task
Produce a technical evaluation matrix contrasting candidate database configurations for a legal discovery platform to determine the optimal engine for mixed vector, full-text, and metadata queries.
Method
- Deconstruct the workload mix from {{search_modality_mix}} into exact keyword, faceted metadata, and vector similarity retrieval requirements.
- Assess the scalability requirements for storing and indexing {{total_document_volume}} within {{existing_infrastructure}}.
- Frame database candidates across three categories: dedicated vector database, integrated relational vector store, and enterprise search cluster.
- Evaluate index rebuild times and query execution times against {{retrieval_latency_sla}}.
- Score candidate databases on strict adherence to {{jurisdictional_residency_rules}} including field-level encryption and immutable audit trails.
- Analyze operational overhead specific to litigation holds and rapid evidence ingestion in {{legal_practice_area}}.
- Compile findings into a comparison matrix with a weighted trade-off analysis.
Constraints
- MUST compare hybrid search capabilities (combining BM25/keyword with dense embeddings).
- MUST evaluate immutability, audit logging, and {{jurisdictional_residency_rules}} for each candidate.
- Do not recommend proprietary hosted solutions that violate {{existing_infrastructure}} deployment constraints.
- Candidate solutions MUST be technically viable for {{total_document_volume}}.
Output format
- Scope Statement (1 brief paragraph defining workload and scale parameters).
- Technical Evaluation Matrix (Markdown table with columns: Candidate Engine, Hybrid Search Quality, Indexing Speed, {{jurisdictional_residency_rules}} Compliance, Operational Complexity, Fit Score [1-10]).
- Trade-off Analysis (3 bullet points highlighting primary trade-offs: latency vs. cost, schema flexibility vs. strict auditability).
- Definitive Architecture Selection (1 short paragraph with final recommendation).
Self-review
- Ensure the matrix explicitly accounts for {{search_modality_mix}} requirements.
- Confirm that {{retrieval_latency_sla}} is specifically evaluated in the engine review.
- Check that every candidate supports required residency and compliance boundaries.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.