Scholarly Repository and Research Discoverability Indexing Checklist
Optimize open-access academic repositories for Google Scholar, PubMed, and commercial crawlers.
Use this checklist when deploying or auditing institutional digital repositories, preprint servers, or university press publications. It guarantees proper metadata harvestability via OAI-PMH and citation indexing across academic search engines.
Role: Scholarly Communications and Academic Search Indexing Specialist with expertise in institutional repository discoverability, Highwire Press tags, and OAI-PMH.
Context
- Repository Name: {{repository_name}}
- Repository Software: {{repository_software}}
- DOI Registration Agency: {{doi_prefix_registry}}
- Focus Research Disciplines: {{primary_research_disciplines}}
- Repository Harvesting Endpoint: {{oai_pmh_endpoint}}
- Academic Discovery Engines: {{target_academic_engines}}
Task
Create a comprehensive technical discoverability checklist to ensure academic papers, dissertations, and technical reports hosted on {{repository_name}} achieve rapid, full-text indexation and correct citation tracking across {{target_academic_engines}}.
Method
- Inspect HTML head meta-tag implementations against Highwire Press, Dublin Core, and PRISM metadata standards supported by {{repository_software}}.
- Verify automated XML sitemap generation protocols for PDF and full-text HTML galleys, confirming compliance with crawl thresholds of {{target_academic_engines}}.
- Validate OAI-PMH metadata feed integrity at {{oai_pmh_endpoint}} across Dublin Core (oai_dc) and discipline-specific schemas for {{primary_research_disciplines}}.
- Audit DOI landing page redirection, handle resolver stability, and persistent identifier linkage governed by {{doi_prefix_registry}}.
- Assess robots.txt and HTTP header directives to ensure bot access to open-access PDFs while barring scraper-induced server exhaustion.
- Verify Schema.org ScholarlyArticle JSON-LD structures, validating author affiliations, license declarations, and funding agency grants.
- Establish crawl failure diagnostic routines for papers missing from {{target_academic_engines}} citation indexes.
Constraints
- MUST require citation_title, citation_author, citation_publication_date, and citation_pdf_url meta tags on every paper view page.
- MUST verify that the {{oai_pmh_endpoint}} responds with valid XML conforming to OAI 2.0 specs without HTTP timeouts.
- MUST NOT allow crawler-blocking X-Robots-Tag headers on full-text asset downloads.
- Deliverables must separate automated server checks from human-mediated metadata audits.
Output format
- Section 1: Machine-Readable Meta Tag Architecture (7-9 verification items with syntax examples)
- Section 2: PDF & Asset Crawlability Protocols (5-7 verification items)
- Section 3: OAI-PMH & Persistent Identifier Audit (5-6 verification items covering {{doi_prefix_registry}})
- Section 4: Schema.org & Linked Open Data Validation (5-7 verification items)
- Section 5: Engine-Specific Verification for {{target_academic_engines}} (4-6 verification items)
Self-review
- Confirm all 6 context variables are deeply integrated into the checklist specifications.
- Verify that academic metadata tags (citation_) are differentiated from commercial SEO tags (og:).
- Ensure instructions include specific HTTP status codes and mime-type headers for academic PDF delivery.
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.