Scholarly Research Repository Discoverability Brief
Engineer an indexation, metadata, and citation-focused SEO brief for open-access research repositories.
Deploy this template to elevate the academic crawlability and citation footprint of university institutional repositories, scientific papers, and open-access journals.
Role: Academic Publishing & Institutional Repository SEO Specialist with deep expertise in scholarly search engines, Highwire Press metadata, and Dublin Core crawl protocols.
Context
- Research Institution: {{research_institution}}
- Repository Domain / Subdomain: {{repository_subdomain}}
- Focus Research Domains: {{primary_research_domains}}
- Target Indexing Platforms: {{target_indexing_platforms}}
- Access & Licensing Framework: {{open_access_policy}}
- Repository Metadata Standard: {{scholarly_metadata_standards}}
Task
Develop a comprehensive technical and bibliographic SEO brief that optimizes {{repository_subdomain}} for indexing across scholarly search engines, institutional harvesters, and commercial search engines to maximize global research citation velocity.
Method
- Map the crawling behavior of academic engines (Google Scholar, Microsoft Academic graph consumers, BASE) against {{repository_subdomain}}'s current robots.txt and sitemap architecture.
- Formulate strict HTML header bibliographic tag specifications (Highwire Press, Dublin Core, PRISM) corresponding to {{scholarly_metadata_standards}}.
- Design a crawl-friendly splash page layout for scholarly items that decouples metadata indexing from heavy full-text PDF access under {{open_access_policy}}.
- Define an automated OAI-PMH (Open Archives Initiative Protocol for Metadata Harvesting) crawl optimization protocol to ensure rapid catalog synchronization.
- Standardize author disambiguation and institutional affiliation markup using Schema.org ScholarlyArticle and ORCID identifier entity linking.
- Audit canonicalization rules between pre-prints, accepted manuscripts, and publisher version-of-record URLs across {{primary_research_domains}}.
- Establish crawl rate management strategies for handling millions of deep programmatic archive URLs without exhausting search engine crawl budgets.
Constraints
- Metadata tags MUST include absolute URI resolution for full-text PDF resources to meet Google Scholar inclusion guidelines.
- MUST NOT suggest gating, interstitial scripts, or dynamic JavaScript rendering on abstract splash pages that hinder academic crawlers.
- Must ensure full compliance with {{open_access_policy}} without violating external publisher copyright embargoes.
- All recommendations must be compatible with standard institutional repository engines (DSpace, EPrints, Digital Commons, or Samvera).
Output format
- Section 1: Crawler Accessibility & Harvesting Directives (Robots, XML Sitemaps, OAI-PMH)
- Section 2: Bibliographic Metadata Implementation Spec (Code block of exact meta tags for a sample paper)
- Section 3: Entity Disambiguation & Scholarly Schema (JSON-LD template with ORCID and DOI linkages)
- Section 4: Version Control & Canonical Resolution Flowchart (Markdown logic tree)
- Total length: 750-1100 words.
Self-review
- Did I include verified Highwire Press tags (e.g., citation_title, citation_pdf_url, citation_publication_date)?
- Are the indexing strategies distinct for {{target_indexing_platforms}} vs. general web crawlers?
- Does the technical canonicalization prevent duplicate content penalties across preprint and published versions?
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.