High-Frequency Trading Runbook Standardization Plan
Plan the overhaul and verification of mission-critical incident response and failover runbooks for electronic trading platforms.
Use this template to plan the restructuring of mission-critical trading infrastructure runbooks. It sets up actionable schedules for documenting disaster recovery workflows, technical escalation paths, and exchange reconnection procedures.
Role: Senior Trading Infrastructure Technical Writer specializing in low-latency systems and operational resilience.
Context
- Brokerage / Trading Firm: {{brokerage_name}}
- Trading Infrastructure Scope: {{trading_system_scope}}
- Incident Severity Tiers: {{incident_severity_tiers}}
- Governance & Regulatory Mandate: {{primary_audit_mandate}}
- Documentation Repository & Toolchain: {{documentation_toolchain}}
- Delivery Window: {{execution_window}}
Task
Construct a comprehensive project plan to rewrite, test, and validate mission-critical incident response and failover runbooks for {{brokerage_name}}'s {{trading_system_scope}}, ensuring alignment with {{primary_audit_mandate}} within {{execution_window}}.
Method
- Inventory existing emergency procedure documentation across {{trading_system_scope}}, identifying outdated diagrams, broken command scripts, and missing network failover paths.
- Standardize a uniform runbook template tailored for high-pressure incident mitigation (symptoms, immediate triage commands, verification commands, and rollback actions).
- Map procedural workflows to specific severity definitions outlined in {{incident_severity_tiers}}.
- Coordinate technical interview sprints with Site Reliability Engineers, quantitative infrastructure leads, and network engineers to capture failover sequences.
- Implement a strict validation mechanism within {{documentation_toolchain}} to test executable code blocks, CLI snippets, and monitoring query links.
- Schedule table-top disaster recovery drills to evaluate the accuracy and execution speed of the newly drafted runbooks.
- Prepare final regulatory artifacts demonstrating procedural resilience for {{primary_audit_mandate}} review.
Constraints
- MUST format all diagnostic and remediation steps into sequential, deterministic command-line actions.
- MUST NOT use ambiguous subjective language (e.g., 'wait a while' or 'check if healthy') without exact timeout values and metric thresholds.
- All runbooks must include rollback procedures for failed manual interventions.
- The execution roadmap must accommodate engineer shift schedules across trading desk operational hours.
Output format
Provide a technical runbook overhaul plan formatted under the following headings:
- Infrastructure Scope & Runbook Inventory (summary table)
- Standardized Runbook Structural Specification (anatomy of a production runbook)
- Authoring, Verification & Table-Top Drill Schedule (weekly progression across {{execution_window}})
- Tooling, Automated Linting & Version Control Plan (integration with {{documentation_toolchain}})
- Audit Compliance & Operational Sign-off Protocol (alignment with {{primary_audit_mandate}})
Self-review
- Does the plan account for all system components listed in {{trading_system_scope}}?
- Are command verification steps integrated into the authoring workflow to prevent stale instructions?
- Is the validation methodology capable of satisfying {{primary_audit_mandate}} without disrupting trading hours?
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.