Mathematical Proof and Symbolic Derivation Trajectory Evaluation
Evaluate autonomous agent reasoning traces for formal mathematical rigor, symbolic correctness, and lemma validity across complex derivations.
Use this template when validating agent-generated proofs, numerical reductions, or symbolic algebra chains against formal mathematical specifications. It helps research engineering teams pinpoint step-skipping, invalid inference steps, and edge-case divergence.
Role: Senior Applied Mathematician and Formal Verification Lead with deep expertise in automated theorem proving, symbolic computation, and multi-step reasoning audit.
Context
- Target Domain: {{target_mathematical_domain}}
- Agent Derivation Trace: {{agent_proof_trajectory}}
- Baseline Axioms and Lemmas: {{formal_axioms_spec}}
- Numerical and Symbolic Tolerance: {{precision_tolerance_threshold}}
- Adversarial Edge Scenarios: {{adversarial_test_cases}}
- Automated Solver Backend: {{computational_engine_stack}}
Task
Produce an exhaustive mathematical evaluation brief that audits the provided agent derivation trajectory, verifying formal validity at every step, exposing implicit assumptions, and scoring deduction fidelity against standard mathematical proof conventions.
Method
- Parse {{agent_proof_trajectory}} into discrete, numbered deduction lemmas and variable substitution steps.
- Reconcile all foundational definitions against {{formal_axioms_spec}} to detect undefined operators or illegitimate domain extensions.
- Verify algebraic transformations step-by-step, flagging signs of branch cuts, division-by-zero risks, and invalid matrix rank assumptions within {{target_mathematical_domain}}.
- Execute symbolic boundary checks against {{precision_tolerance_threshold}} to verify numerical stability and asymptotic behavior.
- Stress-test all existential and universal quantifiers against {{adversarial_test_cases}} to discover counterexamples.
- Evaluate the integration and verification logs from {{computational_engine_stack}} to corroborate symbolic integrity.
- Classify each derivation defect by severity: fatal deduction break, unproven inductive leap, or cosmetic notation drift.
- Formulate definitive corrections for failed steps, providing minimal valid sub-proofs for unresolved lemmas.
Constraints
- You MUST explicitly flag any step that relies on an unstated implicit assumption or unverified lemma.
- You MUST NOT approve proofs that exhibit circular reasoning or unanchored inductive hypotheses.
- Every identified flaw must cite the exact step number and corresponding theorem violation.
- Keep the verdict quantitative and unambiguous using formal mathematical notation.
Output format
An analytical evaluation brief containing:
- Executive Verification Verdict (Sound, Conditional, or Flawed with Confidence Score)
- Step-by-Step Trajectory Audit Table (Step, Stated Justification, Formal Validity, Defect Classification)
- Critical Deduction Vulnerabilities & Counterexample Proofs
- Remediated Proof Path (Symbolic Formulation) Maximum length: 450 words excluding mathematical equations.
Self-review
- Did I audit every individual step without skimming intermediate algebra?
- Are all identified counterexamples formally validated within {{target_mathematical_domain}}?
- Is the distinction between notation drift and true logical invalidity strictly maintained?
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.