Evaluation
AuraScore 81/100

Causal Reasoning and Quantitative Drift Sign-Off

Author an advanced evaluation sign-off email analyzing causal inference stability and numerical drift in quantitative models.

Use this template when evaluating causal inference agents, financial math engines, or quantitative risk models. It produces a detailed evaluation email covering counterfactual validity, numerical drift, and regulatory governance standards.

Template

Role: Chief AI Risk Officer & Quantitative Reasoning Validator

Context

  • System Identifier: {{model_registry_id}}
  • Lead Validator: {{validation_lead_name}}
  • Reasoning Engine: {{causal_inference_framework}}
  • Maximum Variance Allowed: {{drift_tolerance_limit}}
  • Scenario Testbed: {{red_team_scenario_pack}}
  • Governing Entity: {{governance_board}}

Task

Generate a formal validation sign-off email addressed to the {{governance_board}} documenting the quantitative stability, causal fidelity, and numerical precision of {{model_registry_id}} under the stress of {{red_team_scenario_pack}}.

Method

  1. Review baseline structural equation modeling and causal discovery validity.
  2. Evaluate reasoning step integrity using {{causal_inference_framework}}.
  3. Measure numerical calculation precision and floating-point drift under compounding operations.
  4. Stress-test confounding variable isolation within {{red_team_scenario_pack}}.
  5. Compare observed variance metrics directly against {{drift_tolerance_limit}}.
  6. Categorize unfaithful causal assertions versus legitimate model uncertainties.
  7. Evaluate fail-safe mechanisms when encountering indeterminate causal graphs.
  8. Issue the final validation endorsement on behalf of {{validation_lead_name}}.

Constraints

  • MUST provide clear Pass/Conditional Pass/Fail status in the opening paragraph.
  • MUST NOT grant unconditional sign-off if observed drift exceeds {{drift_tolerance_limit}}.
  • Maintain high-level quantitative rigor without omitting statistical proofs.
  • Total output length MUST be between 400 and 650 words.
  • Use bulleted risk matrices for all identified reasoning vulnerabilities.

Output format

Subject: EVALUATION DISPOSITION: {{model_registry_id}} Causal Integrity Audit

  • Addressed to: {{governance_board}}
    1. Executive Disposition & Regulatory Alignment
    1. Quantitative Accuracy & Numerical Drift Analysis
    1. Causal Inference Faithfulness (Evaluated via {{causal_inference_framework}})
    1. Vulnerability Matrix from {{red_team_scenario_pack}}
    1. Sign-off Determination & Operational Conditions
  • Attestation by {{validation_lead_name}}

Self-review

  • Check that mathematical drift is quantitatively contrasted with {{drift_tolerance_limit}}.
  • Confirm the causal framework {{causal_inference_framework}} is correctly characterized.
  • Ensure tone satisfies institutional risk and governance criteria.
AuraScore breakdown
81/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering10/12 · Adequate

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency7/10 · Adequate

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness3/5 · Adequate

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

ai-agents
agents-evaluation
complex-reasoning-analysis-math
causal-inference
quantitative-eval
model-risk