Statistics
AuraScore 83/100

Credit Risk Model Backtesting and Calibration Brief

Evaluate probability of default drift, calibration fit, and stress test reliability across institutional credit portfolios.

Use this template when conducting annual model validation or regulatory backtesting on internal ratings-based (IRB) credit portfolios. It translates statistical test discrepancies into actionable recalibration strategies for risk committees.

Template

Role: Principal Quantitative Risk Modeler specializing in Basel IV IRB credit risk validation.

Context

  • Financial Institution: {{institution_name}}
  • Portfolio Segment: {{portfolio_segment}}
  • Historical Loss Data: {{historical_loss_dataset}}
  • Significance Threshold: {{significance_threshold}}
  • Macroeconomic Phase: {{macroeconomic_cycle_context}}
  • Benchmark Specification: {{benchmark_model_spec}}

Task

Synthesize empirical validation metrics into a rigorous credit risk model backtesting brief that quantifies predictive drift, validates discriminatory power, and delivers concrete recalibration recommendations for supervisory review.

Method

  1. Calculate discriminatory power metrics including the Area Under the Receiver Operating Characteristic (AUROC) and Gini coefficient against {{historical_loss_dataset}}.
  2. Evaluate calibration accuracy using the Hosmer-Lemeshow goodness-of-fit test and binomial test across all rating notches at {{significance_threshold}}.
  3. Quantify population stability and characteristic drift by computing the Population Stability Index (PSI) and System Stability Index relative to {{benchmark_model_spec}}.
  4. Perform transition matrix stability analysis to identify rating migration velocity under {{macroeconomic_cycle_context}}.
  5. Execute Brier score decomposition to disentangle uncertainty, reliability, and resolution components of the current Probability of Default (PD) estimates.
  6. Compare empirical default realization curves against theoretical cumulative default rates across {{portfolio_segment}}.
  7. Formulate statistical override rules and parameter adjustments to reconcile observed under-prediction or over-prediction pockets.

Constraints

  • MUST express all statistical divergence tests with explicit p-values and confidence intervals.
  • MUST NOT recommend qualitative overrides without empirical support from the Brier decomposition or PSI scores.
  • Assumptions regarding loss truncations or right-censoring MUST be mathematically justified.
  • Total brief length MUST NOT exceed 1,000 words.

Output format

Provide a technical brief structured as follows:

  • Executive Summary (150 words max)
  • Discriminatory Power & Separation Analysis (tabular breakdown of AUROC, Gini, and Kolmogorov-Smirnov)
  • Calibration Fit & Goodness-of-Fit Diagnostics (test statistics, p-values, and notch-level findings)
  • Macroeconomic Sensitivity & Stability Indices (PSI scores and transition matrix shifts)
  • Parameter Recalibration Strategy (bulleted tactical steps)

Self-review

  1. Are all statistical metrics computed against the explicit {{significance_threshold}}?
  2. Does the recalibration strategy directly resolve the specific rating notch failures identified?
  3. Are portfolio segment nuances in {{portfolio_segment}} appropriately addressed without generic modeling platitudes?
AuraScore breakdown
83/100Provisional
Instruction clarity15/15 · Strong

Explicit role, a named task, and discrete steps the model can follow.

Context architecture12/12 · Strong

Background, inputs and variables the model needs before it starts.

Constraint engineering12/12 · Strong

Hard boundaries — what the model must and must not do.

Output specification6/14 · Thin

A named, field-level shape for the response.

Reasoning structure10/10 · Strong

Ordered work items that force analysis before an answer.

Model compatibility10/10 · Strong

Length and structure that travel across frontier models.

Token efficiency5/10 · Thin

Signal density — instruction weight without padding.

Reusability7/7 · Strong

Documented variables so the scaffold adapts to new inputs.

Robustness5/5 · Strong

Quality bar, assumptions and behaviour when inputs are thin.

Observed performance1/5 · Thin

How much real usage the template has behind it.

data-analytics
data-statistics
financial-services
credit-risk
backtesting
basel