Cluster Bootstrap
Purpose. Replace and supplement the original analytic Hanley-McNeil AUROC interval with an institution-level (UNITID) bootstrap, so that test rows of the same institution stay together when the AUROC distribution is resampled.
Method
- Restricted to the locked validation cell: FASB × Sector 2, closure-within-five-years label.
- Test cohort: 3,232 rows, 1,644 unique institutions, 117 closure-positive rows from 71 distinct closure-positive institutions.
- Resampling unit: institution (UNITID), with replacement, sample size equal to the number of unique institutions.
- 1,000 bootstrap samples; deterministic RNG seed.
- Predictors compared: joint-weak G+A; G-only; A-only; divergence; the strongest tested single-variable baselines (negated net-position cushion, negated cash operating margin, negated six-year graduation rate).
Key findings
- Joint-weak AUROC point estimate ≈ 0.864. Cluster-bootstrap 95 % percentile interval ≈ [0.806, 0.914]. The original analytic Hanley-McNeil interval was [0.801, 0.928]; the cluster interval is essentially identical at the lower bound and slightly tighter at the upper bound.
- Per-predictor cluster CIs (median, then 2.5 %–97.5 %): joint-weak 0.864 [0.806, 0.914]; G-only 0.798 [0.738, 0.848]; A-only 0.728 [0.646, 0.800]; divergence 0.680 [0.590, 0.759]; negated net-position cushion 0.792 [0.732, 0.846]; negated six-year graduation rate 0.760 [0.699, 0.812]; negated cash operating margin 0.632 [0.554, 0.702].
- Paired-difference distribution, joint-weak minus negated net-position cushion (the strongest baseline): mean ≈ +0.072; 95 % cluster-bootstrap CI ≈ [+0.002, +0.138]; about 98 % of resamples positive.
- Paired-difference vs negated low six-year graduation rate (the simplest A baseline): mean ≈ +0.105; 95 % CI ≈ [+0.022, +0.181]; about 99 % of resamples positive.
How to read this
Higher point estimate, not definitive superiority. The cluster-bootstrap paired-difference 95 % CI vs the strongest baseline has its lower bound at +0.002 — narrowly above zero. The defensible reading is “higher AUROC point estimate over tested single-variable baselines,” not “statistically superior to the baselines.”
Within-sample uncertainty only. Cluster bootstrap addresses the right uncertainty model for repeated-institution panel rows. It does not address other forms of model misspecification, distribution shift across years, or stratum mis-classification. Multi-fold temporal stability is the separate Temporal Folds track.
Limitations
- The resample preserves the locked train/test split; the model is not refit on bootstrapped train data. This is the standard test-metric bootstrap holding the trained model fixed.
- The number of closure-positive UNITIDs is 71; cluster-bootstrap uncertainty is fundamentally constrained by that count.
- The result is for the validated lane only and does not generalise to other strata or other targets.