Validation
This section reports the held-out validation evidence behind the locked higher-ed claim. Numbers come from documented Phase-15 and Phase-19 outputs; none of the modelling was re-run for this site.
Validated lane: FASB × Sector 2 — private nonprofit four-year institutions. Validated target: closure within five years. Primary geometry: joint-weak G+A.
1. Primary closure validation
On the locked test cohort (test years 2017–2018; transforms fit on train years 2014–2016 only), the joint-weak G+A geometry produces:
- AUROC point estimate ≈ 0.864.
- Institution-cluster bootstrap 95 % CI ≈ [0.806, 0.914] (1,000 UNITID-level resamples).
- Original analytic Hanley-McNeil 95 % CI ≈ [0.801, 0.928]; the cluster-bootstrap and analytic intervals are nearly identical at the lower bound.
- Top-decile lift over the test-cohort base rate (≈ 2.4 %) of about 5.96 ×.
The closure label is built strictly from future events: an institution-year T receives a positive label only when the institution’s closure year falls in {T+1, …, T+5}.
2. Temporal multi-fold validation
Four sliding time-block splits (the closure label requires a five-year forward window, so the latest testable T is 2018):
| Fold | Train years | Test years | Test rows | Closure positives | Joint-weak AUROC |
|---|---|---|---|---|---|
| Fold 1 | 2014 | 2015 | 1,075 | 21 | 0.772 |
| Fold 2 | 2014–2015 | 2016 | 1,070 | 21 | 0.840 |
| Fold 3 (locked) | 2014–2016 | 2017–2018 | 2,168 | 52 | 0.864 |
| Fold 4 | 2014–2017 | 2018 | 1,080 | 30 | 0.854 |
Joint-weak AUROC stays in [0.77, 0.86] across all four folds. Within-A/G ordering (joint-weak above G-only above A-only above divergence) holds in every fold.
Honest record on Fold 1. In Fold 1 (2015 test, 21 closure positives), the joint-weak AUROC of 0.772 is essentially tied with the simple low six-year graduation rate baseline (AUROC 0.773). The A/G geometry’s edge over single-variable baselines is uneven across folds; in three of four folds it shows a +0.05 to +0.07 AUROC point-estimate edge, and in one of four folds the edge effectively disappears. Phase-19 reports this rather than choosing a favourable fold to headline.
3. Baseline comparison
On the locked test cohort the joint-weak G+A geometry has the highest AUROC point estimate among tested predictors. The strongest single-variable baseline is the negated net-position cushion (AUROC ≈ 0.792). Cluster-bootstrap paired-difference between joint-weak and that baseline:
- Mean difference ≈ +0.072.
- 95 % cluster-bootstrap CI ≈ [+0.002, +0.138].
- About 98 % of bootstrap samples have a positive paired difference.
For the enrollment-collapse target, simple low six-year graduation rate alone has a higher AUROC point estimate than joint-weak. The A/G framework does not beat baselines for that target. See Temporal Folds for the per-target breakdown.
4. ED Financial Responsibility composite
The federal Financial Responsibility composite was fetched as a baseline comparator via the Urban Institute Education Data Portal API (a public mirror of federal data). Its public coverage in the mirror used here stops at year 2016, while the locked test cohort is 2017–2018.
On the descriptive train cohort (2014–2016 FASB × Sector 2), the negated composite score has AUROC ≈ 0.747 (95 % analytic CI ≈ [0.686, 0.808]) for closure within five years. The simple low-graduation-rate baseline on the same cohort has AUROC ≈ 0.788.
5. Heightened Cash Monitoring (HCM)
The federal Heightened Cash Monitoring list was attempted as a separate distress anchor. The federal endpoints hosting the list were not reachable from the build environment used for this work, and no public mirror of HCM is available in the same data portal that hosts the financial-responsibility composite. As a result, no HCM evidence is presented here.
HCM remains a candidate for future work. See HCM Anchor Note.
6. Lead-time pattern
Among 108 closing institutions in the FASB × Sector 2 panel, joint-weak scores were elevated above the within-cohort median across the entire T-1 to T-5 horizon (median joint-weak ≈ 0.85–1.10, well above the test-cohort median of about 0.06). At the locked top-decile cutoff:
- About 13 % of pre-closure rows are flagged at T-1.
- About 26–27 % of pre-closure rows are flagged at T-4 and T-5.
The signal is therefore a multi-year early-warning pattern, not a coincident T-1 alarm. The flagging rates are consistent with a useful precision–recall trade — not a perfect screen.
7. Bottom line
- The validated lane (FASB × Sector 2 closure within five years) has a strong, narrow, hardened result.
- Beyond that lane, the higher-ed universe remains unvalidated.
- The A/G framework does not beat baselines for enrollment-collapse targets.
- This is a research-grade artifact, not a production system.