Temporal Folds
Purpose. The original locked validation used a single train/test time-block split. This track tests whether the closure signal survives across multiple sliding splits — that is, whether the result is specific to the locked 2017–2018 test window or holds more broadly.
Method
- Restricted to FASB × Sector 2; target = closure within five years.
- Closure label requires a five-year forward window, so the latest testable observation year is 2018 (T+5 ≤ 2023). This bounds the available folds.
- Four sliding folds, each with disjoint train and test years; transforms refit on the fold’s train years only.
- Same predictors as the cluster-bootstrap track (joint-weak, G-only, A-only, divergence, three single-variable baselines).
Per-fold AUROC for joint-weak G+A
| Fold | Train years | Test years | Test rows | Closure positives | Joint-weak AUROC |
|---|---|---|---|---|---|
| Fold 1 | 2014 | 2015 | 1,075 | 21 | 0.772 |
| Fold 2 | 2014–2015 | 2016 | 1,070 | 21 | 0.840 |
| Fold 3 (locked) | 2014–2016 | 2017–2018 | 2,168 | 52 | 0.864 |
| Fold 4 | 2014–2017 | 2018 | 1,080 | 30 | 0.854 |
Joint-weak AUROC stays in [0.77, 0.86] across all four folds; lower CI > 0.65 in every fold. Within-A/G ordering (joint-weak above G-only above A-only above divergence) holds in every fold.
A/G edge over the strongest baseline
- Fold 1 (2015 test): joint-weak ≈ 0.772; the simple low six-year graduation rate baseline ≈ 0.773. Essentially tied.
- Fold 2 (2016 test): joint-weak ≈ 0.840; best baseline ≈ 0.788. Δ ≈ +0.05.
- Fold 3 (2017–2018 test, the locked split): joint-weak ≈ 0.864; best baseline ≈ 0.792. Δ ≈ +0.07.
- Fold 4 (2018 test): joint-weak ≈ 0.854; best baseline ≈ 0.797. Δ ≈ +0.06.
Honest record on Fold 1. The A/G geometry’s point-estimate edge over single-variable baselines is uneven across folds. In three of four folds the edge is +0.05 to +0.07. In Fold 1 the edge effectively disappears because the simple six-year graduation rate already captures most of the closure signal in that early test window. Phase 19 records this rather than choosing a more favourable fold to headline.
Pooled fold summary
A pooled-across-folds metric is intentionally not headlined because Folds 3 and 4 share year 2018 — pooling would double-count. Per-fold reporting is the honest framing.
- 4 of 4 folds show joint-weak AUROC ≥ 0.77 with lower CI > 0.65.
- 3 of 4 folds (2, 3, 4) show A/G higher point estimate than the best single-variable baseline by +0.05 to +0.07.
- 1 of 4 folds (Fold 1, 2015 test) shows A/G essentially tied with the simplest A-only baseline.
Limitations
- Folds 1, 2, and 4 use shorter / single-year test windows (21–30 closure positives). Per-fold CIs are wider than the locked Fold-3 CI as a result.
- Per-fold reporting uses analytic Hanley-McNeil intervals. Cluster-bootstrap uncertainty is computed only for Fold 3 (the locked split) and reported in the Cluster Bootstrap note.
- The Phase-12.5 panel is the same for every fold; only train/test years rotate.