Methodology
The Higher-Ed A/G Research Extension applies a two-layer geometry — a Ground (G) layer of institutional financial structure and an Appearance (A) layer of institutional performance and continuation signals — to U.S. higher-education institutions, using public IPEDS data covering fiscal years 2014 through 2023.
1. Data sources
All inputs are real public NCES/IPEDS releases, fetched with documented provenance and SHA-256 fingerprints. No private data, no third-party scraping, and no synthetic substitution.
- HD — institutional directory (sector, control, level, Title-IV status,
CLOSEDAT,CYACTIVE). - F2 / FASB — finance for private nonprofit institutions.
- F1A / GASB — finance for public institutions.
- F3 — finance for for-profit institutions.
- ADM — admissions counts (applied / admitted / enrolled).
- GR — graduation rates (bachelor’s 150 % cohort).
- EFFY — 12-month enrollment.
- HD
CLOSEDAT/CYACTIVE— closure year used to define the closure-within-five-years label.
The federal Financial Responsibility composite score was fetched as a comparator via the Urban Institute Education Data Portal (a public mirror of federal data). Its coverage stops at 2016 in the public mirror, which limits how it can be compared with the locked test cohort. See ED Financial Responsibility Baseline.
2. Validated lane
The validated lane is FASB × Sector 2 — private nonprofit four-year institutions. This lane was selected because:
- FASB-filed financial statements are internally comparable across institutions.
- Bachelor’s-cohort graduation rates and admissions reporting are well populated for these institutions.
- Closure events are observable and dated through HD
CLOSEDAT.
Other lanes were not validated:
- GASB / public sector institutions — closure events in public four-year institutions are very rare in the panel window; only enrollment-collapse targets could be tested, with mixed results.
- For-profit institutions — endowment-cushion is structurally absent for this family, and admissions/graduation reporting is sparser. The closure regime in this family also differs.
- Two-year and community-college institutions — graduation/admissions reporting is structured around different cohorts, and a two-year-appropriate Appearance layer has not been built.
3. G and A in higher education
The Ground layer summarises institutional financial structure: operating margin, net-position cushion, endowment cushion (where applicable), and leverage. These are computed family-specifically (FASB / GASB / for-profit) because the underlying chart of accounts differs across families. Cross-family numerical pooling is not performed.
The Appearance layer summarises institutional performance and continuation: graduation rate, admissions selectivity, and yield. This Appearance layer is well defined for four-year selective lanes and is not extended to two-year sectors in this work.
4. Geometry
The validated geometry is joint-weak G+A — institutions whose Ground and Appearance scores are both below their within-stratum medians. In the validated lane this geometry produced the strongest held-out closure AUROC and the strongest top-decile lift among the tested predictors.
Two other geometries were tested:
- Divergence (Appearance minus Ground) — empirically the weakest predictor of closure in this lane.
- Single-layer scores (G-only and A-only) — informative but lower point estimates than joint-weak.
For the secondary enrollment-collapse target, simple single-variable baselines (notably low six-year graduation rate) had higher point estimates than the A/G geometries. The Appearance-side baseline is competitive there and the A/G framework does not improve on it. This finding is reported honestly in the validation section.
5. Validation design
- Closure-within-five-years label — built from HD
CLOSEDAT: the label fires for an institution-year T when the institution’s closure year is in {T+1, …, T+5}. The label is strictly forward-looking; closures occurring at or before the observation year are excluded. - Held-out time-block validation — train years are strictly earlier than test years.
- Train-only transform fitting — winsorisation bounds and robust-z parameters (median and MAD-based scale) are computed on training data only and applied to test data.
- Institution-cluster bootstrap — resamples UNITIDs at the cluster level so test rows of the same institution stay together. 1,000 resamples; reproducible RNG seed.
- Temporal multi-fold validation — four sliding train/test splits to test whether the result is specific to a single time window.
- No future leakage — features come from observation-year IPEDS records or earlier; no future-window data enters the predictor set.
6. What was not done
- No universal higher-ed validation. Other lanes were tested and either failed to support a claim or were excluded from scope.
- No public-sector, two-year-sector, or for-profit-sector validity claim.
- No live institution ranking, no institution-level published scores, and no operational risk listing.
- No production model, no automated decision system, no calibrated probability.
- No head-to-head comparison against the federal Financial Responsibility composite on the locked test cohort — the public mirror used here covers 2014–2016 while the locked test cohort is 2017–2018.
- No HCM (Heightened Cash Monitoring) anchor — that source was not reachable in this run.