Synthetic data validation · CLIF 2.1

Does the synthetic data behave like real critical illness?

A full audit of the fully-synthetic CLIF datasets against real CLIF — the ICU cohort (85,248 stays) and the whole-hospital population (365,000 stays) — across missingness, length-of-stay, life support, death, patient flow, lab-value shape, the longitudinal course of illness, and privacy.

ICU: 85,248 stays vs real ICU cohort Full hospital: 365,000 stays vs real full population Synthetic — generated, no real records
Where it should match
matched
LOS distribution (median and tails), all life support, missingness, lab-value shapes, and the illness trajectory land in the real range.
Where it differs by design
4
Chicago race mix, younger age, and network-median mortality — deliberate, not drift.
Lab-value distributions
shape-matched
Empirical inverse-CDF marginals reproduce skew and long tails, not just means — creatinine's kidney tail, lactate's skew.

ICU cohort — in the statistical region of the original

The ICU dataset's length-of-stay, organ support, and measurement density were fit to reproduce real CLIF ICU statistics. They do.

MetricSyntheticRealVerdict
Hospital LOS — median167 h164 hmatched
Hospital LOS — p10 / p90 (tails)45 / 497 h51 / 513 hmatched
Invasive ventilation (IMV)0.4020.412matched
Non-invasive ventilation (NIPPV)0.0580.064matched
High-flow nasal cannula0.0610.069matched
CRRT (renal replacement)0.0410.040exact
Core chemistry present (creatinine, Na, Hgb)0.9850.986exact
Lactate present (per stay)0.6740.692matched
Arterial blood gas present (per stay)0.6900.695matched

Lab values — the shape, not just the rate

Each lab is drawn through an empirical inverse-CDF fit to the real distribution, so skewed and long-tailed shapes are reproduced — a single log-normal would flatten creatinine's kidney-disease tail or lactate's skew. Values shown are the 10th / 50th / 90th percentiles.

LabSynthetic p10 / p50 / p90Real p10 / p50 / p90Verdict
Creatinine (mg/dL)0.53 / 1.0 / 2.860.5 / 1.0 / 3.2tail matched
Lactate (mmol/L)0.9 / 1.7 / 4.410.9 / 1.8 / 4.8matched
Sodium (mmol/L)132 / 138 / 144132 / 138 / 145exact
BUN (mg/dL)8 / 19 / 559 / 22 / 63matched
Platelets (10⁹/L)74 / 205 / 38464 / 193 / 393matched

The longitudinal story — deterioration toward death

The hardest test: does a dying synthetic patient get sicker the way a real one does? Aligned to the last 48 hours before death or discharge, decedents should diverge from survivors — and they do, in both datasets.

Decedents Survivors Synthetic Real CLIF

Different — by design, not by accident

The dataset is deliberately a Chicago-representative, network-median cohort — statistically distinct from its source cohort so no one dataset can be mistaken for the other. CI also locks those network-median rates (LOS, IMV, vaso, CRRT, ABG co-occurrence) against the shareable base_pack on every PR — no DUA required to run the gate.

CharacteristicSyntheticReal CLIFRationale
Race — White / Black42% / 30%67% / 9%Chicago, not Boston
Age — median6367shifted younger
In-hospital mortality0.0950.115CLIF network median
Female53%43%cohort shift

The whole hospital — realistic patient flow

The second dataset is a full hospital population (365,000 stays): most patients never leave the ward, ~15% reach the ICU, and they arrive through a realistic mix of front doors. The population statistics match real; the flow is a deliberately cleaner design (more planned ICU access, fewer floor deteriorations) than the raw source cohort.

MetricSyntheticReal full pop.Verdict
Reach the ICU0.1500.156matched
Hospital LOS — median64 h67 hmatched
Hospital LOS — p10 / p9013 / 226 h13 / 243 hmatched
In-hospital mortality0.0220.021matched
Direct admissions0.200.06by design
Arrive via ER / OR0.55 / 0.180.84 / 0.10by design
Transferred into ICU (deteriorate)0.0930.139by design <10%
Direct / planned ICU admit0.0570.017by design

Front-door mix (synthetic): ED 57% · direct-to-ward 21% · OR/procedural 11% · direct-ICU 6% · stepdown 6%.

Privacy — no record traces back

The generator samples from aggregate parameters and never copies a real record. Three standard privacy measures (synthetic vs real, 4,000 each) confirm no synthetic patient is a stand-in for a real one.

MeasureValueReading
Distance to closest real record — median (p5)3.49 (1.89)no copies even the closest 5% sit ~1.9σ from any real record
Nearest-neighbour distance ratio — median0.97no unique linkage ≈1 means not singled-out to one real record
Identifiability0.25moderate the fidelity/privacy trade-off of matching real value shapes

Honest gaps

Arterial-gas panel union runs slightly high — 0.76 vs 0.70. Real blood gases are co-ordered as a panel, and presence is drawn through a fitted Gaussian copula so co-ordered labs appear together (this corrected an earlier several-fold inflation). One residual remains: SₐO₂, measured less consistently (0.36), still contributes semi-independently, so "any arterial gas present" sits a few points above real.

White-cell counts skew a touch low. WBC p90 is 14 vs 18 real — the empirical marginal reproduces most of the distribution but slightly under-weights the high-count (leukocytosis) tail. Every other lab's value distribution tracks real, including the tails.

Verdict

Both datasets reproduce real CLIF where it counts — missingness, length-of-stay distributions, life support, death, lab-value shapes, and the death-and-recovery trajectory — the ICU cohort against the real ICU population and the full-hospital dataset against the real whole-hospital population. They are intentionally distinct where designed to be: Chicago demographics, network-median mortality, and a cleaner patient flow. And they carry no traceable records. Statistically similar, believably real, but different and safe to share.