— The argument
Prediction is not the same as clinical utility.
The field can predict late-onset sepsis in preterm neonates reasonably well. It has been able to for years. There is no shortage of models with respectable ROC curves, and our group has contributed to that pile.
What no one has done is build the bridge from prediction to clinical utility. We have models. We do not have validated response protocols. We do not have implementation data at scale. We do not have answers to the question that follows every alarm: now what?
That gap is the subject of this series, and the central argument of the talk in Lyon.
— Literature base
Every search, every paper, every choice — archived here.
From Episode 02 onwards, the evidence work behind every essay lives on this page. PubMed queries with their results. Extraction tables with their data. Methodological choices with their reasoning. By the time of the talk, you should be able to see the full chain of reasoning from raw literature to final slide — and copy the workflow if it helps you.
Episode 02 added: the seven academic groups doing continuous-physiology machine-learning work on late-onset sepsis prediction in preterm infants, plus the two systematic reviews that frame the field. Full methodology log with verbatim PubMed queries is on the Episode 02 search page.
Episode 03 added: the eight-dimension comparative deep-read of the four anchor papers — Berg 2023, Kausch 2023, Yang 2024, Meeus 2024 — across cohort design, signal stack, ML methodology, validation, performance, alarm policy, limitations, and contributions. The full grid, with the three supporting-cast papers and licensing notes, is on the Episode 03 comparison page.
Episode 04 added: the same two decisions — signal stack and alarm policy — turned onto our own paper, Berg 2023, read against its supplement. The detection-fraction-versus-precision framing, the sensitivity of the headline recall to the refractory period and true-positive window (supplement Tables S3, S7, S8), the blood-culture-time proxy, and an unresolved supplementary CRP discrepancy are set out on the Episode 04 methodology log.
Episode 05 added: the EU regulatory landscape a model like Berg 2023 has to clear before the bedside — the MDR/IVDR distinction, the Rule 11 classification ladder (Class IIb as the most defensible reading, not a formal determination), the Epic Sepsis Model as a deployed-but-unvalidated cautionary example, and the state of the AI Act overlay and the pending December 2025 MDR/IVDR simplification proposal. Full sourcing is on the Episode 05 methodology log.
Episode 06 added: the clinical-impact rung of the validation ladder — the HeRO randomised trial read strictly against its own registered endpoints, the counterfactual problem that keeps the mechanism unprovable, the design requirements a trial of a model like Berg 2023 would inherit, and the search receipt behind the claim that fifteen years on there is still only one randomised outcome trial in this literature. Full sourcing is on the Episode 06 methodology log.
— Branch A · Continuous-physiology ML for LOS in preterm infants
UVA / Med Predictive Science
US
The HeRO heritage. Only multi-NICU external validation in the field — trained on UVA, validated at Columbia and Washington University. Pulse-rate-equals-ECG finding makes the model deployable on a standalone oximeter.
WKZ Utrecht
NL
Largest single-center matched cohort in the field (292 LOS / 1,497 control). Alarm-fatigue framework with 8-hour shift-length refractory period and multi-threshold escalation, since adopted by Yang 2024 and Meeus 2024. Our paper.
Eindhoven / Máxima MC
NL
Most thorough feature ablation in the field — tested raw waveforms vs 1 Hz vs 1/min vs 1/hour sampling, and HR alone vs HR+RR+SpO2 combinations. Sino-Dutch collaboration (Southeast University Nanjing on methodology, Máxima MC on data).
Antwerp / Innocens BV
BE
Joint LOS and NEC prediction. Only commercial spinoff in the field (Innocens BV), currently navigating European MDR/CE certification.
Karolinska Institutet / KTH Stockholm
SE
Bridges both branches — Naive Bayes combining heart-rate characteristics, respiratory and oxygenation signals, and clinical features. Senior co-authors wrote the Persad 2021 systematic review.
Rennes
FR
Visibility-graph analysis of heart rate variability. Methodologically distinctive in the field — exploits non-linear graph properties rather than time-domain statistics.
Lausanne
CH
The only independent real-world evaluation of the commercial HeRO score outside the US. Showed strong gestational-age-dependent performance — sensitivity 76% below 28 weeks, falling to 25% above 32 weeks.
— Clinical-impact evidence · the HeRO randomised trial and its secondary analyses
Moorman 2011 · the trial
US
The only randomised clinical-outcome trial in this field. 3,003 VLBW infants, nine US NICUs, 2004–2010; HRC index displayed vs masked. Primary endpoint (days alive and ventilator-free at 120 days) not significant, P = .08. Mortality, a secondary endpoint, fell 10.2% to 8.1% (HR 0.78, 95% CI 0.61–0.99, P = .04). Open access via PMC.
Fairchild 2013 · mechanism
US
Secondary analysis of 2,989 trial infants asking whether the mortality benefit was septicaemia-associated. Incidence and organism distribution similar across arms; 30-day mortality after culture-positive LOS lower in the displayed arm. The authors offer earlier diagnosis as a possible explanation rather than a demonstrated one.
Schelonka 2020 · follow-up
US
Neurodevelopmental follow-up at 18–22 months for infants ≤1000 g. The composite of death or NDI was not significantly reduced (38.9% vs 44.4%; RR 0.87, 95% CI 0.73–1.05, P = .17), with the outcome available for 72% of eligible infants. The mortality reduction persisted in this subset.
King 2021 · subgroup
US
Narrower analysis restricted to extremely low birthweight infants who developed sepsis, reporting an absolute reduction in the death-or-NDI composite. A real signal, but subgroup-level and from a single trial. First author is affiliated with the device manufacturer.
Sullivan 2023 · equity
US
Asks whether the score behaved equally for Black and White infants. Prediction performance was similar, and no differential effect on sepsis, mortality, antibiotic days, length of stay or ventilator days. Display did increase blood cultures in White but not Black infants — a differential effect on clinical response rather than on the model itself.
— Systematic reviews
Koppens 2023
NL
Branch A systematic review (HRC monitoring for LOS in preterm). Searched four databases; 15 papers, of which three report the single identified randomised trial. Concludes that methodological weakness and limited generalisability do not justify putting HRC monitoring into routine care, and calls for a large international randomised trial. This is the field's own policy verdict, and the evidential backbone of Episode 06.
Kainth 2024
IN
Branch B systematic review (clinical and laboratory features). 19 studies, 76 ML models. Pooled AUROC 0.94 but 18 of 19 studies at high risk of bias, no external validation, almost all from high-income settings. Explicitly excludes vital-signs-only studies — the methodological line that separates the two branches.
— How this is built
A research method, made visible.
The series is built with Claude as a research and writing partner. Not as decoration — as the actual working method. PubMed searches, paper extraction, evidence tables, regulatory mapping, slide logic.
Every essay includes a "How I used Claude" section showing exactly where the AI helped and where it didn't. The methodology is the second deliverable. If you want to replicate it for your own field, everything you need will be here by October. The Episode 02 methodology log is the first full receipt.
— Destination
Lyon, 31 October 2026.
— The talk
Big data, AI and sepsis prediction opportunities in the NICU
A 20-minute talk for the Omics in Sepsis session at CEPAS 2026 — the first Congress of the European Paediatric Academic Societies.