— Living knowledge base

Road to CEPAS 2026

A 26-week public series on AI and neonatal sepsis prediction, built toward a 20-minute talk in Lyon on 31 October 2026. Twelve essays. One growing literature base. One honest argument.

Episode 09 of 12 · Updated 19 Sep 2026
— The argument

Prediction is not the same as clinical utility.

The field can predict late-onset sepsis in preterm neonates reasonably well. It has been able to for years. There is no shortage of models with respectable ROC curves, and our group has contributed to that pile.

What no one has done is build the bridge from prediction to clinical utility. We have models. We do not have validated response protocols. We do not have implementation data at scale. We do not have answers to the question that follows every alarm: now what?

That gap is the subject of this series, and the central argument of the talk in Lyon.

— The series

Twelve essays through October. One postscript from Lyon.

01
Framing & argument · the gap between prediction and action
02
Live PubMed search with Claude · queries archived · seven-group field map · methodology log
03
Four anchor papers, eight-dimension cross-study comparison · comparison grid
04
Critical self-appraisal of the 2023 LOS prediction paper · signal stack, alarm policy, and the proxy at the centre · methodology log
05
EU MDR classification, Rule 11, and the Epic Sepsis Model as cautionary tale · methodology log
06
What a real clinical-impact trial would take · the HeRO trial, its split endpoints, and the fifteen years since · methodology log
07
Incentives, funding and ownership · why a validated response protocol is a public good nobody is paid to produce · methodology log
08
The cascade that already runs · three decisions, three kinds of evidence, and the one gate a model fires at · methodology log
09
Routine care as the baseline a model competes with · gate 1 counted in an internal audit · replace or add · methodology log
10
Big Data in the NICU — What Actually Exists
Data landscape mapping · registries, EHRs, monitor streams
Upcoming
11
Building the Evidence Website
How this knowledge base was built · the meta-episode
Upcoming
12
The Presentation Takes Shape
Slide logic, narrative structure, the final argument · and see you in Lyon
Upcoming
13
From the Stage
Slides, recording, audience Q&A · post-congress
After CEPAS
— Literature base

Every search, every paper, every choice, archived here.

From Episode 02 onwards, the evidence work behind every essay lives on this page. PubMed queries with their results. Extraction tables with their data. Methodological choices with their reasoning. By the time of the talk, you should be able to see the full chain of reasoning from raw literature to final slide, and copy the workflow if it helps you.

Episode 02 added: the seven academic groups doing continuous-physiology machine-learning work on late-onset sepsis prediction in preterm infants, plus the two systematic reviews that frame the field. Full methodology log with verbatim PubMed queries is on the Episode 02 search page.

Episode 03 added: the eight-dimension comparative deep-read of the four anchor papers, Berg 2023, Kausch 2023, Yang 2024, Meeus 2024, across cohort design, signal stack, ML methodology, validation, performance, alarm policy, limitations, and contributions. The full grid, with the three supporting-cast papers and licensing notes, is on the Episode 03 comparison page.

Episode 04 added: the same two decisions, signal stack and alarm policy, turned onto our own paper, Berg 2023, read against its supplement. The detection-fraction-versus-precision framing, the sensitivity of the headline recall to the refractory period and true-positive window (supplement Tables S3, S7, S8), the blood-culture-time proxy, and an unresolved supplementary CRP discrepancy are set out on the Episode 04 methodology log.

Episode 05 added: the EU regulatory landscape a model like Berg 2023 has to clear before the bedside, the MDR/IVDR distinction, the Rule 11 classification ladder (Class IIb as the most defensible reading, not a formal determination), the Epic Sepsis Model as a deployed-but-unvalidated cautionary example, and the state of the AI Act overlay and the pending December 2025 MDR/IVDR simplification proposal. Full sourcing is on the Episode 05 methodology log.

Episode 06 added: the clinical-impact rung of the validation ladder, the HeRO randomised trial read strictly against its own registered endpoints, the counterfactual problem that keeps the mechanism unprovable, the design requirements a trial of a model like Berg 2023 would inherit, and the search receipt behind the claim that fifteen years on there is still only one randomised outcome trial in this literature. Full sourcing is on the Episode 06 methodology log.

Episode 07 added: the incentive structure around the missing second trial, the 510(k) clearance that preceded the 2011 study and made it commercially unnecessary, the asymmetry between what a model publication costs and what a prospective trial costs, two far better-funded medical-AI groups stopping at the same rung of the validation ladder, the enrolment arithmetic set against BeNeDuctus and STOP-BPD, and the argument that a validated response protocol is a public good that no private actor can recover the cost of. Full sourcing is on the Episode 07 methodology log.

Episode 08 added: a third branch. The response protocol turns out to be a cascade that already runs in every unit, described in the SIBEN consensus, opened as readily by a nurse's concern as by an algorithm, and separating into three gates with different costs and different evidence behind each. Added here: the NeoPInS trial and the discontinuation and biomarker reviews that sit at gates 2 and 3, the Berka–Dierikx design contrast that produces 0.99 retrospectively and 0.77 prospectively, the Bekhof nomogram at 0.828 from clinical signs alone, and the NESCOS core outcome set that keeps mortality as one of nine. The absence at gate 1, where a model actually fires, is recorded as an absence. Full sourcing is on the Episode 08 methodology log.

Episode 09 added: a gate-1 re-reading of the literature. Every model paper this series has used anchors onset to something recorded (a culture, a documented evaluation, the CRASH moment, a ±7-day window around a sent culture) or starts with infants already under evaluation, so none counts gate 1 as an ordinary baseline. Added here: the Griffin event definition behind HeRO, and the Masino, Cabrera-Quiros, RALIS and Mani anchors. Set against them, an unpublished internal audit, described in general terms only, puts documented evaluation closer to once a day than three, mostly a look and mostly without a culture. The funding instruments promised alongside the comparator were drafted and cut, and remain open. Full sourcing is on the Episode 09 methodology log.

— Branch A · Continuous-physiology ML for LOS in preterm infants
UVA / Med Predictive Science US
Kausch SL et al. 2023, Pediatric Research. 10.1038/s41390-022-02444-7
The HeRO heritage. Only multi-NICU external validation in the field, trained on UVA, validated at Columbia and Washington University. Pulse-rate-equals-ECG finding makes the model deployable on a standalone oximeter.
WKZ Utrecht NL
van den Berg M et al. 2023, Computers in Biology and Medicine. 10.1016/j.compbiomed.2023.107156
Largest single-center matched cohort in the field (292 LOS / 1,497 control). Alarm-fatigue framework with 8-hour shift-length refractory period and multi-threshold escalation, since adopted by Yang 2024 and Meeus 2024. Our paper.
Máxima MC / Eindhoven · Erasmus MC · UVA NL / US
Peng Z, Schouten JS, Silvertand D, et al. 2025, IEEE Transactions on Biomedical Engineering. 10.1109/TBME.2025.3618080
Comparative external validation of late-onset sepsis prediction models across multiple clinical environments, with Lake and Sullivan from the HeRO lineage among the authors. Directly relevant to the fifth obstacle in Episode 07: how much of a single-centre model's performance survives a change of centre. IEEE copyright, so paraphrase only.
Eindhoven / Máxima MC NL
Yang M et al. 2024, Computer Methods and Programs in Biomedicine. 10.1016/j.cmpb.2024.108335
Most thorough feature ablation in the field, tested raw waveforms vs 1 Hz vs 1/min vs 1/hour sampling, and HR alone vs HR+RR+SpO2 combinations. Sino-Dutch collaboration (Southeast University Nanjing on methodology, Máxima MC on data).
Antwerp / Innocens BV BE
Meeus M et al. 2024, Journal of Pediatrics. 10.1016/j.jpeds.2023.113869
Joint LOS and NEC prediction. Only commercial spinoff in the field (Innocens BV), currently navigating European MDR/CE certification.
Karolinska Institutet / KTH Stockholm SE
Honoré A et al. 2023, Acta Paediatrica. 10.1111/apa.16660
Bridges both branches, Naive Bayes combining heart-rate characteristics, respiratory and oxygenation signals, and clinical features. Senior co-authors wrote the Persad 2021 systematic review.
Rennes FR
Leon C et al. 2021, IEEE Journal of Biomedical and Health Informatics. 10.1109/JBHI.2020.3021662
Visibility-graph analysis of heart rate variability. Methodologically distinctive in the field, exploits non-linear graph properties rather than time-domain statistics.
Lausanne CH
Rio L et al. 2022, Pediatric Research. 10.1038/s41390-021-01913-9
The only independent real-world evaluation of the commercial HeRO score outside the US. Showed strong gestational-age-dependent performance, sensitivity 76% below 28 weeks, falling to 25% above 32 weeks.
— Clinical-impact evidence · the HeRO randomised trial and its secondary analyses
Moorman 2011 · the trial US
Moorman JR et al. 2011, Journal of Pediatrics. 10.1016/j.jpeds.2011.06.044
The only randomised clinical-outcome trial in this field. 3,003 VLBW infants, nine US NICUs, 2004–2010; HRC index displayed vs masked. Primary endpoint (days alive and ventilator-free at 120 days) not significant, P = .08. Mortality, a secondary endpoint, fell 10.2% to 8.1% (HR 0.78, 95% CI 0.61–0.99, P = .04). Open access via PMC.
Fairchild 2013 · mechanism US
Fairchild KD et al. 2013, Pediatric Research. 10.1038/pr.2013.136
Secondary analysis of 2,989 trial infants asking whether the mortality benefit was septicaemia-associated. Incidence and organism distribution similar across arms; 30-day mortality after culture-positive LOS lower in the displayed arm. The authors offer earlier diagnosis as a possible explanation rather than a demonstrated one.
Schelonka 2020 · follow-up US
Schelonka RL et al. 2020, Journal of Pediatrics. 10.1016/j.jpeds.2019.12.066
Neurodevelopmental follow-up at 18–22 months for infants ≤1000 g. The composite of death or NDI was not significantly reduced (38.9% vs 44.4%; RR 0.87, 95% CI 0.73–1.05, P = .17), with the outcome available for 72% of eligible infants. The mortality reduction persisted in this subset.
King 2021 · subgroup US
King WE et al. 2021, Early Human Development. 10.1016/j.earlhumdev.2021.105419
Narrower analysis restricted to extremely low birthweight infants who developed sepsis, reporting an absolute reduction in the death-or-NDI composite. A real signal, but subgroup-level and from a single trial. First author is affiliated with the device manufacturer.
Sullivan 2023 · equity US
Sullivan BA et al. 2023, Pediatric Research. 10.1038/s41390-023-02470-z
Asks whether the score behaved equally for Black and White infants. Prediction performance was similar, and no differential effect on sepsis, mortality, antibiotic days, length of stay or ventilator days. Display did increase blood cultures in White but not Black infants, a differential effect on clinical response rather than on the model itself.
— Systematic reviews
Koppens 2023 NL
Koppens HJ et al. 2023, Neonatology. 10.1159/000531118
Branch A systematic review (HRC monitoring for LOS in preterm). Searched four databases; 15 papers, of which three report the single identified randomised trial. Concludes that methodological weakness and limited generalisability do not justify putting HRC monitoring into routine care, and calls for a large international randomised trial. This is the field's own policy verdict, and the evidential backbone of Episode 06.
Kainth 2024 IN
Kainth D, Prakash S, Sankar MJ. 2024, Pediatric Infectious Disease Journal. 10.1097/INF.0000000000004409
Branch B systematic review (clinical and laboratory features). 19 studies, 76 ML models. Pooled AUROC 0.94 but 18 of 19 studies at high risk of bias, no external validation, almost all from high-income settings. Explicitly excludes vital-signs-only studies, the methodological line that separates the two branches.
— Branch C · The cascade the alarm opens into · gates 1 to 3
SIBEN 2020 · the cascade itself Ibero-America
Sola A et al. 2020, NeoReviews. 10.1542/neo.21-8-e505
The tenth clinical consensus, and the reference description of what actually happens when neonatal sepsis is suspected: blood drawn, venous access used, antibiotics started, mother and child separated, stay lengthened. Puts the proportion of suspected cases that prove to be sepsis at generally under 10%, and no more than 25–30%. The evidential basis for treating the response protocol as existing infrastructure rather than as something to be designed.
NeoPInS · gate 3 NL / CH / CA / CZ
Stocker M et al. 2017, Lancet. 10.1016/S0140-6736(17)31444-7
1,710 neonates, 18 hospitals, procalcitonin-guided decision making for suspected early-onset sepsis. Antibiotic duration fell from 65.0 to 55.1 hours on intention to treat, but the co-primary non-inferiority endpoint for re-infection or death could not be demonstrated: no sepsis-related deaths and nine possible re-infections in the whole trial. The rare-endpoint arithmetic of Episode 07, arriving from the stopping side.
Feng 2025 · gate 3 review CN
Feng K, Zhang T, Hua Z. 2025, Pediatric Research. 10.1038/s41390-025-04290-9
Discontinuation strategies for suspected early-onset sepsis: 11 randomised trials across nine regimens, 10 at high risk of bias, and no basis for naming an optimal strategy. Notes the near-absence of trials comparing the guideline-recommended 36–48 hour course with anything else.
van Leeuwen 2024 · gate 2 biomarkers NL
van Leeuwen LM et al. 2024, Clinical Microbiology and Infection. 10.1016/j.cmi.2024.03.005
2,296 articles screened, 171 included, 69 in the meta-analysis. Mixed and inconsistent evidence across biomarkers and sample types, with no uniform case definition. Early-onset scope, so cited for the shape of the biomarker evidence rather than as a statement about late-onset practice.
Berka 2021 vs Dierikx 2025 · the design contrast CZ / NL
Berka I et al. 2021, JPIDS. 10.1093/jpids/piab071  ·  Dierikx TH et al. 2025, Pediatric Research. 10.1038/s41390-025-04008-x
Retrospective case-control, 285 infants, IL-6 above 100 ng/L at AUC 0.988 and NPV 99%. Prospective consecutive enrolment, 63 suspected episodes, presepsin at AUC 0.77. Different markers, so not a like-for-like comparison — but the same design gradient this literature base documents across Branch A, now visible in the diagnostic literature.
Bekhof 2013 · the examination NL
Bekhof J et al. 2013, European Journal of Pediatrics. 10.1007/s00431-012-1910-6
142 preterm infants under 34 weeks, 187 episodes, a nomogram from clinical signs with no laboratory tests in it: central venous catheter, increased respiratory support, grey skin, capillary refill. AUC 0.828, inside the band this series finds across every ML group — from a development sample, in a cohort where suspicion had already been raised. Also names the signs that did not discriminate, several of which are exactly what sends a nurse to find a doctor.
— Branch D · The comparator · where each study puts gate 1
Griffin 2001, 2003 · the HeRO event definition US
Griffin MP, Moorman JR. 2001, Pediatrics. 10.1542/peds.107.1.97  ·  Griffin MP et al. 2003, Pediatric Research. 10.1203/01.PDR.0000064904.05313.D2
Sepsis and sepsis-like illness defined as abrupt clinical deterioration that prompted physicians to obtain blood cultures and start antibiotics. The event is defined by the clinical action that followed it, and that definition carried into the randomised trial, which left the response to the clinician and never counted gate 1.
Masino 2019 · anchored to the evaluation US
Masino AJ et al. 2019, PLoS One. 10.1371/journal.pone.0212665
EHR-based models at CHOP, retrospective case-control; case data from the 44-hour window ending four hours before a documented sepsis evaluation. A recorded clinical action is the anchor, so the moment somebody first became uneasy is invisible.
Cabrera-Quiros 2021 · the CRASH moment NL
Cabrera-Quiros L et al. 2021, Critical Care Explorations. 10.1097/CCE.0000000000000302
Monitoring signals in 32 infants with culture-proven late-onset sepsis and 32 matched controls, anchored to Cultures, Resuscitation, and Antibiotics Started Here. Different definition, same structure: retrospective analysis needs a timestamp.
RALIS · Mithal 2018 US / IL
Mithal LB et al. 2018, Early Human Development. 10.1016/j.earlhumdev.2018.01.008
Multi-vital-sign algorithm; suspicion timed as the moment a culture was sent, and an alert counted as associated if within seven days before or after. Clinical suspicion is timestamped, but the window is too wide to locate gate 1.
Mani 2014 · after gate 1 US
Mani S et al. 2014, JAMIA. 10.1136/amiajnl-2013-001854
299 infants already evaluated for late-onset sepsis; the model helps with the decision that follows. Gate 1 has been passed before the paper starts.
Van Calster 2026 · methods background BE / NL
Van Calster B et al. 2026, Annual Review of Statistics and Its Application. 10.1146/annurev-statistics-042324-123749
Review of the statistical and scientific obstacles to reliable and useful clinical prediction models. Cited in Episode 09 alongside the decision-analytic reading of routine care as the treat-all strategy a model has to beat on net benefit. Not indexed in PubMed; resolved through Crossref.
— Outcomes · what a trial of the cascade would measure
Henry 2022 · the problem IE
Henry CJ et al. 2022, Pediatric Research. 10.1038/s41390-021-01883-y
90 randomised trials in neonatal sepsis, 88 distinct outcomes, only 30 explicitly stating a primary or secondary outcome, survival reported in 74%. The systematic review that fed the core outcome set below.
NESCOS 2025 · the answer International
Taneri PE et al. 2025, JAMA Network Open. 10.1001/jamanetworkopen.2024.61554
Real-time Delphi with 306 participants and an 80% agreement threshold, reducing 55 candidate outcomes to nine. Mortality survived, as one of nine, alongside escalation of antimicrobial therapy and multiorgan dysfunction — both far more frequent and far closer to the cascade. The field's own consensus already contains an antibiotic-exposure outcome.
Bell 2022 · the denominator US
Bell EF et al. 2022, JAMA. 10.1001/jama.2021.23580
NICHD Neonatal Research Network, 10,877 infants at 22–28 weeks. Survival to discharge rose from 76.0% in 2008–2012 to 78.3% in 2013–2018, adjusted difference 2.0% (95% CI 1.0–2.9). Supplies the source for the falling-mortality claim Episode 07 made without one. US network data over periods that do not align exactly with HeRO's enrolment, so read as direction rather than as a matched comparison.
— How this is built

A research method, made visible.

The series is built with Claude as a research and writing partner. Not as decoration, as the actual working method. PubMed searches, paper extraction, evidence tables, regulatory mapping, slide logic.

Every essay includes a "How I used Claude" section showing exactly where the AI helped and where it didn't. The methodology is the second deliverable. If you want to replicate it for your own field, everything you need will be here by October. The Episode 02 methodology log is the first full receipt.

— Destination

Lyon, 31 October 2026.

— The talk

Big data, AI and sepsis prediction opportunities in the NICU

A 20-minute talk for the Omics in Sepsis session at CEPAS 2026, the first Congress of the European Paediatric Academic Societies.

Congress CEPAS 2026 · 1st of its kind
Dates 28–31 October 2026
Venue Centre de Congrès, Lyon
Session Omics in Sepsis · Sat 31 Oct, 10:30–11:30 CET
© 2026 Daniel Vijlbrief · Utrecht, NL