A 26-week public series on AI and neonatal sepsis prediction, built toward a 20-minute talk in Lyon on 31 October 2026. Twelve essays. One growing literature base. One honest argument.
The field can predict late-onset sepsis in preterm neonates reasonably well. It has been able to for years. There is no shortage of models with respectable ROC curves, and our group has contributed to that pile.
What no one has done is build the bridge from prediction to clinical utility. We have models. We do not have validated response protocols. We do not have implementation data at scale. We do not have answers to the question that follows every alarm: now what?
That gap is the subject of this series, and the central argument of the talk in Lyon.
From Episode 02 onwards, the evidence work behind every essay lives on this page. PubMed queries with their results. Extraction tables with their data. Methodological choices with their reasoning. By the time of the talk, you should be able to see the full chain of reasoning from raw literature to final slide, and copy the workflow if it helps you.
Episode 02 added: the seven academic groups doing continuous-physiology machine-learning work on late-onset sepsis prediction in preterm infants, plus the two systematic reviews that frame the field. Full methodology log with verbatim PubMed queries is on the Episode 02 search page.
Episode 03 added: the eight-dimension comparative deep-read of the four anchor papers, Berg 2023, Kausch 2023, Yang 2024, Meeus 2024, across cohort design, signal stack, ML methodology, validation, performance, alarm policy, limitations, and contributions. The full grid, with the three supporting-cast papers and licensing notes, is on the Episode 03 comparison page.
Episode 04 added: the same two decisions, signal stack and alarm policy, turned onto our own paper, Berg 2023, read against its supplement. The detection-fraction-versus-precision framing, the sensitivity of the headline recall to the refractory period and true-positive window (supplement Tables S3, S7, S8), the blood-culture-time proxy, and an unresolved supplementary CRP discrepancy are set out on the Episode 04 methodology log.
Episode 05 added: the EU regulatory landscape a model like Berg 2023 has to clear before the bedside, the MDR/IVDR distinction, the Rule 11 classification ladder (Class IIb as the most defensible reading, not a formal determination), the Epic Sepsis Model as a deployed-but-unvalidated cautionary example, and the state of the AI Act overlay and the pending December 2025 MDR/IVDR simplification proposal. Full sourcing is on the Episode 05 methodology log.
Episode 06 added: the clinical-impact rung of the validation ladder, the HeRO randomised trial read strictly against its own registered endpoints, the counterfactual problem that keeps the mechanism unprovable, the design requirements a trial of a model like Berg 2023 would inherit, and the search receipt behind the claim that fifteen years on there is still only one randomised outcome trial in this literature. Full sourcing is on the Episode 06 methodology log.
Episode 07 added: the incentive structure around the missing second trial, the 510(k) clearance that preceded the 2011 study and made it commercially unnecessary, the asymmetry between what a model publication costs and what a prospective trial costs, two far better-funded medical-AI groups stopping at the same rung of the validation ladder, the enrolment arithmetic set against BeNeDuctus and STOP-BPD, and the argument that a validated response protocol is a public good that no private actor can recover the cost of. Full sourcing is on the Episode 07 methodology log.
Episode 08 added: a third branch. The response protocol turns out to be a cascade that already runs in every unit, described in the SIBEN consensus, opened as readily by a nurse's concern as by an algorithm, and separating into three gates with different costs and different evidence behind each. Added here: the NeoPInS trial and the discontinuation and biomarker reviews that sit at gates 2 and 3, the Berka–Dierikx design contrast that produces 0.99 retrospectively and 0.77 prospectively, the Bekhof nomogram at 0.828 from clinical signs alone, and the NESCOS core outcome set that keeps mortality as one of nine. The absence at gate 1, where a model actually fires, is recorded as an absence. Full sourcing is on the Episode 08 methodology log.
Episode 09 added: a gate-1 re-reading of the literature. Every model paper this series has used anchors onset to something recorded (a culture, a documented evaluation, the CRASH moment, a ±7-day window around a sent culture) or starts with infants already under evaluation, so none counts gate 1 as an ordinary baseline. Added here: the Griffin event definition behind HeRO, and the Masino, Cabrera-Quiros, RALIS and Mani anchors. Set against them, an unpublished internal audit, described in general terms only, puts documented evaluation closer to once a day than three, mostly a look and mostly without a culture. The funding instruments promised alongside the comparator were drafted and cut, and remain open. Full sourcing is on the Episode 09 methodology log.
The series is built with Claude as a research and writing partner. Not as decoration, as the actual working method. PubMed searches, paper extraction, evidence tables, regulatory mapping, slide logic.
Every essay includes a "How I used Claude" section showing exactly where the AI helped and where it didn't. The methodology is the second deliverable. If you want to replicate it for your own field, everything you need will be here by October. The Episode 02 methodology log is the first full receipt.
A 20-minute talk for the Omics in Sepsis session at CEPAS 2026, the first Congress of the European Paediatric Academic Societies.