— Episode 08 · The cascade and its gates
Three decisions · different costs · different evidence

Three gates.

Episode 07 left the trial without an owner. This page is the working file behind the turn that follows: the response protocol is not an accessory to the model but a cascade that already exists, that a nurse's concern opens as readily as an algorithm does, and that separates into three decisions with different costs and different evidence behind each one.

Why this page exists.

Episode 08 of Road to CEPAS 2026 picks up the question Episode 07 deferred: what a response protocol has to specify before it can be the thing under test. The Substack post argues that the honest answer is not a specification at all. The cascade already runs, several times a week, in every unit; the question is whether a model belongs on the list of things that open it.

This page is the working file: the consensus description of the cascade that already exists, the three-gate decomposition and what each gate wants from a test, where the published evidence actually sits, the design contrast that produces an AUC of 0.99 in one study and 0.77 in another, the gate-1 absence and how far the search behind that claim can be pushed, and the endpoint arithmetic that puts mortality inside a core outcome set rather than at the head of it.

A note on what is mine and what is read. The three-gate structure and the claim that the cascade is trigger-agnostic are arguments, not findings, and no cited author is responsible for them. The literature below is used to establish where evidence exists and where it does not, which is a weaker and more defensible use than treating any of it as support for the structure. Where a figure comes from a development sample, or from a cohort in which suspicion had already been raised, it is not restated as though it were a figure from practice.

The cascade

The protocol is not a thing to be designed. It is a thing already running.

My first attempt at specifying a response protocol produced thresholds, refractory periods, escalation rules and alert routing, which is the model again in different clothes. The alarm policy is the last thing the model does. The response protocol is everything after it, and almost none of it is about the model.

The tenth clinical consensus of the Ibero-American Society of Neonatology describes what actually happens when neonatal sepsis is suspected: blood is drawn, venous access is used to give antibiotics, mother and child are separated, and the hospital stay lengthens; sometimes an X-ray, a urine sample and a lumbar puncture follow (Sola et al. 2020). The consensus puts the proportion of suspected cases that prove to be sepsis at generally under 10%, and at no more than 25–30%.

Table 1. What already opens the cascade on a normal day, and what none of these triggers has in common with a model. The right-hand column is the point: every entry ends in the same decisions in the same order.
  Where it comes from What follows
A nurse's concern Bedside observation, often unarticulated The same cascade
A parent's concern Bedside observation, from outside the staff The same cascade
A rising apnoea count Monitor and chart The same cascade
Feeds not tolerated, colour change Clinical examination The same cascade
An incidental CRP A sample drawn for something else The same cascade
A prediction model Continuous physiology, an algorithm The same cascade
Takeaway

If the cascade is the same regardless of what opened the door, then the protocol is shared clinical infrastructure that a model plugs into, not a model accessory. It can be specified, and in principle validated, without reference to any particular AI model at all. That is the whole of the argument, and everything below is an attempt to say where the evidence for it is and is not.


The decomposition

Three decisions, not one response.

Read as a pathway rather than as a reaction to an alarm, the cascade separates. The gates have different costs, and, more usefully for anyone designing a study, they want different things from a test.

Table 2. The three gates. The last column is why a single diagnostic performance figure cannot answer the question: a test that is excellent at gate 1 need not be useful at gate 3.
  The decision What it costs What it wants from a test
Gate 1 · Look Something raises concern. Do we go and look? Attention, an examination, possibly a blood draw in an infant who may weigh 700 g Sensitivity, because the point is not to miss the infant deteriorating silently
Gate 2 · Treat Having looked, do we culture and start antibiotics? Intravenous access, broad-spectrum exposure, and a clock that starts Specificity and timeliness together
Gate 3 · Stop At 36 to 48 hours, culture negative and infant well, do we stop? The residual risk of stopping, carried by whoever signs it off Negative predictive value, because this declares it safe to stop
Takeaway

The output of gate 1 is not antibiotics. It is attention. Collapsing the three into a single "response to the alarm" is what made the protocol look like a specification problem, and separating them is what makes each one answerable by a different kind of study.


The evidence map

The literature clusters at gates 2 and 3, and thins to nothing at gate 1.

A caveat that belongs at the top rather than in a footnote: the evidence below moves between early-onset and late-onset sepsis, because that is how the literature is distributed. The gate structure is the same in both; the trials are not interchangeable, and where a finding is about early-onset sepsis it is labelled as such.

Table 3. Where the published evidence sits, gate by gate. The bottom row is the finding this episode turns on.
  Best available evidence State
Gate 3 · Stop NeoPInS, 1,710 neonates, four countries, procalcitonin-guided decision making in suspected early-onset sepsis (Stocker et al. 2017); a systematic review of discontinuation strategies (Feng et al. 2025) One properly powered trial, and a review concluding there is insufficient evidence to name an optimal strategy
Gate 2 · Treat Biomarker literature, including a systematic review of 171 reports on early-onset sepsis (van Leeuwen et al. 2024) and the late-onset studies below Large and inconsistent; no stand-alone biomarker established as reliable
Gate 1 · Look No confirmatory test evaluated in this population at this gate that I could find

NeoPInS is worth stating precisely, because it is the strongest study at any of the three gates and it still could not do what it set out to do. Antibiotic duration fell from 65.0 to 55.1 hours on intention to treat. The co-primary non-inferiority endpoint for re-infection or death could not be demonstrated: there were no sepsis-related deaths and only nine infants with possible re-infection in the whole trial (Stocker et al. 2017).

Feng and colleagues reviewed the discontinuation question directly and found 11 randomised trials across nine regimens, ten of them at high risk of bias, and concluded that there is currently insufficient evidence to determine the optimal strategy (Feng et al. 2025). The van Leeuwen review screened 2,296 articles, included 171 in the systematic review and 69 in the meta-analysis, and reports mixed and inconsistent evidence across biomarkers and sample types (van Leeuwen et al. 2024). It is an early-onset review, and it is cited here for the shape of the biomarker evidence rather than as a statement about late-onset practice.

Takeaway

A well-conducted multicentre neonatal trial could not evaluate its own safety endpoint because the events did not happen. That is the same arithmetic that closed Episode 07, arriving from the other direction, and it recurs below when the endpoint question is taken up directly.


The design contrast

0.99 and 0.77, and what separates them.

Two late-onset studies sit next to each other and tell the story this series has been telling since Episode 03. They use different markers, so this is not a like-for-like comparison and is not offered as one. What travels between them is the design, not the analyte.

Table 4. Two late-onset diagnostic studies. Different markers, so the performance figures are not directly comparable; the design difference is what the row labels are for.
  Berka et al. 2021 Dierikx et al. 2025
Design Single-centre retrospective case-control Multicentre prospective, consecutive enrolment
Population 285 very preterm infants, 66 with late-onset bloodstream infection 63 suspected episodes in 50 infants, 32 classified as late-onset sepsis
Marker and threshold Interleukin-6 above 100 ng/L Presepsin at initial suspicion, before antibiotics
Reported performance Sensitivity 94%, specificity 99%, AUC 0.988, negative predictive value 99% AUC 0.77 (95% CI 0.66–0.89) for all cases; 0.80 for culture-positive cases only
Authors' conclusion High negative predictive value could encourage early discontinuation Insufficient accuracy as a single biomarker to rule out late-onset sepsis
Takeaway

Retrospective case-control designs produce numbers near 0.99. Prospective consecutive enrolment produces numbers near 0.77. This series documented exactly that pattern across the prediction-model literature on the literature base; the diagnostic literature does the same thing. The prospective study is also small, 63 episodes, which widens its interval and is a second reason not to read the gap as a clean effect of design alone.


The examination

0.828, produced by looking at the baby.

Long before any of this, a Dutch group asked whether the bedside examination alone could carry the work. Bekhof and colleagues followed 142 preterm infants under 34 weeks through 187 episodes of suspected infection and built a nomogram from clinical signs, with no laboratory tests in it. Four signs carried it, and the area under the curve was 0.828, 95% CI 0.764–0.892 (Bekhof et al. 2013).

Table 5. The signs that discriminated and the signs that did not, as reported. Odds ratios with 95% confidence intervals for the four that carried the nomogram.
  Signs
Discriminated Central venous catheter (OR 4.6, 2.2–10.0); increased respiratory support (OR 3.6, 1.9–7.1); grey skin (OR 2.7, 1.4–5.5); capillary refill (OR 2.2, 1.1–4.5)
Did not Temperature instability, apnoea, tachycardia, dyspnoea, hyper- and hypothermia, feeding difficulties, irritability

Two things follow, and they point in opposite directions. The first is that 0.828 sits inside the 0.78–0.88 band this series has found across every machine-learning group working on this problem. The second is that several of the signs that did not discriminate are precisely what sends a nurse to find a doctor. The gate-1 trigger list in Table 1 is therefore not a list of equally good reasons to go and look, and part of it was measured as noise more than a decade ago.

Takeaway: read with care, not discounted

The comparison with the model literature is suggestive rather than fair. The 0.828 comes from a development sample, without the external validation the prediction models are increasingly held to, and from a cohort in which suspicion had already been raised, which is a different and easier population than a whole unit under continuous monitoring. It is still a figure produced by examining infants, and that deserves more than a footnote.


The absence

Every one of these studies enrols infants somebody was already worried about.

This is the part that matters most, and it is a structural observation rather than a criticism of any study. Dierikx and colleagues consecutively included infants already started on antibiotics for suspected late-onset sepsis. Berka and colleagues measured markers when infection was clinically suspected. Bekhof and colleagues enrolled episodes of suspected infection. That is a sensible way to study a test that helps you decide whether to continue, and it tells you nothing about gate 1.

Gate 1 is where a model fires. It is the only gate at which a model can do something a clinician has not already done. It is also the gate at which I could not find a confirmatory test evaluated in this population.

On the strength of the absence claim. I also went looking for work quantifying the other gate-1 triggers, nurse concern in particular, since the nurse is the sensor this field is implicitly competing with. What I found sits in adult and paediatric emergency medicine, not in the NICU. One search is not a systematic review, and I would not claim the literature does not exist. The claim on this page is the weaker one: I could not find it, and the asymmetry is itself interesting. We have studied the tests that follow suspicion far more carefully than we have studied what creates it.
Takeaway

Our own unit has counted this in an internal audit that is not published: how often suspicion is raised, what is done when it is, and how often any of it leads anywhere. Those numbers change the shape of the question rather than its direction, and they are held for Episode 09 rather than compressed into a paragraph here, because they need to be set against a comparator and the comparator is what the next episode is about.


The endpoint

Mortality survived the consensus, as one of nine.

Episode 07 closed on mortality as an inherited endpoint. The field has since answered the question formally, and the answer is more useful than either "keep it" or "drop it".

Henry and colleagues reviewed 90 randomised trials in neonatal sepsis and found 88 distinct outcomes, with only 30 of the 90 explicitly stating a primary or secondary outcome, and survival reported in 74% (Henry et al. 2022). That review fed NESCOS, a core outcome set built through a real-time Delphi with 306 participants and an 80% agreement threshold, which reduced 55 candidate outcomes to nine (Taneri et al. 2025).

Table 6. The nine NESCOS outcomes, grouped by how close each sits to the cascade in Table 1. The grouping is mine; the outcome set is theirs.
  Outcomes
Close to the cascade Escalation of antimicrobial therapy; multiorgan dysfunction; central nervous system infection
Downstream All-cause mortality; need for mechanical ventilation; brain injury on imaging; neurologic status at discharge
Long term Neurodevelopmental impairment; quality of life of parents

Mortality survived, and that is correct. But it survived as one of nine, alongside escalation of antimicrobial therapy and multiorgan dysfunction, both far more frequent and far closer to the cascade. The field's own consensus already contains an antibiotic-exposure outcome.

The case against mortality as a primary endpoint is arithmetic, not sentiment, and NeoPInS demonstrates it: a trial powered on an outcome that has become rare cannot answer its own question. And the outcome has become rarer. Survival to discharge among extremely preterm infants in the NICHD Neonatal Research Network rose from 76.0% in 2008–2012 to 78.3% in 2013–2018, an adjusted difference of 2.0% with a 95% CI of 1.0–2.9 (Bell et al. 2022).

Takeaway: direction, not a precise figure

That is US network data over periods that do not line up exactly with the HeRO trial's enrolment, so it is offered as direction rather than as a matched comparison. The direction is the point. HeRO was powered in an era with more deaths in it, missed its primary endpoint at P = 0.083, and produced a mortality signal as a secondary outcome. Anyone inheriting that endpoint today would need a larger trial to detect a smaller effect. Episode 07 made this claim about falling mortality without a citation behind it; this is the citation.


The measurement

Endpoints, gate by gate.

If the cascade is the intervention, then the endpoints are cascade endpoints, and they arrange themselves against the gates rather than against the model.

Table 7. What a trial of the cascade would measure. The second gate-1 quantity is the model's actual contribution, and nothing in a ROC curve measures it.
  Endpoints
Gate 1 · Look Workups triggered per infant-week; and how many of them a clinician had not already initiated
Gate 2 · Treat Antibiotic starts per workup; blood cultures per start; time from concern to culture; time from culture to first dose
Gate 3 · Stop Antibiotic days per start; proportion of culture-negative episodes stopped at 36 to 48 hours
Across all three Escalation of antimicrobial therapy, which is in NESCOS; and adherence to the specified protocol, without which a null result cannot be interpreted
Takeaway: a correction to Episode 04

A point about false positives that this series had wrong until now. In a model with the precision documented in Episode 04, the false positives do not, or should not, land as unnecessary antibiotics. They land at gate 1, as unnecessary examinations and unnecessary blood draws in infants who may weigh 700 grams. That is a real cost and it is not the cost this series has been talking about. It is also, unlike antibiotic days, something the cascade could be designed to keep small.


The convergence

The model has a short half-life. The cascade does not.

Episode 07 reached the conclusion that the finding from this trial would be inherently unownable, and reached it from the commercial logic: a validated protocol is a guideline, and nobody licenses a guideline. This episode arrives at the same place from the clinical side, which is a stronger position than either argument alone.

If gates 1 to 3 are the same regardless of what opened the door, then a validated cascade works with a nurse's concern, a rising apnoea count, an incidental CRP, our model, or somebody else's. That is precisely why nobody can sell it. It is also why it would still be useful in ten years, when the model that prompted the question has been retrained, replaced or quietly switched off — the half-life problem that Episode 07 raised through the two Nature systems and their measurement problem.

Takeaway: a possibility, not a conclusion

Which raises something worth sitting with rather than resolving here: that the thing worth randomising was never the model. This is offered as a live possibility. It has consequences for the talk in Lyon that have not yet been worked through, and working them through is what the remaining episodes are for.


The ledger

What this episode settled, and what it now owes.

Takeaway

Two promises from Episode 07 are discharged here and two new ones are made, both to Episode 09. The gate-1 absence is the one item on this list that is neither settled nor promised, and it is the item the argument leans on hardest.


References & sourcing.

  1. Sola A, Mir R, Lemus L, Fariña D, Ortiz J, Golombek S; Ibero-American Society of Neonatology (SIBEN). Suspected neonatal sepsis: tenth clinical consensus of the Ibero-American Society of Neonatology (SIBEN). NeoReviews 2020;21(8):e505–e534. doi:10.1542/neo.21-8-e505  Paraphrased · no open licence recorded
  2. Stocker M, van Herk W, el Helou S, et al. Procalcitonin-guided decision making for duration of antibiotic therapy in neonates with suspected early-onset sepsis: a multicentre, randomised controlled trial (NeoPIns). Lancet 2017;390(10097):871–881. doi:10.1016/S0140-6736(17)31444-7  Elsevier · paraphrase only
  3. Feng K, Zhang T, Hua Z. Discontinuation of empirical antibiotics in suspected neonatal early-onset sepsis: a systematic review and meta-analysis. Pediatr Res 2025;99(3):871–878. doi:10.1038/s41390-025-04290-9  11 RCTs · 10 at high risk of bias
  4. van Leeuwen LM, Fourie E, van den Brink G, Bekker V, van Houten MA. Diagnostic value of maternal, cord blood and neonatal biomarkers for early-onset sepsis: a systematic review and meta-analysis. Clin Microbiol Infect 2024;30(7):850–857. doi:10.1016/j.cmi.2024.03.005  Early-onset scope · cited for shape, not practice
  5. Berka I, Korček P, Straňák Z. C-reactive protein, interleukin-6, and procalcitonin in diagnosis of late-onset bloodstream infection in very preterm infants. J Pediatric Infect Dis Soc 2021. doi:10.1093/jpids/piab071  Retrospective case-control · n = 285
  6. Dierikx TH, Admiraal J, Nusman CM, van Laerhoven H, et al. The diagnostic accuracy of presepsin for late-onset neonatal sepsis: a multicenter prospective cohort study. Pediatr Res 2025;98(5):1864–1869. doi:10.1038/s41390-025-04008-x  Prospective consecutive · 63 episodes
  7. Bekhof J, Reitsma JB, Kok JH, Van Straaten IHLM. Clinical signs to identify late-onset sepsis in preterm infants. Eur J Pediatr 2013;172(4):501–508. doi:10.1007/s00431-012-1910-6  Development sample · AUC 0.828
  8. Henry CJ, Semova G, Barnes E, et al. Neonatal sepsis: a systematic review of core outcomes from randomised clinical trials. Pediatr Res 2022;91(4):735–742. doi:10.1038/s41390-021-01883-y  90 RCTs · 88 distinct outcomes
  9. Taneri PE, Biesty L, Kirkham JJ, Molloy EJ, et al. Proposed core outcomes after neonatal sepsis: a consensus statement. JAMA Netw Open 2025;8(2):e2461554. doi:10.1001/jamanetworkopen.2024.61554  NESCOS · open access
  10. Bell EF, Hintz SR, Hansen NI, et al. Mortality, in-hospital morbidity, care practices, and 2-year outcomes for extremely preterm infants in the US, 2013–2018. JAMA 2022;327(3):248–263. doi:10.1001/jama.2021.23580  PMC8767441
  11. Moorman JR, Carlo WA, Kattwinkel J, et al. Mortality reduction by heart rate characteristic monitoring in very low birth weight neonates: a randomized trial. J Pediatr 2011;159(6):900–906.e1. doi:10.1016/j.jpeds.2011.06.044  Open access · PMC
  12. van den Berg M, Medina O, Loohuis I, et al. Development and clinical impact assessment of a machine-learning model for early prediction of late-onset sepsis. Comput Biol Med 2023;163:107156. doi:10.1016/j.compbiomed.2023.107156  Open access · CC BY-NC-ND

Citations on this page were resolved and checked against PubMed before inclusion. Figures are quoted as reported in the cited abstracts and papers; where a figure comes from a development sample, from a small prospective cohort, or from a study population in which suspicion had already been raised, that is stated in the surrounding text rather than left to the reader. This page discusses and critiques the sources above and does not reproduce them. The three-gate decomposition, the claim that the cascade is trigger-agnostic, the grouping of the NESCOS outcomes in Table 6, and the reading of where false positives land are the author's own, offered as argument rather than as findings of the original investigators.