Episode 07 left the trial without an owner. This page is the working file behind the turn that follows: the response protocol is not an accessory to the model but a cascade that already exists, that a nurse's concern opens as readily as an algorithm does, and that separates into three decisions with different costs and different evidence behind each one.
Episode 08 of Road to CEPAS 2026 picks up the question Episode 07 deferred: what a response protocol has to specify before it can be the thing under test. The Substack post argues that the honest answer is not a specification at all. The cascade already runs, several times a week, in every unit; the question is whether a model belongs on the list of things that open it.
This page is the working file: the consensus description of the cascade that already exists, the three-gate decomposition and what each gate wants from a test, where the published evidence actually sits, the design contrast that produces an AUC of 0.99 in one study and 0.77 in another, the gate-1 absence and how far the search behind that claim can be pushed, and the endpoint arithmetic that puts mortality inside a core outcome set rather than at the head of it.
My first attempt at specifying a response protocol produced thresholds, refractory periods, escalation rules and alert routing, which is the model again in different clothes. The alarm policy is the last thing the model does. The response protocol is everything after it, and almost none of it is about the model.
The tenth clinical consensus of the Ibero-American Society of Neonatology describes what actually happens when neonatal sepsis is suspected: blood is drawn, venous access is used to give antibiotics, mother and child are separated, and the hospital stay lengthens; sometimes an X-ray, a urine sample and a lumbar puncture follow (Sola et al. 2020). The consensus puts the proportion of suspected cases that prove to be sepsis at generally under 10%, and at no more than 25–30%.
| Where it comes from | What follows | |
|---|---|---|
| A nurse's concern | Bedside observation, often unarticulated | The same cascade |
| A parent's concern | Bedside observation, from outside the staff | The same cascade |
| A rising apnoea count | Monitor and chart | The same cascade |
| Feeds not tolerated, colour change | Clinical examination | The same cascade |
| An incidental CRP | A sample drawn for something else | The same cascade |
| A prediction model | Continuous physiology, an algorithm | The same cascade |
If the cascade is the same regardless of what opened the door, then the protocol is shared clinical infrastructure that a model plugs into, not a model accessory. It can be specified, and in principle validated, without reference to any particular AI model at all. That is the whole of the argument, and everything below is an attempt to say where the evidence for it is and is not.
Read as a pathway rather than as a reaction to an alarm, the cascade separates. The gates have different costs, and, more usefully for anyone designing a study, they want different things from a test.
| The decision | What it costs | What it wants from a test | |
|---|---|---|---|
| Gate 1 · Look | Something raises concern. Do we go and look? | Attention, an examination, possibly a blood draw in an infant who may weigh 700 g | Sensitivity, because the point is not to miss the infant deteriorating silently |
| Gate 2 · Treat | Having looked, do we culture and start antibiotics? | Intravenous access, broad-spectrum exposure, and a clock that starts | Specificity and timeliness together |
| Gate 3 · Stop | At 36 to 48 hours, culture negative and infant well, do we stop? | The residual risk of stopping, carried by whoever signs it off | Negative predictive value, because this declares it safe to stop |
The output of gate 1 is not antibiotics. It is attention. Collapsing the three into a single "response to the alarm" is what made the protocol look like a specification problem, and separating them is what makes each one answerable by a different kind of study.
A caveat that belongs at the top rather than in a footnote: the evidence below moves between early-onset and late-onset sepsis, because that is how the literature is distributed. The gate structure is the same in both; the trials are not interchangeable, and where a finding is about early-onset sepsis it is labelled as such.
| Best available evidence | State | |
|---|---|---|
| Gate 3 · Stop | NeoPInS, 1,710 neonates, four countries, procalcitonin-guided decision making in suspected early-onset sepsis (Stocker et al. 2017); a systematic review of discontinuation strategies (Feng et al. 2025) | One properly powered trial, and a review concluding there is insufficient evidence to name an optimal strategy |
| Gate 2 · Treat | Biomarker literature, including a systematic review of 171 reports on early-onset sepsis (van Leeuwen et al. 2024) and the late-onset studies below | Large and inconsistent; no stand-alone biomarker established as reliable |
| Gate 1 · Look | — | No confirmatory test evaluated in this population at this gate that I could find |
NeoPInS is worth stating precisely, because it is the strongest study at any of the three gates and it still could not do what it set out to do. Antibiotic duration fell from 65.0 to 55.1 hours on intention to treat. The co-primary non-inferiority endpoint for re-infection or death could not be demonstrated: there were no sepsis-related deaths and only nine infants with possible re-infection in the whole trial (Stocker et al. 2017).
Feng and colleagues reviewed the discontinuation question directly and found 11 randomised trials across nine regimens, ten of them at high risk of bias, and concluded that there is currently insufficient evidence to determine the optimal strategy (Feng et al. 2025). The van Leeuwen review screened 2,296 articles, included 171 in the systematic review and 69 in the meta-analysis, and reports mixed and inconsistent evidence across biomarkers and sample types (van Leeuwen et al. 2024). It is an early-onset review, and it is cited here for the shape of the biomarker evidence rather than as a statement about late-onset practice.
A well-conducted multicentre neonatal trial could not evaluate its own safety endpoint because the events did not happen. That is the same arithmetic that closed Episode 07, arriving from the other direction, and it recurs below when the endpoint question is taken up directly.
Two late-onset studies sit next to each other and tell the story this series has been telling since Episode 03. They use different markers, so this is not a like-for-like comparison and is not offered as one. What travels between them is the design, not the analyte.
| Berka et al. 2021 | Dierikx et al. 2025 | |
|---|---|---|
| Design | Single-centre retrospective case-control | Multicentre prospective, consecutive enrolment |
| Population | 285 very preterm infants, 66 with late-onset bloodstream infection | 63 suspected episodes in 50 infants, 32 classified as late-onset sepsis |
| Marker and threshold | Interleukin-6 above 100 ng/L | Presepsin at initial suspicion, before antibiotics |
| Reported performance | Sensitivity 94%, specificity 99%, AUC 0.988, negative predictive value 99% | AUC 0.77 (95% CI 0.66–0.89) for all cases; 0.80 for culture-positive cases only |
| Authors' conclusion | High negative predictive value could encourage early discontinuation | Insufficient accuracy as a single biomarker to rule out late-onset sepsis |
Retrospective case-control designs produce numbers near 0.99. Prospective consecutive enrolment produces numbers near 0.77. This series documented exactly that pattern across the prediction-model literature on the literature base; the diagnostic literature does the same thing. The prospective study is also small, 63 episodes, which widens its interval and is a second reason not to read the gap as a clean effect of design alone.
Long before any of this, a Dutch group asked whether the bedside examination alone could carry the work. Bekhof and colleagues followed 142 preterm infants under 34 weeks through 187 episodes of suspected infection and built a nomogram from clinical signs, with no laboratory tests in it. Four signs carried it, and the area under the curve was 0.828, 95% CI 0.764–0.892 (Bekhof et al. 2013).
| Signs | |
|---|---|
| Discriminated | Central venous catheter (OR 4.6, 2.2–10.0); increased respiratory support (OR 3.6, 1.9–7.1); grey skin (OR 2.7, 1.4–5.5); capillary refill (OR 2.2, 1.1–4.5) |
| Did not | Temperature instability, apnoea, tachycardia, dyspnoea, hyper- and hypothermia, feeding difficulties, irritability |
Two things follow, and they point in opposite directions. The first is that 0.828 sits inside the 0.78–0.88 band this series has found across every machine-learning group working on this problem. The second is that several of the signs that did not discriminate are precisely what sends a nurse to find a doctor. The gate-1 trigger list in Table 1 is therefore not a list of equally good reasons to go and look, and part of it was measured as noise more than a decade ago.
The comparison with the model literature is suggestive rather than fair. The 0.828 comes from a development sample, without the external validation the prediction models are increasingly held to, and from a cohort in which suspicion had already been raised, which is a different and easier population than a whole unit under continuous monitoring. It is still a figure produced by examining infants, and that deserves more than a footnote.
This is the part that matters most, and it is a structural observation rather than a criticism of any study. Dierikx and colleagues consecutively included infants already started on antibiotics for suspected late-onset sepsis. Berka and colleagues measured markers when infection was clinically suspected. Bekhof and colleagues enrolled episodes of suspected infection. That is a sensible way to study a test that helps you decide whether to continue, and it tells you nothing about gate 1.
Gate 1 is where a model fires. It is the only gate at which a model can do something a clinician has not already done. It is also the gate at which I could not find a confirmatory test evaluated in this population.
Our own unit has counted this in an internal audit that is not published: how often suspicion is raised, what is done when it is, and how often any of it leads anywhere. Those numbers change the shape of the question rather than its direction, and they are held for Episode 09 rather than compressed into a paragraph here, because they need to be set against a comparator and the comparator is what the next episode is about.
Episode 07 closed on mortality as an inherited endpoint. The field has since answered the question formally, and the answer is more useful than either "keep it" or "drop it".
Henry and colleagues reviewed 90 randomised trials in neonatal sepsis and found 88 distinct outcomes, with only 30 of the 90 explicitly stating a primary or secondary outcome, and survival reported in 74% (Henry et al. 2022). That review fed NESCOS, a core outcome set built through a real-time Delphi with 306 participants and an 80% agreement threshold, which reduced 55 candidate outcomes to nine (Taneri et al. 2025).
| Outcomes | |
|---|---|
| Close to the cascade | Escalation of antimicrobial therapy; multiorgan dysfunction; central nervous system infection |
| Downstream | All-cause mortality; need for mechanical ventilation; brain injury on imaging; neurologic status at discharge |
| Long term | Neurodevelopmental impairment; quality of life of parents |
Mortality survived, and that is correct. But it survived as one of nine, alongside escalation of antimicrobial therapy and multiorgan dysfunction, both far more frequent and far closer to the cascade. The field's own consensus already contains an antibiotic-exposure outcome.
The case against mortality as a primary endpoint is arithmetic, not sentiment, and NeoPInS demonstrates it: a trial powered on an outcome that has become rare cannot answer its own question. And the outcome has become rarer. Survival to discharge among extremely preterm infants in the NICHD Neonatal Research Network rose from 76.0% in 2008–2012 to 78.3% in 2013–2018, an adjusted difference of 2.0% with a 95% CI of 1.0–2.9 (Bell et al. 2022).
That is US network data over periods that do not line up exactly with the HeRO trial's enrolment, so it is offered as direction rather than as a matched comparison. The direction is the point. HeRO was powered in an era with more deaths in it, missed its primary endpoint at P = 0.083, and produced a mortality signal as a secondary outcome. Anyone inheriting that endpoint today would need a larger trial to detect a smaller effect. Episode 07 made this claim about falling mortality without a citation behind it; this is the citation.
If the cascade is the intervention, then the endpoints are cascade endpoints, and they arrange themselves against the gates rather than against the model.
| Endpoints | |
|---|---|
| Gate 1 · Look | Workups triggered per infant-week; and how many of them a clinician had not already initiated |
| Gate 2 · Treat | Antibiotic starts per workup; blood cultures per start; time from concern to culture; time from culture to first dose |
| Gate 3 · Stop | Antibiotic days per start; proportion of culture-negative episodes stopped at 36 to 48 hours |
| Across all three | Escalation of antimicrobial therapy, which is in NESCOS; and adherence to the specified protocol, without which a null result cannot be interpreted |
A point about false positives that this series had wrong until now. In a model with the precision documented in Episode 04, the false positives do not, or should not, land as unnecessary antibiotics. They land at gate 1, as unnecessary examinations and unnecessary blood draws in infants who may weigh 700 grams. That is a real cost and it is not the cost this series has been talking about. It is also, unlike antibiotic days, something the cascade could be designed to keep small.
Episode 07 reached the conclusion that the finding from this trial would be inherently unownable, and reached it from the commercial logic: a validated protocol is a guideline, and nobody licenses a guideline. This episode arrives at the same place from the clinical side, which is a stronger position than either argument alone.
If gates 1 to 3 are the same regardless of what opened the door, then a validated cascade works with a nurse's concern, a rising apnoea count, an incidental CRP, our model, or somebody else's. That is precisely why nobody can sell it. It is also why it would still be useful in ten years, when the model that prompted the question has been retrained, replaced or quietly switched off — the half-life problem that Episode 07 raised through the two Nature systems and their measurement problem.
Which raises something worth sitting with rather than resolving here: that the thing worth randomising was never the model. This is offered as a live possibility. It has consequences for the talk in Lyon that have not yet been worked through, and working them through is what the remaining episodes are for.
Two promises from Episode 07 are discharged here and two new ones are made, both to Episode 09. The gate-1 absence is the one item on this list that is neither settled nor promised, and it is the item the argument leans on hardest.
Citations on this page were resolved and checked against PubMed before inclusion. Figures are quoted as reported in the cited abstracts and papers; where a figure comes from a development sample, from a small prospective cohort, or from a study population in which suspicion had already been raised, that is stated in the surrounding text rather than left to the reader. This page discusses and critiques the sources above and does not reproduce them. The three-gate decomposition, the claim that the cascade is trigger-agnostic, the grouping of the NESCOS outcomes in Table 6, and the reading of where false positives land are the author's own, offered as argument rather than as findings of the original investigators.