— Episode 09 · The baseline a model competes with
Gate 1 · counted once · set against the literature and the timer

The comparator.

Episode 08 held our unit's count of what happens at gate 1 until there was something to set it against. This page is the working file behind that comparison: what the published models assume about the moment somebody decides to look, what an internal audit found when we counted it, and what that does to the alarm budget, the early-warning figure and the trial question.

Why this page exists.

Episode 09 of Road to CEPAS 2026 starts from a thought experiment: a kitchen timer beside every incubator, set for two hours, and a trained person who examines the infant every time it rings. The Substack post argues that routine care already runs this timer, and that whatever a prediction model adds, it has to add on top of it.

This page is the working file: how each paper this series has used anchors the onset of sepsis, and therefore what it can and cannot say about gate 1; the shape of the audit, kept in general terms because it is unpublished; the three consequences for our own 2023 model; the replace-or-add question; where HeRO and Bekhof sit against the timer; and the decision-analytic framing that makes the timer a strategy rather than a metaphor.

A note on what is mine and what is read. The timer, the three gates and the decision to treat the audit as the baseline rather than as one more set of findings are arguments, not findings, and no cited author is responsible for them. The audit and the comparison between alarms and documented evaluations are unpublished internal work; they are described here in general terms only, and no figure from them should be read as a published estimate.

The timer

The comparator arm already exists, whether the protocol names it or not.

Routine neonatal care runs in cycles of a few hours, and every cycle puts hands on the infant: temperature, nappy, position, feed, lines. A model that fires at three in the morning is competing with the three o'clock care round and with whoever is standing at that incubator anyway.

Takeaway

Routine care is part of the comparator arm of any trial of a sepsis alarm. The question that follows is basic, and the literature turns out not to answer it: what does anyone think currently happens at that incubator?


Gate 1 in the literature

Every paper anchors to something recorded. None counts the moment before.

Gate 1 is the moment somebody decides an infant needs looking at. The papers this series has leaned on were re-read for one thing only: where each puts the onset of the episode, and what that anchor lets it say about gate 1.

Table 1. Where each study anchors the episode, and what that means for gate 1. The right-hand column is the finding: none of these counts gate 1 as part of an ordinary baseline before an alarm exists.
  Anchor What it can say about gate 1
van den Berg 2023 (our model) Blood culture, explicitly standing for first clinical suspicion Alarm limit of three a day for the unit set from an estimate of routine sepsis checks; never measured. Limitations acknowledge suspicion may precede the culture
Griffin & Moorman 2001, 2003 Abrupt deterioration prompting cultures and antibiotics Event defined by the clinical action that followed it
Moorman 2011 (HeRO trial) Score displayed; response left to the clinician No mandated response, no count of how often anyone went to look. Gate 1 visible only downstream, as more work-ups
Masino 2019 Documented sepsis evaluation Recorded clinical action as anchor
Cabrera-Quiros 2021 CRASH moment: Cultures, Resuscitation, and Antibiotics Started Here Recorded clinical action as anchor
RALIS (Mithal 2018) Culture sent; alert counted if within 7 days before or after Window too wide to locate gate 1
Mani 2014 Infants already evaluated for sepsis Gate 1 passed before the cohort starts
Bekhof 2013 Episodes in which infection was already suspected Records what was visible once somebody was concerned; cannot say how often concern arose

The detail behind two rows. Masino and colleagues took case data from the 44-hour window ending four hours before a sepsis evaluation (Masino et al. 2019). Cabrera-Quiros and colleagues defined a CRASH moment for each infant with culture-proven late-onset sepsis and an equivalent moment in matched controls (Cabrera-Quiros et al. 2021). The RALIS study counted an alert as associated with an episode if it occurred within seven days before or after infection was suspected, with suspicion timed as the moment a culture was sent (Mithal et al. 2018).

Takeaway

Retrospective models need a timestamp, so a recorded clinical action becomes the anchor. Suspicion-cohort studies begin after somebody is already concerned. The one prospective trial added a score to an unmeasured baseline and reported what changed downstream. Models are trained and tested against recorded actions, then discussed as tools for a decision that often happened before those actions were recorded.


The audit

Gate 1 is frequent, cheap, and usually ends with a look.

We tried to count gate 1, in an internal audit of infants born at 32 weeks or less over one year. It is unpublished, so it is described here in general terms only.

On what the audit does and does not settle. It comes from the same unit and population as the 2023 paper, in a more recent year. The definitions are not identical: considering sepsis during a round is broader than an evaluation that gets written down. The audit does not settle which estimate is right. It does mean the paper's sentence that a bedside visit at every alarm would not exceed current practice no longer fully survives.
Takeaway: what this does to training data

Episodes that end with a blood draw are a selected subset of the moments at which somebody became concerned. The much larger group of bedside evaluations that stop at "looks fine" is hard to distinguish retrospectively from ordinary care, and that is precisely the group the culture-anchored literature cannot see.


The model

Three consequences for our own paper.

Table 2. The two operating points that matter, read against the audit. The published early-warning figure belongs to the first row; the row nearest the measured baseline has none.
  ~3 alarms a day, unit-wide 1 alarm a day, unit-wide
Relation to audit baseline Above documented evaluation rate Close to documented evaluation rate
Overall alarm performance Multi-threshold setting behind the headline figure About half of culture-positive episodes detected within the alarm window; precision about 5%, roughly twenty alarms per alarm associated with sepsis
Early warning before culture About half of infants with sepsis flagged before the culture, derived from the change in cumulative recall between 24 hours before culture and the culture itself Not published
Takeaway: a correction of emphasis, not of the paper

None of this contradicts what van den Berg et al. 2023 reported, and the paper states its alarm limit as an estimate. What changes is which operating point a trial should be built around, and the published figure that goes with it. This continues the self-critique of Episode 04 from the side of the comparator rather than the metric.


The trial question

Replace or add.

Table 3. The two things a model could do to a unit's gate 1. The limited overlap between alarms and existing evaluations makes the second the more plausible reading.
  Replace Add
What happens Roughly the same number of looks per day, but the model chooses who One alarm a day on top of roughly one documented evaluation a day: about double the occasions somebody is sent to look
What it has to show That the model chooses better than the nurse That the extra looks come early enough, in the right infants, to be worth the attention
What we know Nothing published; the literature does not know whom the nurse would have chosen. Unpublished comparison suggests model and clinicians often choose different infants The number that would justify it does not yet exist
Takeaway

The trial question reduces to one sentence: does the model improve a look that was going to happen anyway, or does it create a useful look that otherwise would not have happened?


HeRO and Bekhof

A number beside the cot is not an alarm that sends someone to it.

HeRO is the one randomised trial on the list, and it sits mostly on the first branch. There was no discrete alarm budget: the score was displayed continuously, updated hourly, and available to whoever was already caring for the infant. Mortality fell from 10.2% to 8.1%; blood cultures rose by about ten percent; antibiotic days were five percent higher, not statistically significant (Moorman et al. 2011). In a secondary analysis, mortality in the thirty days after a first episode of proven sepsis was lower when the score was displayed (Fairchild et al. 2013). That is consistent with the number changing what clinicians did once already attending to an infant. Because gate 1 was never counted, whether anyone looked earlier remains an inference.

Bekhof answers a different question. In 187 episodes of already-suspected infection, four readily available variables — increased respiratory support, prolonged capillary refill, grey skin and a central venous catheter — reached an AUC of 0.828 (Bekhof et al. 2013). It is not a ceiling for the timer: the infants had already crossed gate 1, the outcome combined culture-proven and clinical sepsis, and two of the four predictors are clinical context rather than examination findings.

Takeaway: what "subclinical" has to mean

Bekhof shows that once somebody is concerned, ordinary bedside information already carries substantial signal. A model claiming earlier detection has to find information available before that concern exists, in heart rate, saturation or their dynamics, that the routine examination has not yet made visible. That was the original promise of heart-rate-characteristics monitoring, fifteen years on and still resting on one large trial. Bekhof's nomogram never became routine either: the pattern predates machine learning.


The decision analysis

The timer is the treat-all strategy.

If the intervention under consideration is "go and examine the infant", the timer is the analogue of treat-all in decision-curve terms: examine everyone at the chosen interval. A model has clinical utility only if using it to decide whom to examine produces more net benefit than that default, or than examining nobody outside routine care. For the methodological background on why discrimination alone does not establish usefulness, see the review by Van Calster et al. 2026, which the post cites at this point.

No formal decision curve is drawn for our model here, because the individual predictions from the published analysis are not readily available in a form that makes one useful. The timer makes the same problem visible without the graph.

Table 4. If the model's advantage is in choosing rather than seeing, it is an attention-allocation tool and should be measured as one. The last row is the cheap alternative.
  What to measure
Workload Examinations per infant-week
Yield Deterioration missed per nurse-hour
Comparison Other ways of spending the same attention
One such alternative A structured examination during routine care, using a small set of predefined signs. Not glamorous; very cheap to test
Takeaway

This refines the gate-1 row of Table 7 on the Episode 08 page: "workups triggered per infant-week" now has a measured baseline beside it, and "how many a clinician had not already initiated" is the replace-or-add question in numbers.


The ledger

What this episode settled, and what it still owes.

Takeaway

One promise from Episode 08 is discharged and one is not. The argument now rests on two unpublished pieces of internal work, which is honest to say on a public page and a constraint on what can be put on a slide.


References & sourcing.

  1. van den Berg M, Medina O, Loohuis I, et al. Development and clinical impact assessment of a machine-learning model for early prediction of late-onset sepsis. Comput Biol Med 2023;163:107156. doi:10.1016/j.compbiomed.2023.107156  Open access · CC BY-NC-ND
  2. Griffin MP, Moorman JR. Toward the early diagnosis of neonatal sepsis and sepsis-like illness using novel heart rate analysis. Pediatrics 2001;107(1):97–104. doi:10.1542/peds.107.1.97  Deterioration-defined event
  3. Griffin MP, O'Shea TM, Bissonette EA, Harrell FE, Lake DE, Moorman JR. Abnormal heart rate characteristics preceding neonatal sepsis and sepsis-like illness. Pediatr Res 2003;53(6):920–926. doi:10.1203/01.PDR.0000064904.05313.D2  Two NICUs
  4. Moorman JR, Carlo WA, Kattwinkel J, et al. Mortality reduction by heart rate characteristic monitoring in very low birth weight neonates: a randomized trial. J Pediatr 2011;159(6):900–906.e1. doi:10.1016/j.jpeds.2011.06.044  Open access · PMC
  5. Fairchild KD, Schelonka RL, Kaufman DA, et al. Septicemia mortality reduction in neonates in a heart rate characteristics monitoring trial. Pediatr Res 2013;74(5):570–575. doi:10.1038/pr.2013.136  Secondary analysis · n = 2,989
  6. Masino AJ, Harris MC, Forsyth D, et al. Machine learning models for early sepsis recognition in the neonatal intensive care unit using readily available electronic health record data. PLoS One 2019;14(2):e0212665. doi:10.1371/journal.pone.0212665  Open access · evaluation anchor
  7. Cabrera-Quiros L, Kommers D, Wolvers MK, et al. Prediction of late-onset sepsis in preterm infants using monitoring signals and machine learning. Crit Care Explor 2021;3(1):e0302. doi:10.1097/CCE.0000000000000302  Open access · CRASH anchor
  8. Mithal LB, Yogev R, Palac HL, Kaminsky D, Gur I, Mestan KK. Vital signs analysis algorithm detects inflammatory response in premature infants with late onset sepsis and necrotizing enterocolitis. Early Hum Dev 2018;117:83–89. doi:10.1016/j.earlhumdev.2018.01.008  RALIS · ±7-day window
  9. Mani S, Ozdas A, Aliferis C, et al. Medical decision support using machine learning for early detection of late-onset neonatal sepsis. J Am Med Inform Assoc 2014;21(2):326–336. doi:10.1136/amiajnl-2013-001854  Already under evaluation
  10. Bekhof J, Reitsma JB, Kok JH, Van Straaten IHLM. Clinical signs to identify late-onset sepsis in preterm infants. Eur J Pediatr 2013;172(4):501–508. doi:10.1007/s00431-012-1910-6  Development sample · AUC 0.828
  11. Van Calster B, van Smeden M, van Amsterdam W, Coemans M, Wynants L, Steyerberg EW. The enemies of reliable and useful clinical prediction models: a review of statistical and scientific challenges. Annu Rev Stat Appl 2026;13:465–492. doi:10.1146/annurev-statistics-042324-123749  Crossref · not in PubMed

Citations 1–10 were resolved against PubMed before inclusion; the anchor definitions in Table 1 were checked against the abstracts and, for RALIS, the full text. Citation 11 is a statistics journal outside PubMed and was resolved through Crossref. HeRO figures are as reported in the trial and its secondary analysis. Figures for our own model are as reported in the 2023 paper. The internal audit and the alarm-versus-evaluation comparison are unpublished and appear only in general terms. The timer, the gate structure, the replace-or-add framing and the reading of the timer as a treat-all strategy are the author's own, offered as argument rather than as findings of the original investigators.