Episode 08 held our unit's count of what happens at gate 1 until there was something to set it against. This page is the working file behind that comparison: what the published models assume about the moment somebody decides to look, what an internal audit found when we counted it, and what that does to the alarm budget, the early-warning figure and the trial question.
Episode 09 of Road to CEPAS 2026 starts from a thought experiment: a kitchen timer beside every incubator, set for two hours, and a trained person who examines the infant every time it rings. The Substack post argues that routine care already runs this timer, and that whatever a prediction model adds, it has to add on top of it.
This page is the working file: how each paper this series has used anchors the onset of sepsis, and therefore what it can and cannot say about gate 1; the shape of the audit, kept in general terms because it is unpublished; the three consequences for our own 2023 model; the replace-or-add question; where HeRO and Bekhof sit against the timer; and the decision-analytic framing that makes the timer a strategy rather than a metaphor.
Routine neonatal care runs in cycles of a few hours, and every cycle puts hands on the infant: temperature, nappy, position, feed, lines. A model that fires at three in the morning is competing with the three o'clock care round and with whoever is standing at that incubator anyway.
Routine care is part of the comparator arm of any trial of a sepsis alarm. The question that follows is basic, and the literature turns out not to answer it: what does anyone think currently happens at that incubator?
Gate 1 is the moment somebody decides an infant needs looking at. The papers this series has leaned on were re-read for one thing only: where each puts the onset of the episode, and what that anchor lets it say about gate 1.
| Anchor | What it can say about gate 1 | |
|---|---|---|
| van den Berg 2023 (our model) | Blood culture, explicitly standing for first clinical suspicion | Alarm limit of three a day for the unit set from an estimate of routine sepsis checks; never measured. Limitations acknowledge suspicion may precede the culture |
| Griffin & Moorman 2001, 2003 | Abrupt deterioration prompting cultures and antibiotics | Event defined by the clinical action that followed it |
| Moorman 2011 (HeRO trial) | Score displayed; response left to the clinician | No mandated response, no count of how often anyone went to look. Gate 1 visible only downstream, as more work-ups |
| Masino 2019 | Documented sepsis evaluation | Recorded clinical action as anchor |
| Cabrera-Quiros 2021 | CRASH moment: Cultures, Resuscitation, and Antibiotics Started Here | Recorded clinical action as anchor |
| RALIS (Mithal 2018) | Culture sent; alert counted if within 7 days before or after | Window too wide to locate gate 1 |
| Mani 2014 | Infants already evaluated for sepsis | Gate 1 passed before the cohort starts |
| Bekhof 2013 | Episodes in which infection was already suspected | Records what was visible once somebody was concerned; cannot say how often concern arose |
The detail behind two rows. Masino and colleagues took case data from the 44-hour window ending four hours before a sepsis evaluation (Masino et al. 2019). Cabrera-Quiros and colleagues defined a CRASH moment for each infant with culture-proven late-onset sepsis and an equivalent moment in matched controls (Cabrera-Quiros et al. 2021). The RALIS study counted an alert as associated with an episode if it occurred within seven days before or after infection was suspected, with suspicion timed as the moment a culture was sent (Mithal et al. 2018).
Retrospective models need a timestamp, so a recorded clinical action becomes the anchor. Suspicion-cohort studies begin after somebody is already concerned. The one prospective trial added a score to an unmeasured baseline and reported what changed downstream. Models are trained and tested against recorded actions, then discussed as tools for a decision that often happened before those actions were recorded.
We tried to count gate 1, in an internal audit of infants born at 32 weeks or less over one year. It is unpublished, so it is described here in general terms only.
Episodes that end with a blood draw are a selected subset of the moments at which somebody became concerned. The much larger group of bedside evaluations that stop at "looks fine" is hard to distinguish retrospectively from ordinary care, and that is precisely the group the culture-anchored literature cannot see.
| ~3 alarms a day, unit-wide | 1 alarm a day, unit-wide | |
|---|---|---|
| Relation to audit baseline | Above documented evaluation rate | Close to documented evaluation rate |
| Overall alarm performance | Multi-threshold setting behind the headline figure | About half of culture-positive episodes detected within the alarm window; precision about 5%, roughly twenty alarms per alarm associated with sepsis |
| Early warning before culture | About half of infants with sepsis flagged before the culture, derived from the change in cumulative recall between 24 hours before culture and the culture itself | Not published |
None of this contradicts what van den Berg et al. 2023 reported, and the paper states its alarm limit as an estimate. What changes is which operating point a trial should be built around, and the published figure that goes with it. This continues the self-critique of Episode 04 from the side of the comparator rather than the metric.
| Replace | Add | |
|---|---|---|
| What happens | Roughly the same number of looks per day, but the model chooses who | One alarm a day on top of roughly one documented evaluation a day: about double the occasions somebody is sent to look |
| What it has to show | That the model chooses better than the nurse | That the extra looks come early enough, in the right infants, to be worth the attention |
| What we know | Nothing published; the literature does not know whom the nurse would have chosen. Unpublished comparison suggests model and clinicians often choose different infants | The number that would justify it does not yet exist |
The trial question reduces to one sentence: does the model improve a look that was going to happen anyway, or does it create a useful look that otherwise would not have happened?
HeRO is the one randomised trial on the list, and it sits mostly on the first branch. There was no discrete alarm budget: the score was displayed continuously, updated hourly, and available to whoever was already caring for the infant. Mortality fell from 10.2% to 8.1%; blood cultures rose by about ten percent; antibiotic days were five percent higher, not statistically significant (Moorman et al. 2011). In a secondary analysis, mortality in the thirty days after a first episode of proven sepsis was lower when the score was displayed (Fairchild et al. 2013). That is consistent with the number changing what clinicians did once already attending to an infant. Because gate 1 was never counted, whether anyone looked earlier remains an inference.
Bekhof answers a different question. In 187 episodes of already-suspected infection, four readily available variables — increased respiratory support, prolonged capillary refill, grey skin and a central venous catheter — reached an AUC of 0.828 (Bekhof et al. 2013). It is not a ceiling for the timer: the infants had already crossed gate 1, the outcome combined culture-proven and clinical sepsis, and two of the four predictors are clinical context rather than examination findings.
Bekhof shows that once somebody is concerned, ordinary bedside information already carries substantial signal. A model claiming earlier detection has to find information available before that concern exists, in heart rate, saturation or their dynamics, that the routine examination has not yet made visible. That was the original promise of heart-rate-characteristics monitoring, fifteen years on and still resting on one large trial. Bekhof's nomogram never became routine either: the pattern predates machine learning.
If the intervention under consideration is "go and examine the infant", the timer is the analogue of treat-all in decision-curve terms: examine everyone at the chosen interval. A model has clinical utility only if using it to decide whom to examine produces more net benefit than that default, or than examining nobody outside routine care. For the methodological background on why discrimination alone does not establish usefulness, see the review by Van Calster et al. 2026, which the post cites at this point.
No formal decision curve is drawn for our model here, because the individual predictions from the published analysis are not readily available in a form that makes one useful. The timer makes the same problem visible without the graph.
| What to measure | |
|---|---|
| Workload | Examinations per infant-week |
| Yield | Deterioration missed per nurse-hour |
| Comparison | Other ways of spending the same attention |
| One such alternative | A structured examination during routine care, using a small set of predefined signs. Not glamorous; very cheap to test |
This refines the gate-1 row of Table 7 on the Episode 08 page: "workups triggered per infant-week" now has a measured baseline beside it, and "how many a clinician had not already initiated" is the replace-or-add question in numbers.
One promise from Episode 08 is discharged and one is not. The argument now rests on two unpublished pieces of internal work, which is honest to say on a public page and a constraint on what can be put on a slide.
Citations 1–10 were resolved against PubMed before inclusion; the anchor definitions in Table 1 were checked against the abstracts and, for RALIS, the full text. Citation 11 is a statistics journal outside PubMed and was resolved through Crossref. HeRO figures are as reported in the trial and its secondary analysis. Figures for our own model are as reported in the 2023 paper. The internal audit and the alarm-versus-evaluation comparison are unpublished and appear only in general terms. The timer, the gate structure, the replace-or-add framing and the reading of the timer as a treat-all strategy are the author's own, offered as argument rather than as findings of the original investigators.