Episode 05 ended on the rung a retrospective model cannot climb by itself: does acting on the alarm leave the baby better off? That is a trial question, and in this corner of the literature it has been asked properly exactly once. This page is the working file behind Episode 06 — how the 2011 trial was designed, what it did and did not measure, and what a genuine impact trial of a model like ours would cost.
Episode 06 of Road to CEPAS 2026 follows the thread Episode 05 left hanging. The Substack post argues that the distance between a good model and a proven clinical pathway is measured in years and thousands of patients, and that the one trial the field leans on is both more important and more fragile than its citation count suggests.
This page is the working file: the three-stage validation ladder set out explicitly, the design and outcomes of the HeRO trial laid side by side, the counterfactual problem stated in full, the design requirements a trial of a model like Berg 2023 would inherit, and the search receipt behind the claim that fifteen years on there is still only one.
The discussion in van den Berg et al. 2023 sketches the road a prediction model has to travel before it earns a place at the bedside. It is worth setting out as a ladder, because the rung a study sits on determines what it is entitled to claim.
| The question it answers | Care changed? | |
|---|---|---|
| Retrospective development | Would the alarm have fired before the clinician noticed, replaying saved data minute by minute? | No — nothing was done differently |
| Prospective silent validation | Does the score stay accurate on new patients, while clinicians cannot yet see it? | No — deliberately masked |
| Clinical-impact trial | Does displaying the score, or acting on it, change what happens to the baby? | Yes — this is the point |
Our published model sits on the first rung. A retrospective model is a simulation: it looks back over years of saved oxygen-saturation and heart-rate data and asks how often it would have been early. The answer is known. But a replay is not a trial, and no baby was treated differently because the number went up.
Between April 2004 and May 2010, across nine American NICUs, 3,003 very low birth weight infants were randomised to have a continuously updated heart-rate risk score either displayed at the bedside or recorded but hidden (Moorman et al., 2011). The score was the HeRO index — a number expressing the fold-increase in the risk of sepsis over the next 24 hours, derived from reduced heart-rate variability and transient decelerations that tend to appear before a preterm infant looks clinically septic.
It is not our model. It is the same species of thing: a continuous physiological risk score, updated hourly, sitting on a monitor next to a sick baby.
| Detail | |
|---|---|
| Period | April 2004 – May 2010 |
| Centres | Nine US NICUs |
| Participants | 3,003 very low birth weight infants |
| Randomisation | Individual infants; score displayed at the bedside vs recorded but masked |
| Intervention | Display of the score. No specific clinical response required by protocol |
The result everyone remembers is a mortality reduction. The result the trial was built to measure is a different one, and it did not reach significance. Both are true, and the gap between them is the most instructive thing in the paper.
| Outcome | Result | |
|---|---|---|
| Primary | Days alive and off the ventilator in the 120 days after randomisation | +2.3 days in the displayed arm, P = .08 — not significant |
| Secondary | All-cause mortality | 8.1% vs 10.2%; HR 0.78, P = .04; NNM 48 |
| Pre-specified subgroup | All-cause mortality, infants below 1,000 g | 13.2% vs 17.6%; NNM 23 |
Read strictly, the canonical trial in this field missed on the outcome it was built to measure and hit on one listed beneath. Two readings stay open at once. A composite can obscure a real mortality effect if its components move in different directions, the ventilator half adding noise that drags the whole below the threshold. Equally, a significant secondary finding after a negative primary endpoint can itself be a chance result, and mortality was one of several secondaries. The confidence interval for the mortality benefit only just excluded no effect.
A Dutch systematic review weighing exactly this rated the trial's risk of bias high — on performance and detection bias, and on the absence of correction for multiple testing — and concluded that the evidence does not justify putting heart-rate-characteristic monitoring into routine care without a further trial (Koppens et al., 2023).
None of this makes the finding worthless. It makes it what it is: clinically important, internally coherent, and statistically less secure than its prominence in the later literature suggests. The point that survives either reading is the design lesson — how much rides on a decision made before a single baby is enrolled, which endpoint you name as primary, and how hard that is to know in advance.
The trial deliberately did not tell anyone what to do. Clinicians were taught what the score meant and told that a rising number should prompt them to look at the baby and consider tests, but no specific action was required by the protocol. The authors themselves call this a debatable weakness. It is the exact gap this series is about: the trial tested the display of a number, not a defined response to it. If the mortality effect was causal at all, it operated through nine units' worth of clinicians each responding in their own way.
And the response cannot be cleanly measured. To know whether the alarm helped, you would want to know how much earlier sepsis was caught — but the biological moment a preterm infant's sepsis begins is unobservable. There is no clock that starts. Time-to-onset, the thing you would most want to anchor to, cannot be measured directly, and the trial says as much: the mechanism of the benefit could be inferred, not proven.
What the 2011 team could show is indirect, and it is worth laying out, because it is the texture a real impact trial lives in.
| Finding | What it does not settle | |
|---|---|---|
| Sepsis detected | 358 affected infants displayed vs 379 control — no meaningful difference | Displaying the score did not catch more sepsis |
| Mortality after proven sepsis | 10.0% vs 16.1% in the 30 days following | Whether clinicians acted sooner, or something else differed |
| Diagnostic cost | Slightly more blood cultures, of borderline significance | The size of the burden, given the borderline estimate |
| Treatment cost | Slightly more antibiotic days, not statistically significant | Whether there is any real antibiotic penalty at all |
"Act sooner on the infections already there" is as far as the strongest trial in the field can take the claim. That is not a failure of the trial. It is the limit of what a trial of this design can prove — and the same limit any future study, including a study of our model, would run into.
A new trial still cannot see the true onset of sepsis. It can, however, measure the operational chain the alert is meant to shorten: time from alert to bedside assessment, to blood culture, to antibiotics; physiological deterioration in the hours after an alert; escalation to organ support; adherence to the specified response protocol. None of that pins down biological onset. All of it would make the mechanism far less opaque than it was in 2011.
Lay the 2011 experience against our 2023 paper and the shape of the required study comes into focus. It is sobering.
None of this is a reason not to do it. It is a description of what doing it properly costs — and of why the distance between a good model and a proven pathway is measured in years and thousands of patients rather than in another analysis of the data already in hand.
Tethering back to Episode 05's regulatory thread: the device entered the trial already 510(k)-cleared — for its intended use as an ECG-derived measure of reduced heart-rate variability and transient decelerations. Not for diagnosing sepsis. Not for improving survival. That clearance rested, as the 510(k) route does, on substantial equivalence to a legally marketed predicate device and on adequate safety and performance within the stated intended use.
Regulatory clearance and clinical benefit are different claims, established by different means. The randomised trial addressed a question the clearance never touched — whether showing the number to a clinician at three in the morning changes the night for the better. Only the benefit claim necessarily calls for a clinical-impact trial, though what regulators demand in practice varies with the device, its intended use, and the jurisdiction.
When the smallest survivors were followed up at 18 to 22 months, the composite outcome of death or neurodevelopmental impairment was not significantly reduced — roughly 39% in the displayed group against 44% in controls, a relative risk near 0.87 that did not clear significance (Schelonka et al., 2020).
Two caveats carry weight. The outcome was available for only about 72% of eligible infants, and attrition on that scale limits what can be concluded. And the analysis is of a composite of death or impairment, not of the survivors themselves; asking whether the extra survivors were better off means conditioning on survival, which brings its own selection problem. What can fairly be said is narrower: the follow-up did not show that the mortality benefit was matched by an improvement on the death-or-impairment composite. A later analysis did find a benefit in the narrower subgroup of extremely preterm infants who developed sepsis (King et al., 2021) — a real but subgroup-level signal that a single trial cannot settle.
The question recurses. A study suggesting that displaying the score saves lives raises a further one it was never built to answer: did the improved survival come with improved, or at least not worse, neurodevelopmental outcomes?
Before leaning on the word "only," it is worth showing the work — and others have looked too.
It stands alone not because the question stopped mattering — the same review calls plainly for a large international randomised trial — but because a trial like this is enormous, slow and hard to fund, and because a result positive enough to be quoted let the field feel the question was answered when it was answered only halfway.
The number can already tell you the baby's risk has risen. One large randomised trial suggests that displaying this particular score may reduce mortality — a real and hard-won signal, but a single fragile result, not a settled clinical pathway. What no trial has yet shown, for HeRO and still less for a single-centre model built on a replay, is which move at the bedside turns a rising number into a better morning. With the trial mapped, Episode 07 asks who would ever run it, and what building toward a validated response protocol means when the field's incentives reward the next model over the missing trial.
This page discusses and critiques the sources above; it does not reproduce them. Reported figures are paraphrased and pointed back to the primary source. Readings of the trial's endpoint structure and of the strength of the mortality signal are the author's own, offered as interpretation rather than as findings of the original investigators.