, Episode 07 · Incentives and ownership
Nothing forbids it · nothing rewards it · nobody owns it

Who would run it.

Episode 06 costed the trial. This page is the working file behind the question that follows, which turns out not to have a person or an institution as its answer. The regulatory route exists, the model has an owner, and the capability is not in doubt. What is missing is that the trial is nobody's job, and the reasons for that are structural rather than personal.

Why this page exists.

Episode 07 of Road to CEPAS 2026 picks up where Episode 06 stopped. The Substack post argues that the absence of a second impact trial in this field is not a failure of nerve but an accurate response to a set of incentives, and that the thing most worth proving is precisely the thing nobody can own.

This page is the working file: the sequence that put the HeRO device on the market before the trial began, the asymmetry between what a model publication costs and what a trial costs, two much better-funded groups stopping at the same rung we stop at, the enrolment arithmetic set against two Dutch-Belgian trials, and the argument that a validated response protocol is a public good.

A note on register. This page describes a set of incentives, not a set of villains. Every group named here did work that advanced the field, and the 2011 team did the hardest version of it first. Where a claim is an interpretation rather than a reported finding, it is marked as one. Where a figure comes from an abstract rather than from the full text, it is not restated as though it were verified.

The sequence

The market did not require the trial.

The assumption worth discarding first is that the 2011 randomised trial existed because somebody needed it in order to sell something. The order of events runs the other way. The HeRO monitoring system already held 510(k) clearance when the trial began, for reporting heart rate characteristics as a measure of variability and transient decelerations, and the device was supplied by its manufacturer for use in the study (Moorman et al. 2011).

The commercial pathway was therefore already complete. A cleared device had reached the bedside by a route that never asked anyone to demonstrate that displaying the number changed anything for a baby. The trial was surplus to that route. It ran because a group wanted to know.

Table 1. What each claim actually required. Clearance and clinical benefit are established by different means; only the third row obliges anyone to run a trial.
  What it needed Trial required?
Reaching the market Substantial equivalence to a predicate device, within a stated intended use No, settled in 2004
Displaying the score Clearance plus a clinician willing to read it No
Claiming patient benefit Prospective comparative evidence against a named endpoint Yes, and it has been done once
Takeaway

The regulatory bar in Europe is demanding, but it is not a bar that requires proof of clinical benefit either. A trial of the alarm is not what stands between a score and a bedside. This removes the most comfortable explanation for the missing second trial, which is that someone is being prevented from running it.


The reward

Doing it first, and being remembered for the wrong line.

The trial was powered to detect a two-day difference in days alive and off the ventilator, which is what set the sample size near three thousand. It found 2.3 days at P = 0.083. The primary endpoint did not reach significance. The result the field remembers, the reduction in all-cause mortality from 10.2% to 8.1%, was a secondary outcome, and the Dutch systematic review flagged the absence of correction for multiple testing among its reasons for rating the trial at high risk of bias (Koppens et al. 2023). The endpoint structure is set out in full on the Episode 06 page.

Read as a career, rather than as a result: a group did the hardest thing available in this field, did it first, did it at a scale nobody has matched in the fifteen years since, and the headline was a miss. The trial is now cited constantly, most often for the line beneath the one it was built to test.

Takeaway: an interpretation, not a finding

For anyone deciding where to spend six years of a working life, that is the worked example in front of them. The observation is offered as a reading of the incentives, not as a claim about what any individual investigator concluded.


The asymmetry

What the field pays for instead.

The two options are not close to each other in cost, in risk, or in the probability of producing something publishable.

Table 2. A new model against a prospective trial, on the dimensions that decide which one gets done. The AUC band is the convergence documented across Branch A on the literature base.
  Another model A prospective trial
Time to output About a year Six years, on the one precedent available
Data Already collected Thousands of infants, prospectively enrolled
Probability of a publishable result High, AUCs in this literature land in a narrow band Uncertain, with a plausible null
Exposure No hypothesis named in advance A primary endpoint fixed before enrolment
Takeaway

The field is not choosing models over trials out of timidity. It is responding accurately to what it is funded and promoted for. Treating this as a failure of individual will misdiagnoses it, and misdiagnosis leads to the wrong remedy.


The control case

Two of the best-funded groups in medical AI stop at the same rung.

If money were the binding constraint, the constraint should relax where the money is abundant. Two systems published in Nature in the same issue suggest it does not.

MIRA is an autonomous agent operating inside a sandboxed electronic health record, able to take histories, order and interpret laboratory, imaging and microbiology tests, generate differential diagnoses and formulate treatment plans including prescriptions and admissions. On simulations built from real patient cases it outperformed physicians on diagnostic accuracy while making guideline-concordant and medication-safe decisions (Ferber et al. 2026). The paper closes by stating that generalisation, safety and governance still need to be established through prospective, real-world studies.

AMIE was compared with 21 primary care physicians across 100 multivisit case scenarios in a randomised, blinded virtual OSCE. It was non-inferior on management reasoning as assessed by specialists, and scored better on the preciseness of the treatment and investigation it proposed and on its grounding in clinical guidelines (Liévin et al. 2026). The paper likewise notes that further research is needed before real-world translation.

The accompanying commentary names the underlying difficulty as a measurement problem: the instruments used to evaluate these systems are changing faster than the studies that use them (Nature 2026). Base models improve every few months, which makes it hard to say where a system's capability came from, and a result about one version judged one way may not survive the next release. Both papers show scaffolding that helped an older model doing little for a newer one.

Takeaway

Applied to us, the same logic bites harder. A trial of a 2023 model begun today reports in 2032, by which time the model may have been superseded several times over. Everyone contemplating such a trial already knows this, and it is a rational reason to wait rather than a lazy one. It is also not a reason that improves with time.


The scale

The enrolment arithmetic, without the excuses.

Two recent trials in this region give an objective sense of the distance, and both are properly conducted studies that answered real questions. The comparison is about what a mortality endpoint in this population demands, not about any country's capacity.

Table 3. Enrolment and duration for three neonatal randomised trials. Figures as reported in the cited publications.
  Infants randomised Setting and period
BeNeDuctus 273 Netherlands and Belgium; registered 2015, published end of 2022 (Hundscheid et al. 2023)
STOP-BPD 372 19 NICUs, Netherlands and Belgium; enrolling November 2011 to December 2016 (Onland et al. 2019)
HeRO 3,003 Nine US NICUs; April 2004 to May 2010, screening nearly six thousand
Takeaway

An order of magnitude separates the first two from the third. That gap is set by the endpoint, not by ambition or by geography.


The commercial logic

Closed both ways.

A trial on that scale would cost millions, so it is worth asking who recovers the money. There are two candidate things it could be testing, and neither route pays.

Table 4. The two things a trial could test, against whether a trial is needed and whether the result can be owned. The second row is the argument this series has been making since Episode 01.
  Needs a trial? Can it be sold?
The alarm No, clearance does not ask for one Yes, as a device
The response to the alarm Yes, nothing else can establish it No, a validated protocol is a guideline

A validated response protocol gets published, and then it gets copied, and nobody pays a licence fee to follow a guideline. Worse, from the point of view of whoever funded the trial, a protocol that works on our score would very likely work on a competitor's score, or on thresholds applied to raw vital signs with no model underneath at all. The funder would be handing the result to the field.

Takeaway

This is the structure of a public good, and private actors underinvest in public goods reliably, not out of malice but because the benefit escapes them. It follows that this trial was never going to arrive through a commercial route, and that waiting for industry to fund it is waiting for something the incentives do not produce.


The endpoint

Mortality is an inheritance, not a choice.

Mortality was never selected as the endpoint for heart rate characteristics monitoring. It was a secondary outcome that cleared significance narrowly, in a trial whose primary outcome did not, and without correction for multiple testing. Fifteen years on, the field treats it as the benchmark a new trial would have to meet.

That creates a bind with no comfortable exit. The only effect estimate available for a sample size calculation is the one from 2011, and effect sizes that just cross the significance threshold tend to be overstated, so powering on it risks a study underpowered for the true effect. Powering instead for something smaller and more plausible pushes the sample size past anything realistically assembled.

Takeaway: and the turn to Episode 08

The endpoint, not the ambition or the funding climate, is what makes the trial impossible as currently conceived. An outcome nobody chose on purpose is setting the size of a study nobody can run. Episode 08 takes up what should replace it, and what a response protocol has to specify before it can be the thing under test.


The honest ledger

Six obstacles, which multiply rather than add.

Nothing forbids this trial, nothing rewards it, and nobody is responsible for it. Those are three separate problems, and the value of writing the series has been learning to see them separately. The same applies to the obstacles themselves, which had been sitting in one undifferentiated sense that the thing is too hard.

Takeaway

Any one of these would justify hesitation. Together they multiply. Separating them is what makes some of them addressable, and makes clear that several are not mine to solve. That is a more useful position than one undifferentiated obstacle, and it is still a long way short of a plan.

On open release. The model was made publicly available after publication, which was the right decision. It also illustrates the ownership problem in its sharpest form: publishing a model puts it where anyone may use it and nobody must prove it. The obligation to test the thing does not travel with the download. That is a description of what happens to accountability when the next step costs six years, not an argument against open science.

References & sourcing.

  1. Moorman JR, Carlo WA, Kattwinkel J, et al. Mortality reduction by heart rate characteristic monitoring in very low birth weight neonates: a randomized trial. J Pediatr 2011;159(6):900–906.e1. doi:10.1016/j.jpeds.2011.06.044  Open access · PMC
  2. Koppens HJ, Onland W, Visser DH, Denswil NP, van Kaam AH, Lutterman CA. Heart rate characteristics monitoring for late-onset sepsis in preterm infants: a systematic review. Neonatology 2023;120(5):548–557. doi:10.1159/000531118  PMC10614451
  3. Hundscheid T, Onland W, Kooi EMW, et al. Expectant management or early ibuprofen for patent ductus arteriosus. N Engl J Med 2023;388(11):980–990. doi:10.1056/NEJMoa2207418  BeNeDuctus · n = 273
  4. Onland W, Cools F, Kroon A, et al. Effect of hydrocortisone therapy initiated 7 to 14 days after birth on mortality or bronchopulmonary dysplasia among very preterm infants receiving mechanical ventilation: a randomized clinical trial. JAMA 2019;321(4):354–363. doi:10.1001/jama.2018.21443  STOP-BPD · n = 372
  5. Ferber D, Hilgers L, Höper C, et al. Towards autonomous medical artificial intelligence agents. Nature 2026;655(8125):1282–1291. doi:10.1038/s41586-026-10675-5  Paraphrased · no licence recorded
  6. Liévin V, Palepu A, Weng WH, et al. Towards conversational artificial intelligence for disease management. Nature 2026;655(8125):1292–1299. doi:10.1038/s41586-026-10764-5  Paraphrased · no licence recorded
  7. Medical AI has a measurement problem. Nature 2026 (news and views, accompanying the two papers above). doi:10.1038/d41586-026-02125-z  Paraphrased throughout
  8. Peng Z, Schouten JS, Silvertand D, Long X, Lake DE, Taal HR, Niemarkt HJ, Andriessen P, Sullivan B, van Pul C. External validation complexities: a comparative study of late-onset sepsis prediction models across multiple clinical environments. IEEE Trans Biomed Eng 2025. doi:10.1109/TBME.2025.3618080  IEEE · all rights reserved · paraphrase only
  9. van den Berg M, Medina O, Loohuis I, et al. Development and clinical impact assessment of a machine-learning model for early prediction of late-onset sepsis. Comput Biol Med 2023;163:107156. doi:10.1016/j.compbiomed.2023.107156  Open access · CC BY-NC-ND

Citations on this page were resolved and checked against PubMed before inclusion; the IEEE reference, which PubMed does not index, was verified against the publisher record. This page discusses and critiques the sources above and does not reproduce them. The two Nature papers and the accompanying commentary carry no open licence recorded in PubMed and are paraphrased throughout, with no figures reproduced. Readings of the incentive structure, of the commercial logic, and of mortality as an inherited endpoint are the author's own, offered as interpretation rather than as findings of the original investigators.