← Journal

20.96 months against 9.92, from a chart review not a trial

Stanford Medicine's levetiracetam figure comes from a retrospective chart review. Digital Science found 12 paper mill studies cited in clinical guidelines.

Stanford Medicine reported median overall survival of 20.96 months in children with diffuse midline glioma who took levetiracetam, against 9.92 months in those who did not, published in Nature Medicine. It is the largest clinical effect on today's record. It is also not a trial result: the human data are a retrospective chart review, and the association did not hold for gliomas originating elsewhere in the brain.

A doubled survival median is the kind of figure that travels on its own, detached from the sentence that qualifies it. Today's science digest holds several of these. It also holds, on the same day, a measurement of where such figures end up when nobody checks them.

The number and the design are separate facts

  • The New England Journal of Medicine published complete regression of chemotherapy-resistant hepatoblastoma in a 3-year-old boy after two doses of CAR T cell therapy, a single case rather than a trial result.
  • Rockefeller University and its collaborators found that HIV rebounding in the 68 participants of the phase 2 RIO trial had selected for resistance to the antibody 10-1074-LS but not to 3BNC117-LS, a result that complicates dual broadly neutralising antibody strategies rather than validating them.
  • City of Hope tested a pancreatic cancer liquid biopsy in nearly 1,800 patients across the United States, Europe and Asia, to check whether its earlier detection results held up in diverse clinical settings.

Only the last two are the checking step, and neither yields a headline number. The RIO result is a phase 2 trial reporting that a strategy got harder. The City of Hope work is a replication across three continents whose entire purpose is to find out whether an earlier result survives contact with varied clinical settings. This is the slow, expensive half of the pipeline, and it is the half that decides whether the fast half meant anything.

What travels when nothing checks

Digital Science analysts traced 1,975 articles by authors affiliated with a suspected paper mill and found 480 of them cited in patents, 57 in policy documents and 12 in clinical guidelines. Twelve is a small number until you consider what a clinical guideline is for.

The same digest records two cases of the check running late, and costing something to run. Technology, Mind, and Behavior retracted a paper on how generative AI affects confidence in work tasks after its sole author declined to give editors the underlying data, and after Middlesex University stated the research was not approved, conducted or supervised by it. Sage is heading toward an arbitration decision in a suit brought by ten researchers whose three papers on the risks of mifepristone it retracted in 2024 over undeclared conflicts of interest.

Those are the cases where someone eventually looked. The paper mill figures are what the record looks like when nobody does.

The qualifier is the finding

Read the Stanford line again with that in mind. The qualifier is not a hedge appended to the result, it is part of the result: a retrospective review of children given an anti-seizure drug for reasons the charts cannot show is a reason to run a trial, not a reason to act on 20.96 months. The version of that sentence with its second half removed is the version that gets cited in a guideline.

Insilico Medicine published LongevityBench, an open benchmark that scored 18 frontier AI systems from six developers on reasoning across aging biology data, with the Buck Institute, Harvard Medical School and Liquid AI as collaborators. A benchmark is the same instinct moved upstream: an attempt to establish what a claim is worth before it starts travelling.

Built from the digest of 2026-09-18-science.