Researchers from Boston Children’s Hospital’s Manton Center for Orphan Disease Research, Harvard, and OpenAI published a study in NEJM AI on Wednesday in which the OpenAI o3 Deep Research reasoning model was given access to de-identified clinical and genomic information from 376 rare-disease cases that had already been worked up by human specialists and remained unsolved. After expert review of the model’s outputs, additional confirmatory testing, and standard clinical workflow, physicians established new diagnoses in 18 of those cases. The reported additional diagnostic yield is 4.8 percent on top of what expert analysis had already extracted.
The headline framing of “AI beats doctors” is the wrong frame here and the paper is careful about saying so. The cases in this cohort were not unsolved because nobody had looked at them. They were unsolved because the world’s best pediatric-genetics teams had spent meaningful hours per case and reached an honest dead end. What o3 Deep Research did was synthesize across the medical literature, the clinical findings, and the variant data at a breadth and recall that a human team cannot sustain across hundreds of patients, and surface candidate phenotype-genotype matches that the original workups had not landed on. Every diagnosis was confirmed by a clinician through the normal process before it counted. The model is a search and synthesis layer with very high recall over the long tail. That is the actual product.
A 4.8 percent yield on cases that were already considered unsolvable is, in a real clinical sense, a lot. The rare-disease space is the canonical setting where small absolute numbers are large humanitarian numbers, because the population of unsolved cases globally runs into the millions and each diagnosis is the difference between a family with a name for what their kid has and a family without one. The harder caveat OpenAI included up front is that o3 Deep Research is not approved for direct customer use in diagnosis. It is a research tool feeding into a clinician workflow, and the publication is partly a credibility move to keep it framed that way. That framing is going to come under pressure the moment a parent of a child on a five-year diagnostic odyssey reads the press release and asks their pediatrician to run it.
The other useful read on this paper is what it implies about where reasoning-model evaluations land next. NEJM AI publishing it at all is a signal. The journal exists because the medical establishment now considers structured peer-reviewed evaluation of clinical AI a category that needs its own venue, and the first batch of papers in that venue is starting to look like benchmark setting for what counts as a defensible deployment story in 2026. Rare-disease diagnosis is the easy domain to make this argument in, because the alternative is no answer. The harder argument, in cancer-care decision support or in primary care, comes next.