In 88 of 98 patients, the diagnosis that doctors ultimately settled on appeared somewhere in the first seven candidates on Google’s AMIE chatbot’s ranked list. That is the 90 percent figure at the center of AMIE’s first paper in The Lancet, and it measures something narrower than the phrase “matched the doctors” suggests: the list, not the single best guess. AMIE’s top pick alone was right in 56 percent of cases, and its top three in 75 percent.
Google announced the Lancet publication on October 8, 2026, in a post by research lead Mike Schaekermann. The study was run with Beth Israel Deaconess Medical Center in Boston, and the Lancet paper is titled “Conversational diagnostic artificial intelligence in ambulatory primary care: a prospective feasibility study.”
A Text Chat Before the Appointment
The design, as laid out in Google Research’s study post, was a pre-registered, single-center, single-arm feasibility trial. Adults with new, non-emergency problems at the hospital’s ambulatory primary care clinic chatted with AMIE by text up to five days before their visit, with a physician supervising each chat live. Patients had to consent before the transcript and summary went to their own clinician.
A hundred adults completed the chat, and 98 went on to attend the appointment, which makes 98 the analysed sample. According to the full text of the same trial, 114 patients began the AMIE session, 100 finished it, and the two who missed their primary care appointment dropped out of the analysis. The sample skewed younger than the clinic’s usual urgent care population.
Final Diagnosis by Chart Review, Scored Top-Seven
The yardstick for “doctors’ diagnoses” was not what the treating physician wrote at the visit. A three-internist panel, blinded to AMIE’s output, set the final diagnosis by reviewing the patient’s chart about eight weeks after the appointment, including follow-up labs, imaging and specialist notes. Two internists then rated whether each AMIE candidate suggested that diagnosis or something very close to it.
Scored that way, the final diagnosis fell within AMIE’s first seven candidates in 88 of 98 cases, within the first three in 73, and in first position in 55. Forty-six of the 98 cases had a final diagnosis confirmed by a diagnostic test; the other 52 were presumptive. Google’s own blog sentence says AMIE’s differential diagnoses “matched the doctors’ final diagnoses 90% of the time,” but the Google Research post labels that figure top-seven in one passage, and the paper’s counts fix the cutoff.
Adam Rodman, listed as the paper’s last author, and first author Peter Brodeur are named in the arXiv full text, by initials, among the internists who set and graded the diagnoses; the paper’s author list runs past fifty names, with Schaekermann and Alan Karthikesalingam writing the Google Research post that announced the work in March 2026 and updated it on October 8 to note publication in The Lancet.
Primary Care Physicians Ahead on Practicality
Three clinical evaluators blindly graded each case’s differential and management plan, and the median grade was used. On the quality of the differential, AMIE and primary care physicians scored alike (p = 0.6), and on how appropriate (p = 0.1) and how safe (p = 1.0) the management plans were, the two were also similar. Physicians were rated better on how practical (p = 0.003) and how cost-effective (p = 0.004) their plans were.
That split is the clearest sign of what a text-only assistant lacks.
Google Research attributes the gap to AMIE having no access to the electronic health record, no physical exam and no images or other multimodal input. It also lists the study’s own limits: text only, no controlled comparison group, live rather than delayed supervision, and incomplete exploration of how health literacy and familiarity with chatbots shaped results. Without a control arm, the paper cannot say AMIE made care better, only that the setup was workable.
Supervisors could halt a chat under four predefined criteria, and none did, which the Google post puts as “not a single conversation needed to be interrupted.” Patients’ attitudes toward AI grew more positive after using AMIE (p < 0.001), and Google reports that its summaries helped clinicians prepare for the visit in 75 percent of cases and influenced their approach to care in more than half.
A companion argument appeared in a Nature Medicine comment by authors from Google, Harvard Medical School and Stanford, summarised by AI Weekly, saying trust in clinical AI has to come from prospective studies in real clinics and not from benchmark scores. Google’s own announcement says larger clinical trials are needed, and the authors of the arXiv version of the paper describe the result as initial feasibility, safety and user acceptance in a real-world setting; the published version is on The Lancet’s site. With 98 patients at one academic clinic, the study cannot say how AMIE’s seven-name list would perform across a wider population.
This article was produced with the assistance of AI and reviewed by Morning Overview editors prior to publication.
More from Morning Overview
- The NSA says three phone features should be off whenever you aren’t using them
- Hurricane Hunter radar shows four warning signs that a tropical cyclone is about to strengthen, a University of Miami study found
- Tropical Storm Rachel is dumping up to 12 inches on four Mexican states on its way to major hurricane strength
- The FTC says Lens.com doubled the price shoppers saw in Google ads