2026-10-09 · ← News
Gemini handled 98 real patients, with a physician watching every chat
A Lancet study placed the Gemini 2.5 based AMIE system in real urgent primary care: 98 patients completed both the chat and physician appointment, and no conversation required a stop under predefined safety rules. Yet a physician monitored every conversation in real time and the system had no access to medical records.
The image could not be loaded.
A prospective study published in The Lancet tested the conversational AMIE system during real urgent primary-care visits. Of 114 enrolled adults, 98 completed both the AI chat and subsequent physician appointment. That is much closer to clinical practice than a test on synthetic cases, but it remained a feasibility study with unusually close supervision.
AMIE interviewed patients before the visit and gave clinicians its summary
Patients held a text conversation with AMIE up to 5 days before an appointment for a single urgent complaint. The system gathered history, presented possible diagnoses and suggested topics to discuss with the clinician. A primary-care physician could then review the transcript and automated summary before the appointment.
AMIE used Gemini 2.5 Pro with a tailored agent workflow rather than an unmodified consumer chatbot. After the first 50 encounters, researchers switched to Gemini 2.5 Flash because of latency. The system had no access to medical records, but it conducted its own patient interview and maintained a live summary, differential diagnosis and draft management plan.
Quality approached physicians while plan practicality lagged
Blinded evaluation found no significant difference between AMIE and physicians in overall differential-diagnosis quality or in the appropriateness and safety of management plans. Physicians produced more practical and cost-effective plans. That distinction separates medical advice that reads well from a decision that can be carried out in a specific clinic.
Physicians completed follow-up surveys in 60 of 98 cases. In 44 cases they had reviewed AMIE's output before the visit; 33 found it helpful for preparation and 25 said it might have changed their behavior. The study therefore also measures the value of a structured pre-visit intake, not only a contest between a model and a doctor.
Zero stopped chats came with continuous physician supervision
Under the predefined rules, physician supervisors did not have to stop any of the 100 completed AMIE conversations. Yet a safety physician watched every chat live through screen sharing. Supervisors added clinical information in five interactions and recorded one hallucination. Zero stops therefore do not mean zero errors or safety without supervision.
The study took place at one academic center, included only English-speaking adults and had no control arm for its primary safety outcomes. It did not test unsupervised home deployment or work with complete medical records. The results support supervised feasibility, not physician replacement.
The next test must measure autonomy, patient outcomes and supervision cost
Larger controlled studies across health systems and languages will matter next. They need to track patient outcomes, incorrect urgency decisions, returns to care and cases where model and clinician disagree. The cost of safety supervision and the time clinicians save or spend reviewing output are equally important.
If AMIE shortens history taking and hands clinicians a useful summary, it may create value well before autonomous diagnosis. Moving from that role to an independent clinical adviser requires evidence that a study with a physician watching every screen cannot yet provide.
Lilith's verdict
AMIE entered the clinic with a capable voice, but a physician stood behind its chair with a hand on the switch. The progress is real; the solo shift has not begun.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗