Plate Nº 16 · recorded October 10, 2026
Health & Medicine ResearchReported finding
AI Chatbot Passes First Real-World Safety Test in Primary Care
In a first-of-its-kind real-world trial at BIDMC, a medical AI chatbot held 98 patient chats with zero safety stops—though trust concerns and untested health outcomes remain.
By James Calloway4 min read878 words
In brief
- 98 patient conversations with the AMIE chatbot required zero safety-stop interventions during the April–November 2025 study
- Supervising physicians identified one hallucination and provided clinical clarification in 5 of 98 encounters
- In 75% of 44 reviewable cases, physicians said AI-generated summaries helped them prepare for visits
- Patient trust in the chatbot's confidentiality and honesty remained a concern despite favorable conversation ratings
- The study, published in The Lancet, assessed feasibility only, not whether AI improves health outcomes

Not one of 98 real patient conversations with a medical AI chatbot required a safety intervention from supervising physicians, according to what researchers believe is the first prospective real-world study of a patient-facing conversational AI system in primary care. The findings, published in The Lancet, mark an early but meaningful step toward testing medical AI with the same rigor applied to other health care interventions.
Researchers at Beth Israel Deaconess Medical Center (BIDMC), working with scientists at Google, tested the Articulate Medical Intelligence Explorer (AMIE), a medical AI chatbot designed to prepare patients for upcoming primary care visits. The study ran between April and November 2025 and enrolled 114 patients.
What did patients actually experience?
Patients used AMIE from home after scheduling an urgent primary care appointment with a BIDMC physician. Through a secure text-chat interface, the system asked about symptoms, gathered medical histories, offered possible diagnoses for patients to review with their doctor, and generated a summary for the clinician before the visit.
Of the 114 enrollees, 98 completed both the AI interaction and their primary care appointment. Every conversation was monitored in real time by a board-certified internal medicine physician who could step in if a patient appeared at risk of harm, showed significant emotional distress, asked to end the session, or if the physician spotted a safety concern—such as a need to clarify symptoms or give emergency care instructions.
Across all 98 completed encounters, none of the conversations required a safety-stop intervention. Supervising physicians identified one hallucination—a case where the AI generated incorrect information—and provided additional clinical clarification in five cases.
"Despite the theoretical performance of patient-facing AI in experimental settings, it is unclear how conversationally safe these encounters are in a real-world setting," said Adam Rodman, M.D., director of AI programs at the Carl J. Shapiro Institute for Research and Education at BIDMC. "Before we begin evaluating them as workload interventions, AI tools must first be shown to operate safely and acceptably in real care workflows."
How did patients and doctors rate the conversations?
Patients appeared receptive to the technology. They rated the quality of AMIE's conversations favorably—as did physician evaluators—across most measures. Patient attitudes toward AI improved after interacting with the chatbot and remained elevated after their appointment with a clinician.
"One of the most important questions was simply how patients would engage with a conversational AI system in a real clinical setting," said co-first author Jacob Koshy, M.D., M.P.H., a hospitalist in the Division of General Medicine at BIDMC. "By prospectively observing these interactions, we were able to better understand both the opportunities and the limitations of the technology in the context of routine patient care."
Where did trust fall short?
Patients gave high marks to the system's ability to listen, explain information and help them feel at ease. But two concerns persisted:
- Trust in the confidentiality of information shared with the system
- Trust in the chatbot's honesty and overall trustworthiness
The researchers said these findings underscore the importance of continued work to build patient trust as AI tools move into clinical care. "Research suggests that patient trust in medical AI systems flows from relationships with their human providers," Rodman said. "Future research will need to explore which interaction characteristics can build patient trust and how AI interactions can enhance the patient-physician relationship."
Did the AI help doctors prepare for visits?
The findings also point to possible practical value in clinical workflows. Primary care providers were able to review an AI-generated transcript or summary before seeing their patients in 44 cases. The reported results:
- In 75% of cases, physicians said reviewing the material helped them prepare for the visit
- In 57% of cases, they reported it may have influenced their clinical approach
- In one case, a provider called the interaction somewhat harmful, citing concern that a patient may have experienced anxiety after AMIE included lymphoma among its possible diagnoses
What the study does not show
Rodman and colleagues emphasized that the study was designed to evaluate feasibility—not to determine whether AI improves health outcomes. That question remains open, and the small sample of 98 completed encounters at a single medical center limits how far the results generalize.
"This study helps establish the baseline characteristics of such real-world conversations," said co-first author Peter Brodeur, M.D., a clinical fellow in cardiovascular medicine at BIDMC. "These baseline characteristics provide a foundation for all future prospective trials aimed at discovering when human-on-the-loop workflows may be appropriate."
Co-senior author Marc L. Cohen, M.D., clinical chief of primary care in the Division of General Medicine at BIDMC, framed the study as a template for how AI safety might be assessed before clinical implementation. "We believe the future involves continued involvement of human clinicians in patient care, but see a future where the triad of patient, human clinician and AI helps to provide our patients with the very best evidence-based care possible," Cohen said.
The paper is published as Peter G. Brodeur et al., "Conversational diagnostic artificial intelligence in ambulatory primary care: a prospective feasibility study," The Lancet (2026). DOI: 10.1016/s0140-6736(26)01535-7.
via Medical Xpress (Source)
More from James Calloway
Show full bio
Staff writer covering marketplaces and e-commerce at SciBeat.
205 articles
Nearby plates
- Doctors Can't Trace AI Chatbot Health Harm, Researchers Warn
- New AWARE Framework Helps Psychiatrists Ask About Patient AI Use
- AI Companions May Cause Lasting Psychological Harm, Study Finds
- Nature Examines AI Agents That Help Design Clinical Trials
- AI System Aims to Speed Up Clinical Trial Patient Matching