Google has published new AMIE research that moves its medical AI work from one-off diagnostic conversations toward longer disease-management scenarios. The system, short for Articulate Medical Intelligence Explorer, is still a research model rather than a clinical product, but the study points to where medical AI evaluation is heading.

The Nature paper tested AMIE in a randomized, blinded virtual clinical exam against 21 primary care physicians. The setup covered 100 multi-visit case scenarios across five specialties and was designed around UK NICE guidance and BMJ Best Practice references. Specialist physician evaluators reviewed the management plans.

Google says AMIE was non-inferior to the physicians on management reasoning. The paper also reports stronger scores for treatment and investigation precision, plus alignment with and grounding in clinical guidelines. A separate medication-reasoning benchmark, RxQA, was used to test difficult prescribing questions derived from drug formularies.

The practical caveat is important. These were simulated cases, not live clinical deployments with real patients, messy records, liability constraints, or local prescribing workflows. Google says more work is needed before a system like AMIE could be used in care settings, including real-world studies and safety validation.

That makes the result notable but narrow. The advance is less about replacing clinicians and more about testing whether conversational models can follow a patient across multiple visits, cite clinical guidance, and maintain a coherent management plan under specialist review.