Microsoft reported an eye-catching result in June 2025: its experimental diagnostic orchestrator, paired with OpenAI’s o3 model, correctly solved as many as 85.5% of 304 exceptionally difficult medical cases adapted from the New England Journal of Medicine. The 21 experienced physicians recruited for the comparison averaged 19.9% across the cases they completed.

This is research reporting, not medical advice, and the system is not a substitute for clinical care. It was tested on a controlled, text-based benchmark, not on real patients. The work was released as a preprint, which means it had not completed journal peer review.

The date attached to those numbers matters. They come from the original Microsoft announcement and June 2025 manuscript. A major revision posted in November 2025 expanded the experiment and reported different figures: maximum accuracy of 84.5% for the AI system and an average of 36.1% for physicians on the journal cases. The requested headline preserves the widely reported launch result; the newer analysis is explained below.

What the original benchmark actually tested

The researchers built a benchmark called SDBench from 304 consecutive clinicopathological conference cases published between 2017 and 2025. These are not routine complaints pulled randomly from a waiting room. The New England Journal of Medicine’s Case Records are educational investigations that often revolve around rare conditions, ambiguous evidence, or surprising diagnoses.

That selection makes the exercise useful for testing diagnostic breadth, but it also changes how an accuracy percentage should be interpreted. Performance on a collection enriched for difficult, unusual cases does not reveal how a system would behave when common conditions dominate and false alarms can send patients into unnecessary testing.

SDBench converted each published case into a sequence. The diagnostic system started with limited information, requested tests or further details, revised its differential diagnosis, and eventually committed to an answer. A separate gatekeeper model held the complete case and disclosed information in response to those requests. This made the task more demanding than selecting an answer from a list, while remaining very different from bedside medicine.

The most recent 56 cases, published in 2024 and 2025, formed a hidden test set. The AI was evaluated across all 304 cases. The physicians were evaluated on cases assigned from that 56-case set.

How the orchestrator made one model act like a panel

The Microsoft AI Diagnostic Orchestrator, called MAI-DxO, did not simply ask a chatbot for a diagnosis. It directed one underlying language model to take on five jobs. A hypothesis agent proposed possible diagnoses. A challenger searched for weaknesses. A checklist agent looked for missing evidence. A test-selection agent chose the next investigation, and a stewardship agent weighed whether more testing was justified.

Those roles created a structured loop of proposing, criticizing, investigating, and revising. An additional judge assessed whether the final diagnosis matched the published answer. In the original manuscript, the most accurate configuration used an ensemble of o3 runs and reached 85.5%.

Accuracy came with substantial simulated expense. The authors estimated that the maximum-accuracy configuration used about $7,184 in model and test costs across the benchmark. A cost-conscious configuration reached 79.9% for roughly $2,396. These are experimental accounting estimates, not hospital budgets, but they show that performance depended partly on how much deliberation and testing the system was allowed.

Why 85.5 versus 20 is not a bedside contest

The physician group included 17 primary-care doctors and four hospital generalists from the United States and United Kingdom. Their median experience was 12 years. Each completed about 36 cases on average, producing 764 diagnoses, and spent an average of 11.8 minutes per case.

They worked without search engines, reference databases, or language models. Each doctor reasoned alone, while MAI-DxO simulated a multi-role panel. The AI processed all 304 cases, whereas the human comparison came from the held-out subset. The paper also used different interaction formats for practical reasons, although the authors reported that switching the AI to the physicians’ single-turn format did not reduce its accuracy on the test set.

More fundamentally, both sides encountered curated text. There was no physical examination, live patient conversation, imaging review, pathology slide, fragmented electronic record, treatment decision, or follow-up outcome. A correct name for a rare diagnosis is valuable, but clinical practice also requires deciding what is urgent, explaining uncertainty, responding to changing evidence, and avoiding harm.

The numbers changed when the paper was revised

The November 2025 revision, also available through the study’s version history, broadened the evaluation by adding 325 emergency-department cases from MIMIC-IV-ED and testing newer model configurations. Its maximum accuracy was 84.5% on the journal cases and 80.9% on the emergency-department set. The reported physician averages were 36.1% and 46.8%, respectively.

That does not necessarily mean the launch figures were wrong. A revised protocol, expanded dataset, different model, and updated comparison can change the outcome. It does mean that the bare claim “85.5% versus 20%” is incomplete without a version date. Preprints are useful because results can be examined quickly, but they can also change substantially before publication.

The revised manuscript still describes the work as a research benchmark and says it was submitted for external peer review. It also retains the central limitations: simulated text-only interaction, a concentration of pathological cases, no real electronic health records or multimodal evidence, and no measurement of patient outcomes.

What the result still demonstrates

Even with those caveats, the experiment offers a meaningful technical lesson. Breaking diagnostic reasoning into explicit roles can outperform a one-shot answer from the same class of model. The improvement suggests that orchestration, criticism, and disciplined information gathering may matter almost as much as the raw capabilities of the model underneath.

It also shows why a strong standalone score is only the beginning. A 2024 randomized clinical trial in JAMA Network Open found that giving 50 physicians access to GPT-4 did not significantly improve their diagnostic reasoning over conventional resources, even though the model alone scored higher in an exploratory comparison. Tool design, training, trust, and workflow determine whether capability becomes useful collaboration.

There are more clinically grounded ways to evaluate medical AI. For example, a large screening study covered previously by ScienceBlog measured cancer detection and radiologist workload inside a real care program. Such studies answer different questions from a benchmark of rare diagnostic puzzles.

What would count as clinical evidence

A convincing next stage would test systems prospectively with representative patients and real clinical teams. It would include common illnesses as well as rare ones, incomplete records, images and laboratory data, and patients whose conditions evolve. Researchers would need to track missed diagnoses, unnecessary tests, delays, cost, fairness across groups, clinician overreliance, and outcomes that matter to patients.

Independent replication would also be essential. The FDA’s framework for evaluating new medical AI uses stresses that evidence must match a product’s intended use and may require both nonclinical and clinical testing for safety and effectiveness. A retrospective benchmark cannot supply that evidence by itself.

Researchers would also need to publish failure analyses, not just average scores. A system that is excellent overall could still be unsafe if its misses cluster around particular populations, if it recommends risky tests too readily, or if its explanations make weak conclusions sound certain. Calibration and error severity matter alongside the percentage of final diagnoses marked correct.

Microsoft’s result is impressive as a demonstration of structured AI reasoning. It is not proof that an AI system can replace a physician, nor that it would improve care if dropped into a hospital tomorrow. The productive question is whether carefully evaluated tools can help clinical teams reach safer decisions, while leaving responsibility and patient care firmly anchored in medicine.