Search traffic for “AI in hospitals” is sold a demo. On 19 August 2026, Nature Medicine published the honest sibling of that demo: a prospective DECIDE-AI stage 1 evaluation of SHAKED, a multi-model large-language-model clinical decision support system, in the emergency department of Rambam Health Care Campus in Haifa, with Technion co-authors. The accuracy number is not the line the authors lead with. Use fell.
What happened
Over four weeks the team analysed 1,138 patients across two parallel units: one with SHAKED, one on routine rotations. ClinicalTrials.gov: NCT06902675. Corresponding author Shahar Shelly, Rambam neurology and Technion faculty. Equal first authors Liron Leibovitch and Adi Ahituv. Received 1 March 2026; accepted 22 July 2026; published 19 August 2026. DOI 10.1038/s41591-026-04601-5.
SHAKED is not a single foundation model with a chat window glued on. Extended data describe an event-driven pipeline off the EHR: an event server and scheduler pull a structured history, a retrieval-augmented generation module sits behind a clinician chat interface, and inference ran on AWS Bedrock. Funding was AWS Partner support plus $20,000 in cloud credits. The funder had no role in design or analysis. The authors declare no competing interests.
| Measure | Result | Source |
|---|---|---|
| Patients | 1,138 in 4 weeks, two parallel wings | abstract |
| Expert review of sampled outputs | 99 of 100 clinically appropriate | abstract |
| Adverse events | none detected | abstract |
| ED length of stay | 4.9 hours in both wings, P = 0.99 | abstract |
| Consultation cycle (ITT) | −9.4 minutes, P = 0.077 (non-significant) | abstract |
| Clinical adoption | 68% → 30% | abstract |
| Workload disengagement | OR = 0.72 per shift hour (95% CI 0.62–0.83) | abstract |
| Preference for radiology consults | OR = 2.98 (95% CI 1.58–5.63) | abstract |
Wing assignment was not a randomised cluster in the sense that would support causal claims about length of stay. A propensity-score model for wing assignment had AUC 0.543 (n = 1,133) — close to chance — which is what you want if the wings were similar. After inverse-probability weighting, standardised mean differences fell inside ±0.10 (maximum 0.032). Per-protocol estimates of consultation time are flagged in the extended data as possibly biased by selective adoption. The paper’s last line is the editorial constraint: the findings inform randomised-trial design and “do not justify clinical deployment of AI clinical decision support at this stage.”
Analysis code is public (MIT) at github.com/llironlibo/SHAKED-analysis and Zenodo 10.5281/zenodo.20736931. Patient-level data are controlled-access under Israeli health privacy law and Rambam IRB.
Why it matters
Most hospital-AI coverage still treats the model as the bottleneck. This paper measured the bottleneck as whether a tired clinician still opens the tool in hour eight. Adoption fell from 68 percent to 30 percent, and the odds ratio per shift hour is 0.72. Radiology consults were the use case physicians preferred. Length of stay did not move. The consultation-cycle trend did not clear P = 0.05 on intention-to-treat.
That combination is the Good Signal version of an AI story: a named hospital, a registered protocol, a number that survived peer review, and a limit the authors wrote themselves. It is the honest sibling of the LiON liver-CT trial published the same day in the same journal. LiON changed 37 reports as an extra reader on a CT queue. SHAKED asked emergency physicians to keep engaging a chat interface under load. One deployed as a second set of eyes on images; the other asked for a click in a crowded bay. The engagement curves are the news.
DECIDE-AI is a reporting guideline for early-stage clinical evaluation of AI decision support (Nat Med 2022). Calling this stage 1 is a discipline, not a hedge. No adverse events in four weeks is not a long-term safety file. Ninety-nine of 100 sampled outputs is not a claim that the 1,138-patient run was error-free. It is a sample.
What to watch next
- A randomised trial that treats engagement as a primary endpoint, not an afterthought. Until then the verified claim is 99/100 appropriate outputs, unchanged length of stay, and use falling from 68% to 30%.
- Shift-hour design. If the OR of 0.72 is real, a tool that only works in hour one is not an ED tool. Watch whether the next protocol randomises at the shift or the wing.
- Do not deploy from this paper. The authors already wrote that sentence. Quote it.
Sources
- Leibovitch, L. et al. Prospective evaluation of a large language model clinical decision support system in the emergency department. Nat Med (2026). https://doi.org/10.1038/s41591-026-04601-5
- ClinicalTrials.gov NCT06902675 — https://clinicaltrials.gov/study/NCT06902675
- Analysis code — https://github.com/llironlibo/SHAKED-analysis and https://doi.org/10.5281/zenodo.20736931



