U

UCSF AI matches physicians at ED triage, correctly identifying urgent cases 89% of the time

“UCSF AI matches physicians at ED triage, correctly identifying urgent cases 89% of the time” documents a Patient Flow & Hospital Operations deployment in Hospital & Health System at UC San Francisco. www.news-medical.net reports triage accuracy (ai, full sample): 89%; this directory has not independently verified that result.

Maintained by Peter Korpak, Founder & Chief AnalystHow evidence is checked

Evidence at a glance

Evidence status:
Automated evidence gate passed
Deployment timeframe:
Not reported by source
Reported outcome metrics:
3 cited below
Directory entry published:
Source link checked:

The source-link check confirms reachability, not independent re-verification of every claim.

89%Triage Accuracy (AI, full sample)
88%Triage Accuracy (AI, physician sub-sample)
86%Triage Accuracy (Physician, sub-sample)

Source-reported figures — cited source: www.news-medical.net

The Challenge

Emergency departments nationwide are overcrowded and overtaxed, creating pressure on nurses and physicians to accurately triage patients at intake. The Emergency Severity Index triage process is resource-intensive and clinicians frequently face simultaneous urgent demands, making consistent prioritization difficult.

The Solution

UCSF researchers evaluated ChatGPT-4 (accessed via UCSF's secure generative AI platform with broad privacy protections) on its ability to extract symptoms from clinical notes and determine urgency. The model was tested against 251,000 de-identified adult ED visit records and benchmarked against physician performance on a 500-pair sub-sample.

Results

The LLM correctly identified which patient in a matched pair had the more serious condition 89% of the time across 10,000 pairs. In a 500-pair sub-sample evaluated by both the AI and a physician, the AI scored 88% accuracy versus 86% for the physician. Researchers note the model is not yet ready for clinical deployment without further validation.

Key Takeaways

  • LLMs can match or slightly exceed physician-level accuracy on ED triage prioritization using only symptom text from clinical notes.
  • Using real-world clinical data (251,000 visits) rather than simulated scenarios strengthens the validity of findings, but bias in training data remains an unresolved concern.
  • Clinical deployment requires additional validation, bias mitigation, and prospective trials before responsible use.

Share:

Details

Company Size
Enterprise
Evidence status
Automated evidence gate passed
Deployment timeframe
Not reported by source
Directory entry published
Source link checked

Have a similar implementation?

Share your customer's AI results and link it to your vendor profile.

Submit a case study →