Comparison Of Quality, Accuracy, And Empathy Of Physician And AI Responses (PAIR) to Real Patient Questions: Physicians Vs. Various LLM

Event: PSTM 2024
Sat, 9/28/2024: 7:45 AM - 7:50 AM
42376 
Abstracts 
SDCC 
Purpose: to evaluate whether responses from ChatGPT and other LLMs could provide comparable quality, accuracy, and empathy to physician answers after controlling for potentially confounding variables such as the length of responses.

Method: This cross-sectional study used RealSelf.com, an online public platform where patients pose questions publicly, which are answered by verified physicians in non-anonymous accounts. Questions from the site were entered as prompts into ChatGPT and Bard and the responses were compared to the responses from physicians on RealSelf.com. Three blinded physicians evaluated matchups between different response pairs (i.e. RealSelf vs. ChatGPT, RealSelf vs. Bard, or ChatGPT vs. Bard). They chose "which response was better" and rated the quality, accuracy, and empathy of each response. The next generation generative AI bot, MedLM was also tested. MedLM was designed for healthcare and built on Med-PalM2 that had previously scored an 85% on the USMLE.

Results: Initial data indicates that physician responses trended higher than ChatGPT scores, but did not show significant difference. Physician responses were significantly higher quality, more accurate, and more empathetic than Bard responses. However, ChatGPT responses, while trending higher than Bard responses, were not significantly higher. The proportion of responses with unacceptable quality, accuracy, and empathy was indistinguishable between physicians, ChatGPT and Bard. MedLM is currently being tested.

Conclusions: In this study, after accounting for verified physician responses and response length, ChatGPT and Bard generated acceptable quality, accurate and empathetic responses to patient questions posed in an online forum. However, overall the physician answers were preferred over chatbot responses in 1:1 matchups.

Tracks

Research and Technology
PSTM 2024