Comparison of Ophthalmologist and Large Language Model Chatbot Responses to Online Patient Eye Care Questions

Overview

Journal JAMA Netw Open

Publisher American Medical Association

Specialty General Medicine

Date 2023 Aug 22

PMID 37606922

Authors

Isaac A Bernstein

Youchen Victor Zhang

Devendra Govil

Iyad Majid

Robert T Chang

Yang Sun

Ann Shue

Jonathan C Chou

Emily Schehlein

Karen L Christopher

Sylvia L Groth

Cassie Ludwig

Sophia Y Wang

Affiliations

Soon will be listed here.

Abstract

Importance: Large language models (LLMs) like ChatGPT appear capable of performing a variety of tasks, including answering patient eye care questions, but have not yet been evaluated in direct comparison with ophthalmologists. It remains unclear whether LLM-generated advice is accurate, appropriate, and safe for eye patients.

Objective: To evaluate the quality of ophthalmology advice generated by an LLM chatbot in comparison with ophthalmologist-written advice.

Design, Setting, And Participants: This cross-sectional study used deidentified data from an online medical forum, in which patient questions received responses written by American Academy of Ophthalmology (AAO)-affiliated ophthalmologists. A masked panel of 8 board-certified ophthalmologists were asked to distinguish between answers generated by the ChatGPT chatbot and human answers. Posts were dated between 2007 and 2016; data were accessed January 2023 and analysis was performed between March and May 2023.

Main Outcomes And Measures: Identification of chatbot and human answers on a 4-point scale (likely or definitely artificial intelligence [AI] vs likely or definitely human) and evaluation of responses for presence of incorrect information, alignment with perceived consensus in the medical community, likelihood to cause harm, and extent of harm.

Results: A total of 200 pairs of user questions and answers by AAO-affiliated ophthalmologists were evaluated. The mean (SD) accuracy for distinguishing between AI and human responses was 61.3% (9.7%). Of 800 evaluations of chatbot-written answers, 168 answers (21.0%) were marked as human-written, while 517 of 800 human-written answers (64.6%) were marked as AI-written. Compared with human answers, chatbot answers were more frequently rated as probably or definitely written by AI (prevalence ratio [PR], 1.72; 95% CI, 1.52-1.93). The likelihood of chatbot answers containing incorrect or inappropriate material was comparable with human answers (PR, 0.92; 95% CI, 0.77-1.10), and did not differ from human answers in terms of likelihood of harm (PR, 0.84; 95% CI, 0.67-1.07) nor extent of harm (PR, 0.99; 95% CI, 0.80-1.22).

Conclusions And Relevance: In this cross-sectional study of human-written and AI-generated responses to 200 eye care questions from an online advice forum, a chatbot appeared capable of responding to long user-written eye health posts and largely generated appropriate responses that did not differ significantly from ophthalmologist-written responses in terms of incorrect information, likelihood of harm, extent of harm, or deviation from ophthalmologist community standards. Additional research is needed to assess patient attitudes toward LLM-augmented ophthalmologists vs fully autonomous AI content generation, to evaluate clarity and acceptability of LLM-generated answers from the patient perspective, to test the performance of LLMs in a greater variety of clinical contexts, and to determine an optimal manner of utilizing LLMs that is ethical and minimizes harm.

Citing Articles

Evaluating Artificial Intelligence in Spinal Cord Injury Management: A Comparative Analysis of ChatGPT-4o and Google Gemini Against American College of Surgeons Best Practices Guidelines for Spine Injury.

Yu A, Li A, Ahmed W, Saturno M, Cho S Global Spine J. 2025; :21925682251321837.

PMID: 39959933 PMC: 11833805. DOI: 10.1177/21925682251321837.

Large Language Models for Chatbot Health Advice Studies: A Systematic Review.

Huo B, Boyle A, Marfo N, Tangamornsuksan W, Steen J, McKechnie T JAMA Netw Open. 2025; 8(2):e2457879.

PMID: 39903463 PMC: 11795331. DOI: 10.1001/jamanetworkopen.2024.57879.

Current applications and challenges in large language models for patient care: a systematic review.

Busch F, Hoffmann L, Rueger C, van Dijk E, Kader R, Ortiz-Prado E Commun Med (Lond). 2025; 5(1):26.

PMID: 39838160 PMC: 11751060. DOI: 10.1038/s43856-024-00717-2.

Assessing the possibility of using large language models in ocular surface diseases.

Ling Q, Xu Z, Zeng Y, Hong Q, Qian X, Hu J Int J Ophthalmol. 2025; 18(1):1-8.

PMID: 39829624 PMC: 11672086. DOI: 10.18240/ijo.2025.01.01.

Large language models for accurate disease detection in electronic health records: the examples of crystal arthropathies.

Burgisser N, Chalot E, Mehouachi S, Buclin C, Lauper K, Courvoisier D RMD Open. 2025; 10(4).

PMID: 39794274 PMC: 11664341. DOI: 10.1136/rmdopen-2024-005003.

References

Van Bulck L, Moons P . What if your patient switches from Dr. Google to Dr. ChatGPT? A vignette-based survey of the trustworthiness, value, and danger of ChatGPT-generated responses to health questions. Eur J Cardiovasc Nurs. 2023; 23(1):95-98. DOI: 10.1093/eurjcn/zvad038. View

Yan A, McAuley J, Lu X, Du J, Chang E, Gentili A . RadBERT: Adapting Transformer-based Language Models to Radiology. Radiol Artif Intell. 2022; 4(4):e210258. PMC: 9344353. DOI: 10.1148/ryai.210258. View

Lee P, Bubeck S, Petro J . Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine. N Engl J Med. 2023; 388(13):1233-1239. DOI: 10.1056/NEJMsr2214184. View

Sallam M . ChatGPT Utility in Healthcare Education, Research, and Practice: Systematic Review on the Promising Perspectives and Valid Concerns. Healthcare (Basel). 2023; 11(6). PMC: 10048148. DOI: 10.3390/healthcare11060887. View

Yeo Y, Samaan J, Ng W, Ting P, Trivedi H, Vipani A . Assessing the performance of ChatGPT in answering questions regarding cirrhosis and hepatocellular carcinoma. Clin Mol Hepatol. 2023; 29(3):721-732. PMC: 10366809. DOI: 10.3350/cmh.2023.0089. View

Virtanen P, Gommers R, Oliphant T, Haberland M, Reddy T, Cournapeau D . SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods. 2020; 17(3):261-272. PMC: 7056644. DOI: 10.1038/s41592-019-0686-2. View

Jeblick K, Schachtner B, Dexl J, Mittermeier A, Stuber A, Topalis J . ChatGPT makes medicine easy to swallow: an exploratory case study on simplified radiology reports. Eur Radiol. 2023; 34(5):2817-2825. PMC: 11126432. DOI: 10.1007/s00330-023-10213-1. View

Grunebaum A, Chervenak J, Pollet S, Katz A, Chervenak F . The exciting potential for ChatGPT in obstetrics and gynecology. Am J Obstet Gynecol. 2023; 228(6):696-705. DOI: 10.1016/j.ajog.2023.03.009. View

Selivanov A, Rogov O, Chesakov D, Shelmanov A, Fedulova I, Dylov D . Medical image captioning via generative pretrained transformers. Sci Rep. 2023; 13(1):4171. PMC: 10010644. DOI: 10.1038/s41598-023-31223-5. View

10.

Rasmussen M, Larsen A, Subhi Y, Potapenko I . Artificial intelligence-based ChatGPT chatbot responses for patient and parent questions on vernal keratoconjunctivitis. Graefes Arch Clin Exp Ophthalmol. 2023; 261(10):3041-3043. DOI: 10.1007/s00417-023-06078-1. View

11.

Danilov G, Kotik K, Shevchenko E, Usachev D, Shifrin M, Strunina Y . Length of Stay Prediction in Neurosurgery with Russian GPT-3 Language Model Compared to Human Expectations. Stud Health Technol Inform. 2022; 289:156-159. DOI: 10.3233/SHTI210882. View

12.

Potapenko I, Boberg-Ans L, Hansen M, Klefter O, van Dijk E, Subhi Y . Artificial intelligence-based chatbot patient information on common retinal diseases using ChatGPT. Acta Ophthalmol. 2023; 101(7):829-831. DOI: 10.1111/aos.15661. View

13.

Hagan 3rd J, Kutryb M . Internet eye questions. Ophthalmology. 2009; 116(10):2036. DOI: 10.1016/j.ophtha.2009.05.008. View

14.

Sinha R, Deb Roy A, Kumar N, Mondal H . Applicability of ChatGPT in Assisting to Solve Higher Order Problems in Pathology. Cureus. 2023; 15(2):e35237. PMC: 10033699. DOI: 10.7759/cureus.35237. View

15.

Almazyad M, Aljofan F, Abouammoh N, Muaygil R, Malki K, Aljamaan F . Enhancing Expert Panel Discussions in Pediatric Palliative Care: Innovative Scenario Development and Summarization With ChatGPT-4. Cureus. 2023; 15(4):e38249. PMC: 10143975. DOI: 10.7759/cureus.38249. View

16.

Ayers J, Poliak A, Dredze M, Leas E, Zhu Z, Kelley J . Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Intern Med. 2023; 183(6):589-596. PMC: 10148230. DOI: 10.1001/jamainternmed.2023.1838. View

17.

Singh S, Djalilian A, Ali M . ChatGPT and Ophthalmology: Exploring Its Potential with Discharge Summaries and Operative Notes. Semin Ophthalmol. 2023; 38(5):503-507. DOI: 10.1080/08820538.2023.2209166. View

18.

Patel S, Lam K . ChatGPT: the future of discharge summaries?. Lancet Digit Health. 2023; 5(3):e107-e108. DOI: 10.1016/S2589-7500(23)00021-3. View

19.

Xie Y, Seth I, Hunter-Smith D, Rozen W, Ross R, Lee M . Aesthetic Surgery Advice and Counseling from Artificial Intelligence: A Rhinoplasty Consultation with ChatGPT. Aesthetic Plast Surg. 2023; 47(5):1985-1993. PMC: 10581928. DOI: 10.1007/s00266-023-03338-7. View

20.

Johnson S, King A, Warner E, Aneja S, Kann B, Bylund C . Using ChatGPT to evaluate cancer myths and misconceptions: artificial intelligence and cancer information. JNCI Cancer Spectr. 2023; 7(2). PMC: 10020140. DOI: 10.1093/jncics/pkad015. View