Skip to content

AI versus MD: Evaluating the surgical decision-making accuracy of ChatGPT-4 (18 May 2024)

Summary

This study in Surgery aimed to investigate the accuracy of ChatGPT-4’s surgical decision-making compared with general surgery residents and attending surgeons. Five clinical scenarios were created from actual patient data based on common general surgery diagnoses. Scripts were developed to sequentially provide clinical information and ask decision-making questions. Responses to the prompts were scored based on a standardised rubric for a total of 50 points. Each clinical scenario was run through Chat GPT-4 and sent electronically to all general surgery residents and attendings at a single institution. Scores were compared using Wilcoxon rank sum tests.

The results showed that, when faced with surgical patient scenarios, ChatGPT-4 performed superior to junior residents and equivalent to senior residents and attendings. The authors argue that large language models, such as ChatGPT, may have the potential to be an educational resource for junior residents to develop surgical decision-making skills.

AI versus MD: Evaluating the surgical decision-making accuracy of ChatGPT-4 (18 May 2024) https://www.surgjournal.com/article/S0039-6060(24)00227-7/abstract

User Feedback

Recommended Comments

There are no comments to display.

Create an account or sign in to comment

Registered address: Patient Safety Learning, China Works SB203, 100 Black Prince Road Vauxhall, London, SE1 7SJ

Account

Navigation

Search

Search

Configure browser push notifications

Chrome (Android)
  1. Tap the lock icon next to the address bar.
  2. Tap Permissions → Notifications.
  3. Adjust your preference.
Chrome (Desktop)
  1. Click the padlock icon in the address bar.
  2. Select Site settings.
  3. Find Notifications and adjust your preference.