A Quantitative Assessment of ChatGPT as a Neurosurgical Triaging Tool

Neurosurgery 95:487–495, 2024

ChatGPT is a natural language processing chatbot with increasing applicability to the medical workflow. Although ChatGPT has been shown to be capable of passing the American Board of Neurological Surgery board examination, there has never been an evaluation of the chatbot in triaging and diagnosing novel neurosurgical scenarios without defined answer choices. In this study, we assess ChatGPT’s capability to determine the emergent nature of neurosurgical scenarios and make diagnoses based on information one would find in a neurosurgical consult.

METHODS: Thirty clinical scenarios were given to 3 attendings, 4 residents, 2 physician assistants, and 2 subinterns. Participants were asked to determine if the scenario constituted an urgent neurosurgical consultation and what the most likely diagnosis was. Attending responses provided a consensus to use as the answer key. Generative pretraining transformer (GPT) 3.5 and GPT 4 were given the same questions, and their responses were compared with the other participants.

RESULTS: GPT 4 was 100% accurate in both diagnosis and triage of the scenarios. GPT 3.5 had an accuracy of 92.59%, slightly below that of a PGY1 (96.3%), an 88.24% sensitivity, 100% specificity, 100% positive predictive value, and 83.3% negative predicative value in triaging each situation. When making a diagnosis, GPT 3.5 had an accuracy of 92.59%, which was higher than the subinterns and similar to resident responders.

CONCLUSION: GPT 4 is able to diagnose and triage neurosurgical scenarios at the level of a senior neurosurgical resident. There has been a clear improvement between GPT 3.5 and 4. It is likely that the recent updates in internet access and directing the functionality of ChatGPT will further improve its utility in neurosurgical triage.

Predicting Spinal Surgery Candidacy From Imaging Data Using Machine Learning

Neurosurgery 89:116–121, 2021

The referral process for consultation with a spine surgeon remains inefficient, given a substantial proportion of referrals to spine surgeons are nonoperative.

OBJECTIVE: To develop a machine-learning-based algorithm which accurately identifies patients as candidates for consultation with a spine surgeon, using only magnetic resonance imaging (MRI).

METHODS: We trained a deep U-Net machine learning model to delineate spinal canals on axial slices of 100 normal lumbar MRI scans which were previously delineated by expert radiologists and neurosurgeons. We then tested the model against lumbar MRI scans for 140 patients who had undergone lumbar spine MRI at our institution (60 of whom ultimately underwent surgery, and 80 of whom did not). The model generated automated segmentations of the lumbar spinal canals and calculated a maximum degree of spinal stenosis for each patient,which served as our biomarker for surgical pathology warranting expert consultation.

RESULTS: Themachine learning model correctly predicted surgical candidacy (ie, whether patients ultimately underwent lumbar spinal decompression) with high accuracy (area under the curve = 0.88), using only imaging data from lumbar MRI scans.

CONCLUSION: Automated interpretation of lumbar MRI scans was sufficient to correctly determine surgical candidacy in nearly 90% of cases. Given that a significant proportion of referrals placed for spine surgery evaluation fail to meet criteria for surgical intervention, our model could serve as a valuable tool for patient triage and thereby address some of the inefficiencies within the outpatient surgical referral process.