A Quantitative Assessment of ChatGPT as a Neurosurgical Triaging Tool

Neurosurgery 95:487–495, 2024

ChatGPT is a natural language processing chatbot with increasing applicability to the medical workflow. Although ChatGPT has been shown to be capable of passing the American Board of Neurological Surgery board examination, there has never been an evaluation of the chatbot in triaging and diagnosing novel neurosurgical scenarios without defined answer choices. In this study, we assess ChatGPT’s capability to determine the emergent nature of neurosurgical scenarios and make diagnoses based on information one would find in a neurosurgical consult.

METHODS: Thirty clinical scenarios were given to 3 attendings, 4 residents, 2 physician assistants, and 2 subinterns. Participants were asked to determine if the scenario constituted an urgent neurosurgical consultation and what the most likely diagnosis was. Attending responses provided a consensus to use as the answer key. Generative pretraining transformer (GPT) 3.5 and GPT 4 were given the same questions, and their responses were compared with the other participants.

RESULTS: GPT 4 was 100% accurate in both diagnosis and triage of the scenarios. GPT 3.5 had an accuracy of 92.59%, slightly below that of a PGY1 (96.3%), an 88.24% sensitivity, 100% specificity, 100% positive predictive value, and 83.3% negative predicative value in triaging each situation. When making a diagnosis, GPT 3.5 had an accuracy of 92.59%, which was higher than the subinterns and similar to resident responders.

CONCLUSION: GPT 4 is able to diagnose and triage neurosurgical scenarios at the level of a senior neurosurgical resident. There has been a clear improvement between GPT 3.5 and 4. It is likely that the recent updates in internet access and directing the functionality of ChatGPT will further improve its utility in neurosurgical triage.

Vertebral Bone Quality Score as a Predictor of Adjacent Segment Disease After Lumbar Interbody Fusion

Neurosurgery 95:284–296, 2024

With lumbar spine fusion being one of the most commonly performed spinal surgeries, investigating common complications such as adjacent segment disease (ASD) is a high priority. To the authors’ knowledge, there are no previous studies investigating the utility of the preoperative magnetic resonance imaging based vertebral bone quality (VBQ) score in predicting radiographic and surgical ASD after lumbar spine fusion. We aimed to investigate the predictive factors for radiographic and surgical ASD, focusing on the predictive potential of the VBQ score.

METHODS: A single-center retrospective analysis was conducted of all patients who underwent 1–3 level lumbar or lumbosacral interbody fusion for lumbar spine degenerative disease between 2014 and 2021 with a minimum 12 months of clinical and radiographic follow-up. Demographic data were collected, along with patient medical, and surgical data. Preoperative MRI was assessed in the included patients using the VBQ scoring system to identify whether radiographic ASD or surgical ASD could be predicted.

RESULTS: A total of 417 patients were identified (mean age, 59.8 ± 12.4 years; women, 54.0%). Eighty-two (19.7%) patients developed radiographic ASD, and 58 (13.9%) developed surgical ASD. A higher VBQ score was a significant predictor of radiographic ASD in univariate analysis (2.4 ± 0.5 vs 3.3 ± 0.4; P < .001) and multivariate analysis (odds ratio, 1.601; 95% CI, 1.453-1.763; P < .001). For surgical ASD, a significantly higher VBQ score was seen in univariate analysis (2.3 ± 0.5 vs 3.3 ± 0.4; P < .001) and served as an independent risk factor in multivariate analysis (odds ratio, 1.509; 95% CI, 1.324-1.720; P < .001). We also identified preoperative disk bulge and preoperative existence of adjacent segment disk degeneration to be significant predictors of both radiographic and surgical ASD. Furthermore, 3-level fusion was also a significant predictor for surgical ASD.

CONCLUSION: The VBQ scoring system might be a useful adjunct for predicting radiographic and surgical ASD.