Machine learning models for predicting patient satisfaction after adult spinal deformity surgery

J Neurosurg Spine 44:457–468, 2026

This clinical study develops and internally validates machine learning–guided logistic regression models to predict patient satisfaction 24 months after adult spinal deformity (ASD) surgery, using 213 patients and three feature-selection methods. Nine routinely measurable predictors—including postoperative WOMAC function, frailty, pelvic compensation, imaging MCID achievement, rFCSA, and SVA—were identified and ranked by SHAP for their influence on satisfaction.

The model showed strong discrimination (AUROC 0.846) and calibration, yielded a nomogram for individualized prognostication, and emphasizes modifiable targets for perioperative care and rehabilitation. Limitations include single-center retrospective design, modest sample size, and inclusion of postoperative variables limiting preoperative decision use.

Goal Develop and internally validate models to predict patient satisfaction 24 months after adult spinal deformity (ASD) surgery, using SRS-22r satisfaction (high satisfaction defined as score ≥ 4.5).

Cohort 213 ASD patients met criteria; 128 (60%) used for training and 85 (40%) for internal test validation.

Pipeline Used three ML feature-selection methods—LASSO, recursive feature elimination (RFE), and Boruta—and retained variables consistently selected by all three.

Final predictors Nine key indicators were retained: rFCSA, fatty infiltration, frailty, pelvic compensation, postoperative SVA, imaging MCID achievement, postoperative subtotal score, postoperative WOMAC function, and change in WOMAC function.

Model Built an interpretable logistic regression model from these predictors; binary cutoff optimized via ROC/Youden index, with SHAP used to rank feature importance.

Performance In the test set, the model achieved AUROC 0.846 and accuracy 0.812 (also reported AUPRC 0.894 and Brier score 0.153).

Top drivers (SHAP order) Higher postoperative WOMAC function, absence of frailty, imaging MCID achieved, larger WOMAC function improvement, higher rFCSA, higher postoperative subtotal, lower postoperative SVA, successful pelvic compensation, and lower fatty infiltration increased satisfaction likelihood.

Implication/limitation Intended mainly to identify modifiable factors to guide postoperative rehabilitation; practical preoperative counseling is limited because key inputs include postoperative variables, and external multicenter validation is still needed.

AtlasGPT: a language model grounded in neurosurgery with domain-specific data and document retrieval

J Neurosurg 143:560–567, 2025

AtlasGPT, a neurosurgery-specific large language model grounded in expert-verified sources and retrieval-augmented generation, outperformed GPT-4 and Gemini Advanced on a neurosurgery board exam, showed greater resistance to medical misinformation, and generated more comprehensive, relevant, and well-referenced answer explanations than standard preparation materials.

• AtlasGPT is a neurosurgery-specific large language model (LLM) built on GPT-4 with retrieval-augmented generation (RAG) from trusted neurosurgical sources.

• AtlasGPT outperformed GPT-4 and Gemini Advanced on a 149-question neurosurgery board exam (accuracy: 90.6% vs 80.5%).

• AtlasGPT showed the highest accuracy on spine and imaging-based questions, even without access to image data.

• In adversarial testing, AtlasGPT was more robust to misinformation, being fooled only 14% of the time, compared to 44% for GPT-4 and 68% for Gemini Advanced.

• Expert neurosurgeons rated AtlasGPT’s explanations as more comprehensive, relevant, and better referenced than official board prep materials.

• AtlasGPT did not produce hallucinations or harmful content in its responses.

• The study suggests domain-specific LLMs like AtlasGPT can enhance medical education, decision-making, and exam preparation in complex fields.

• Limitations include use of a single question bank and need for broader source material in future work.

Stratifying trigeminal neuralgia and characterizing an abnormal property of brain functional organization: a resting-state fMRI and machine learning study

J Neurosurg 143:74–82, 2025

Resting-state fMRI and machine learning revealed distinct brain connectivity and activity differences between classical and idiopathic trigeminal neuralgia (TN) and controls. These findings identify potential neuroimaging biomarkers for TN subtypes, aiding diagnosis and understanding of TN pathophysiology.

Primary trigeminal neuralgia (TN) includes classical (CTN) and idiopathic (ITN) types, sharing clinical features but differing in neurovascular compression (NVC) presence.

• Resting-state fMRI and machine learning were used to analyze brain functional connectivity and spontaneous activity in 50 TN patients (28 CTN, 22 ITN) and 43 controls.

• TN patients showed increased connectivity between the medial prefrontal cortex (mPFC) and left planum temporale, and decreased connectivity between mPFC and left superior frontal gyrus.

• CTN patients had further reduced connectivity between the left insula and left occipital pole, and decreased activity in the right temporal pole compared to ITN.

• TN patients exhibited heightened neural activity in frontal regions compared to controls.

• Machine learning (support vector machine) distinguished TN patients from controls with moderate accuracy (AUC 0.80).

• Findings suggest potential fMRI biomarkers for TN subtypes, aiding understanding of pathophysiology and improving diagnosis.

• Study limitations include small sample size and exclusion of bilateral/secondary TN, warranting further research.

Artificial Intelligence for Automatic Analysis of Shunt Treatment in Presurgery and Postsurgery Computed Tomography Brain Scans of Patients With Idiopathic Normal Pressure Hydrocephalus

Neurosurgery 95:1329–1337, 2024

Ventriculo-peritoneal shunt procedures can improve idiopathic normal pressure hydrocephalus (iNPH) symptoms. However, there are no automated methods that quantify the presurgery and postsurgery changes in the ventricular volume for computed tomography scans. Hence, the main goal of this research was to quantify longitudinal changes in the ventricular volume and its correlation with clinical improvement in iNPH symptoms. Furthermore, our objective was to develop an end-to-end graphical interface where surgeons can directly drag-drop a brain scan for quantified analysis.

METHODS: A total of 15 patients with 47 longitudinal computed tomography scans were taken before and after shunt surgery. Postoperative scans were collected between 1 and 45 months. We use a UNet-based model to develop a fully automated metric. Center slices of the scan that are most representative (80%) of the ventricular volume of the brain are used. Clinical symptoms of gait, balance, cognition, and bladder continence are studied with respect to the proposed metric.

RESULTS: Fifteen patients with iNPH demonstrate a decrease in ventricular volume (as shown by our metric) postsurgery and a concurrent clinical improvement in their iNPH symptomatology. The decrease in postoperative central ventricular volume varied between 6 cc and 33 cc (mean: 20, SD: 9) among patients who experienced improvements in gait, bladder continence, and cognition. Two patients who showed improvement in only one or two of these symptoms had <4 cc of cerebrospinal fluid drained. Our artificial intelligence–based metric and the graphical user interface facilitate this quantified analysis.

CONCLUSION: Proposed metric quantifies changes in ventricular volume before and after shunt surgery for patients with iNPH, serving as an automated and effective radiographic marker for a functioning shunt in a patient with iNPH.

Subtemporal Approach for the Treatment of Ruptured and Unruptured Distal Basilar Artery Aneurysms: Is There a Contemporary Use?

Operative Neurosurgery 27:581–596, 2024

Distal basilar artery aneurysms (DBAs) are high-risk lesions for which endovascular treatment is preferred because of their deep location, yet indications for open clipping nonetheless remain. The subtemporal approach allows for early proximal control and direct visualization of critical posterior perforating arteries, especially for posterior-projecting aneurysms. Our objective was to describe our clinical experience with the subtemporal approach for clipping DBAs in the evolving endovascular era.

METHODS: This was a retrospective, single-institution case series of patients with DBAs treated with microsurgery over a 21-year period (2002-2023). Demographic, clinical, and surgical data were collected for analysis.

RESULTS: Twenty-seven patients underwent clipping of 11 ruptured and 16 unruptured DBAs with a subtemporal approach (24 female; mean age 53 years). Ten patients had expanded craniotomies for treatment of additional aneurysms. The aneurysm occlusion rate was 100%. Good neurological outcomes as defined by the modified Rankin Scale score ≤2 and Glasgow Outcome Scale score ≥4 were achieved in 21/27 patients (78%). Two patients died before hospital discharge, one from vasospasm-induced strokes and another from an intraoperative myocardial infarction.

CONCLUSION: These results demonstrate that microsurgical clip ligation of DBAs using the subtemporal approach remains a viable option for complex lesions not amenable to endovascular management.

A Quantitative Assessment of ChatGPT as a Neurosurgical Triaging Tool

Neurosurgery 95:487–495, 2024

ChatGPT is a natural language processing chatbot with increasing applicability to the medical workflow. Although ChatGPT has been shown to be capable of passing the American Board of Neurological Surgery board examination, there has never been an evaluation of the chatbot in triaging and diagnosing novel neurosurgical scenarios without defined answer choices. In this study, we assess ChatGPT’s capability to determine the emergent nature of neurosurgical scenarios and make diagnoses based on information one would find in a neurosurgical consult.

METHODS: Thirty clinical scenarios were given to 3 attendings, 4 residents, 2 physician assistants, and 2 subinterns. Participants were asked to determine if the scenario constituted an urgent neurosurgical consultation and what the most likely diagnosis was. Attending responses provided a consensus to use as the answer key. Generative pretraining transformer (GPT) 3.5 and GPT 4 were given the same questions, and their responses were compared with the other participants.

RESULTS: GPT 4 was 100% accurate in both diagnosis and triage of the scenarios. GPT 3.5 had an accuracy of 92.59%, slightly below that of a PGY1 (96.3%), an 88.24% sensitivity, 100% specificity, 100% positive predictive value, and 83.3% negative predicative value in triaging each situation. When making a diagnosis, GPT 3.5 had an accuracy of 92.59%, which was higher than the subinterns and similar to resident responders.

CONCLUSION: GPT 4 is able to diagnose and triage neurosurgical scenarios at the level of a senior neurosurgical resident. There has been a clear improvement between GPT 3.5 and 4. It is likely that the recent updates in internet access and directing the functionality of ChatGPT will further improve its utility in neurosurgical triage.

Development and validation of an artificial intelligence model to accurately predict spinopelvic parameters

J Neurosurg Spine 41:88–96, 2024

Achieving appropriate spinopelvic alignment has been shown to be associated with improved clinical symptoms. However, measurement of spinopelvic radiographic parameters is time-intensive and interobserver reliability is a concern. Automated measurement tools have the promise of rapid and consistent measurements, but existing tools are still limited to some degree by manual user-entry requirements. This study presents a novel artificial intelligence (AI) tool called SpinePose that automatically predicts spinopelvic parameters with high accuracy without the need for manual entry.

METHODS SpinePose was trained and validated on 761 sagittal whole-spine radiographs to predict the sagittal vertical axis (SVA), pelvic tilt (PT), pelvic incidence (PI), sacral slope (SS), lumbar lordosis (LL), T1 pelvic angle (T1PA), and L1 pelvic angle (L1PA). A separate test set of 40 radiographs was labeled by four reviewers, including fellowship-trained spine surgeons and a fellowship-trained radiologist with neuroradiology subspecialty certification. Median errors relative to the most senior reviewer were calculated to determine model accuracy on test images. Intraclass correlation coefficients (ICCs) were used to assess interrater reliability.

RESULTS SpinePose exhibited the following median (interquartile range) parameter errors: SVA 2.2 mm (2.3 mm) (p = 0.93), PT 1.3° (1.2°) (p = 0.48), SS 1.7° (2.2°) (p = 0.64), PI 2.2° (2.1°) (p = 0.24), LL 2.6° (4.0°) (p = 0.89), T1PA 1.1° (0.9°) (p = 0.42), and L1PA 1.4° (1.6°) (p = 0.49). Model predictions also exhibited excellent reliability at all parameters (ICC 0.91–1.0).

CONCLUSIONS SpinePose accurately predicted spinopelvic parameters with excellent reliability comparable to that of fellowship-trained spine surgeons and neuroradiologists. Utilization of predictive AI tools in spinal imaging can substantially aid in patient selection and surgical planning.

Invention of an Online Interactive Virtual Neurosurgery Simulator With Audiovisual Capture for Tactile Feedback

Operative Neurosurgery 24:194–200, 2023

BACKGROUND: Present neurosurgical simulators are not portable.

OBJECTIVE: To maximize portability of a virtual surgical simulator by providing online learning and to validate a unique psychometric method (“audiovisual capture”) to provide tactile information without force feedback probes.

METHODS: An online interactive neurosurgical simulator of a posterior petrosectomy was developed. The difference in the hardness of compact vs cancellous bone was presented with audiovisual effects as inclinations of the drilling speed and sound based on engineering perspectives. Three training methods (the developed simulator, lectures and review of slides, and dissection of a 3-dimensional printed temporal bone model [D3DPM]) were evaluated by 10 neurosurgical residents. They all first attended a lecture and were randomly allocated to 2 groups by the training D3DPM (A: simulator; B: review of slides, no simulator). In D3DPM, objective measures (required time, quality of completion, injury scores of important structures, and the number of instructions provided) were compared between groups. Finally, the residents answered questionnaires.

RESULTS: The objective measures were not significantly different between groups despite a younger tendency in group A (graduate year À2.4 years, 95% confidence interval À5.3 to 0.5, P = .081). The mean perceived hardness of cancellous bone on the simulator was 70% of that of compact bone, matching the intended profile. The simulator was superior to lectures and review of slides in feedback and repeated practices and to D3DPM in adaptability to multiple learning environments.

CONCLUSION: A novel online interactive neurosurgical simulator was developed, and satisfactory validity was shown. Audiovisual capture successfully transmitted the tactile information.

An externally validated deep learning model for the accurate segmentation of the lumbar paravertebral muscles

European Spine Journal (2022) 31:2156–2164

Imaging studies about the relevance of muscles in spinal disorders, and sarcopenia in general, require the segmentation of the muscles in the images which is very labour-intensive if performed manually and poses a practical limit to the number of investigated subjects. This study aimed at developing a deep learning-based tool able to fully automatically perform an accurate segmentation of the lumbar muscles in axial MRI scans, and at validating the new tool on an external dataset.

Methods A set of 60 axial MRI images of the lumbar spine was retrospectively collected from a clinical database. Psoas major, quadratus lumborum, erector spinae, and multifidus were manually segmented in all available slices. The dataset was used to train and validate a deep neural network able to segment muscles automatically. Subsequently, the network was externally validated on images purposely acquired from 22 healthy volunteers.

Results The median Jaccard index for the individual muscles calculated for the 22 subjects of the external validation set ranged between 0.862 and 0.935, demonstrating a generally excellent performance of the network, although occasional failures were noted. Cross-sectional area and fat fraction of the muscles were in agreement with published data.

Conclusions The externally validated deep neural network was able to perform the segmentation of the paravertebral muscles in an accurate and fully automated manner, although it is not without limitations. The model is therefore a suitable research tool to perform large-scale studies in the field of spinal disorders and sarcopenia, overcoming the limitations of non-automated methods.

Artificial intelligence in predicting early‑onset adjacent segment degeneration following anterior cervical discectomy and fusion

European Spine Journal (2022) 31:2104–2114

Anterior cervical discectomy and fusion (ACDF) is a common surgical treatment for degenerative disease in the cervical spine. However, resultant biomechanical alterations may predispose to early-onset adjacent segment degeneration (EO-ASD), which may become symptomatic and require reoperation. This study aimed to develop and validate a machine learning (ML) model to predict EO-ASD following ACDF.

Methods Retrospective review of prospectively collected data of patients undergoing ACDF at a quaternary referral medical center was performed. Patients > 18 years of age with > 6 months of follow-up and complete pre- and postoperative X-ray and MRI imaging were included. An ML-based algorithm was developed to predict EO-ASD based on preoperative demographic, clinical, and radiographic parameters, and model performance was evaluated according to discrimination and overall performance.

Results In total, 366 ACDF patients were included (50.8% male, mean age 51.4 ± 11.1 years). Over 18.7 ± 20.9 months of follow-up, 97 (26.5%) patients developed EO-ASD. The model demonstrated good discrimination and overall performance according to precision (EO-ASD: 0.70, non-ASD: 0.88), recall (EO-ASD: 0.73, non-ASD: 0.87), accuracy (0.82), F1-score (0.79), Brier score (0.203), and AUC (0.794), with C4/C5 posterior disc bulge, C4/C5 anterior disc bulge, C6 posterior superior osteophyte, presence of osteophytes, and C6/C7 anterior disc bulge identified as the most important predictive features.

Conclusions Through an ML approach, the model identified risk factors and predicted development of EO-ASD following ACDF with good discrimination and overall performance. By addressing the shortcomings of traditional statistics, ML techniques can support discovery, clinical decision-making, and precision-based spine care.

Can artificial intelligence support or even replace physicians in measuring sagittal balance?

European Spine Journal (2022) 31:1943–1951

Sagittal balance (SB) plays an important role in the surgical treatment of spinal disorders. The aim of this research study is to provide a detailed evaluation of a new, fully automated algorithm based on artificial intelligence (AI) for the determination of SB parameters on a large number of patients with and without instrumentation.

Methods Pre- and postoperative sagittal full body radiographs of 170 patients were measured by two human raters, twice by one rater and by the AI algorithm which determined: pelvic incidence, pelvic tilt, sacral slope, L1-S1 lordosis, T4-T12 thoracic kyphosis (TK) and the spino-sacral angle (SSA). To evaluate the agreement between human raters and AI, the mean error (95% confidence interval (CI)), standard deviation and an intra- and inter-rater reliability was conducted using intra-class correlation (ICC) coefficients.

Results ICC values for the assessment of the intra- (range: 0.88–0.97) and inter-rater (0.86–0.97) reliability of human raters are excellent. The algorithm is able to determine all parameters in 95% of all pre- and in 91% of all postoperative images with excellent ICC values (PreOP-range: 0.83–0.91, PostOP: 0.72–0.89). Mean errors are smallest for the SSA (PreOP: −0.1° (95%-CI: −0.9°–0.6°); PostOP: −0.5° (−1.4°–0.4°)) and largest for TK (7.0° (6.1°–7.8°); 7.1° (6.1°–8.1°)).

Conclusion A new, fully automated algorithm that determines SB parameters has excellent reliability and agreement with human raters, particularly on preoperative full spine images. The presented solution will relieve physicians from timeconsuming routine work of measuring SB parameters and allow the analysis of large databases efficiently.