Artificial intelligence–based deep learning model for evaluating procedural consistency in microvascular anastomosis

J Neurosurg 144:1–10, 2026

This study presents an LSTM-based deep learning model that objectively evaluates microvascular anastomosis performance by predicting hand-motion trajectories from MediaPipe-derived hand landmarks. It quantifies consistency using Kullback-Leibler divergence and validates complementary metrics—economy and flow of motion—comparing two expert neurosurgeons (repeat sessions) and one trainee in simulated end-to-side anastomoses.

Results show low KL divergence for experts versus higher divergence for the trainee, reflecting greater consistency and efficiency. The authors discuss methodology, model architecture choices, limitations in generalizability, and potential integration into microsurgical training workflows for objective skill assessment.

Deep Learning Model: An LSTM-based neural network was developed to objectively assess consistency and precision in microvascular anastomosis by predicting and comparing suturing hand movements using video-based hand landmark tracking, eliminating the need for physical sensors.

Hand Tracking Technology: The model utilized MediaPipe Hand Landmarker, a CNN-based system that detects 21 hand landmarks from standard video, enabling detailed, sensor-free motion analysis during microsurgical simulation.

Performance Metrics: Three primary metrics were used: Kullback-Leibler (KL) divergence for consistency, economy of motion (mean Euclidean distance of hand movement), and flow of motion (median time per suture), providing quantitative, objective evaluation of surgical skill.

Experimental Setup: Two expert neurosurgeons performed microanastomosis simulations (interrupted and continuous suturing) in two sessions one year apart, and a trainee performed the same task for comparison; all sessions were recorded and analyzed using the AI pipeline.

Results and Interpretation: Experts showed low KL divergence (high consistency) and efficient, rhythmic motion, while the trainee had higher KL divergence, longer suture intervals, and more variable motion, reflecting less developed skill.

Model Application: The approach enables rapid, automated assessment of multiple trainees using standard video equipment, supporting objective tracking of skill progression and facilitating feedback in training environments.

Model Rationale: LSTM architecture was chosen for its ability to model long-term temporal dependencies in sequential hand movement data, making it suitable for predicting surgical motion patterns over extended timeframes.

Limitations and Future Directions: Current findings are based on a small sample of experts and one trainee in a simulated environment; broader validation, metric standardization (especially for KL divergence), and extension to real operative settings are needed for generalizability.

Comprehensive Morphometric Analysis to Identify Key Neuroimaging Biomarkers for the Diagnosis of Adult Hydrocephalus Using Artificial Intelligence

Neurosurgery 96:1386–1396, 2025

This study used AI and SHAP analysis to identify five key, easily measurable 1-D neuroimaging biomarkers—FTHR, MEI, MCMI, SMLH, and CPCA—for accurately diagnosing adult non-normal pressure hydrocephalus, offering a practical, standardized, and interpretable approach to improve early detection and clinical decision-making.

• Hydrocephalus diagnosis is often inconsistent due to reliance on clinical and qualitative radiological assessments.

• This study used artificial intelligence (AI) and SHAP analysis to identify key, easily measurable neuroimaging biomarkers for adult non-normal pressure hydrocephalus (non-NPH).

• A comprehensive set of 21 morphometric features was analyzed from MRI images of 42 adult non-NPH patients and 40 healthy controls.

• Gradient Boosting was the best-performing AI classifier, achieving 0.94 accuracy and 0.97 AUC.

• Ventricular volume is the most important biomarker, but measurement is complex for clinicians.

• Five key 1-D biomarkers identified are: frontal-temporal horn ratio (FTHR), modified Evans index (MEI), modified cella media index (MCMI), sagittal maximum lateral ventricle height (SMLH), and coronal posterior callosal angle (CPCA).

• These five markers are easily measurable and provide high diagnostic accuracy, supporting practical clinical use.

• The approach addresses multicollinearity and improves diagnostic standardization, but further validation with larger datasets is needed.

Decoding Glioblastoma Heterogeneity: Neuroimaging Meets Machine Learning

Neurosurgery 96:1181–1192, 2025

This review highlights how advanced neuroimaging and machine learning, especially radiomics and deep learning models, are transforming the noninvasive diagnosis, molecular characterization, and prognosis prediction in IDH-wildtype glioblastoma, offering improved patient stratification and personalized treatment strategies while emphasizing the need for further clinical integration.

• Neuroimaging and machine learning have greatly improved diagnosis, classification, and prognosis of IDH-wildtype glioblastoma, a highly heterogeneous and aggressive brain tumor.

• Advanced MRI techniques, including diffusion tensor imaging (DTI) and radiomics, provide noninvasive insights into tumor infiltration, metabolic profiles, and microstructural changes.

• Machine learning algorithms, especially CNNs, enhance glioblastoma characterization, enabling accurate prediction of genetic mutations, IDH status, tumor subtypes, and survival outcomes.

• Radiomics extracts quantitative features from neuroimages, serving as potential biomarkers for tumor classification, prognosis, and guiding treatment strategies.

• Integration of radiomics and machine learning helps differentiate pseudoprogression from true tumor progression and predicts patterns of tumor invasion and recurrence.

• Imaging biomarkers and machine learning models are promising but remain complementary to molecular diagnostics and are not yet standard in clinical practice.

• Ongoing research aims to refine models, integrate emerging imaging techniques, and better link imaging features to underlying molecular processes for personalized therapy.

• The synergy of neuroimaging and AI is expected to enable noninvasive, precision management and better outcomes for glioblastoma patients.

Artificial intelligence as a modality to enhance the readability of neurosurgical literature for patients

J Neurosurg 142:1189–1195, 2025

The study evaluates ChatGPT 3.5 and GPT4’s ability to generate readable, accurate summaries of neurosurgical literature, enhancing patient comprehension. GPT4 showed higher readability and accuracy, suggesting its potential in improving patient education and bridging the gap between medical findings and public understanding.

Study Overview

Objective: Assess ChatGPT’s ability to generate readable, accurate neurosurgical summaries.

Methods: Analyzed 150 abstracts from top neurosurgical journals.

Models Used: GPT3.5 and GPT4.

Findings

Readability Improvement: GPT4 summaries more readable than original abstracts.

Scientific Accuracy: 84.2% of GPT4 summaries maintained moderate accuracy.

Readability Metrics: GPT4 outperformed GPT3.5 in multiple readability scores.

Implications

Patient Education: GPT4 can enhance neurosurgical literature comprehension for patients.

Health Literacy: Potential to improve health literacy nationwide.

Limitations and Future Research

Accessibility: GPT4’s restricted access limits broader application.

Future Studies: Explore GPT4’s use in other medical specialties.

Microscope-integrated optical coherence tomography for in vivo human brain tumor detection with artificial intelligence

J Neurosurg 141:1343–1351, 2024

It has been shown that optical coherence tomography (OCT) can identify brain tumor tissue and potentially be used for intraoperative margin diagnostics. However, there is limited evidence on its use in human in vivo settings, particularly in terms of its applicability and accuracy of residual brain tumor detection (RTD). For this reason, a microscope-integrated OCT system was examined to determine in vivo feasibility of RTD after resection with automated scan analysis.

METHODS Healthy and diseased brain was 3D scanned at the resection edge in 18 brain tumor patients and investigated for its informative value in regard to intraoperative tissue classification. Biopsies were taken at these locations and labeled by a neuropathologist for further analysis as ground truth. Optical OCT properties were obtained, compared, and used for separation with machine learning. In addition, two artificial intelligence–assisted methods were utilized for scan classification, and all approaches were examined for RTD accuracy and compared to standard techniques.

RESULTS In vivo OCT tissue scanning was feasible and easily integrable into the surgical workflow. Measured backscattered light signal intensity, signal attenuation, and signal homogeneity were significantly distinctive in the comparison of scanned white matter to increasing levels of scanned tumor infiltration (p < 0.001) and achieved high values of accuracy (85%) for the detection of diseased brain in the tumor margin with support vector machine separation. A neuronal network approach achieved 82% accuracy and an autoencoder approach 85% accuracy in the detection of diseased brain in the tumor margin. Differentiating cortical gray matter from tumor tissue was not technically feasible in vivo.

CONCLUSIONS In vivo OCT scanning of the human brain has been shown to contain significant value for intraoperative RTD, supporting what has previously been discussed for ex vivo OCT brain tumor scanning, with the perspective of complementing current intraoperative methods for this purpose, especially when deciding to withdraw from further resection toward the end of the surgery.

Deep learning algorithm for fully automated measurement of sagittal balance in adult spinal deformity

European Spine Journal (2024) 33:4119–4124

Deep learning (DL) algorithms can be used for automated analysis of medical imaging. The aim of this study was to assess the accuracy of an innovative, fully automated DL algorithm for analysis of sagittal balance in adult spinal deformity (ASD).

Material and methods Sagittal balance (sacral slope, pelvic tilt, pelvic incidence, lumbar lordosis and sagittal vertical axis) was evaluated in 141 preoperative and postoperative radiographs of patients with ASD. The DL, landmark-based measurements, were compared with the ground truth values from validated manual measurements.

Results The DL algorithm showed an excellent consistency with the ground truth measurements. The intra-class correlation coefficient between the DL and ground truth measurements was 0.71–0.99 for preoperative and 0.72–0.96 for postoperative measurements. The DL detection rate was 91.5% and 84% for preoperative and postoperative images, respectively.

Conclusion This is the first study evaluating a complete automated DL algorithm for analysis of sagittal balance with high accuracy for all evaluated parameters. The excellent accuracy in the challenging pathology of ASD with long construct instrumentation demonstrates the eligibility and possibility for implementation in clinical routine.

A Quantitative Assessment of ChatGPT as a Neurosurgical Triaging Tool

Neurosurgery 95:487–495, 2024

ChatGPT is a natural language processing chatbot with increasing applicability to the medical workflow. Although ChatGPT has been shown to be capable of passing the American Board of Neurological Surgery board examination, there has never been an evaluation of the chatbot in triaging and diagnosing novel neurosurgical scenarios without defined answer choices. In this study, we assess ChatGPT’s capability to determine the emergent nature of neurosurgical scenarios and make diagnoses based on information one would find in a neurosurgical consult.

METHODS: Thirty clinical scenarios were given to 3 attendings, 4 residents, 2 physician assistants, and 2 subinterns. Participants were asked to determine if the scenario constituted an urgent neurosurgical consultation and what the most likely diagnosis was. Attending responses provided a consensus to use as the answer key. Generative pretraining transformer (GPT) 3.5 and GPT 4 were given the same questions, and their responses were compared with the other participants.

RESULTS: GPT 4 was 100% accurate in both diagnosis and triage of the scenarios. GPT 3.5 had an accuracy of 92.59%, slightly below that of a PGY1 (96.3%), an 88.24% sensitivity, 100% specificity, 100% positive predictive value, and 83.3% negative predicative value in triaging each situation. When making a diagnosis, GPT 3.5 had an accuracy of 92.59%, which was higher than the subinterns and similar to resident responders.

CONCLUSION: GPT 4 is able to diagnose and triage neurosurgical scenarios at the level of a senior neurosurgical resident. There has been a clear improvement between GPT 3.5 and 4. It is likely that the recent updates in internet access and directing the functionality of ChatGPT will further improve its utility in neurosurgical triage.

Development and validation of an artificial intelligence model to accurately predict spinopelvic parameters

J Neurosurg Spine 41:88–96, 2024

Achieving appropriate spinopelvic alignment has been shown to be associated with improved clinical symptoms. However, measurement of spinopelvic radiographic parameters is time-intensive and interobserver reliability is a concern. Automated measurement tools have the promise of rapid and consistent measurements, but existing tools are still limited to some degree by manual user-entry requirements. This study presents a novel artificial intelligence (AI) tool called SpinePose that automatically predicts spinopelvic parameters with high accuracy without the need for manual entry.

METHODS SpinePose was trained and validated on 761 sagittal whole-spine radiographs to predict the sagittal vertical axis (SVA), pelvic tilt (PT), pelvic incidence (PI), sacral slope (SS), lumbar lordosis (LL), T1 pelvic angle (T1PA), and L1 pelvic angle (L1PA). A separate test set of 40 radiographs was labeled by four reviewers, including fellowship-trained spine surgeons and a fellowship-trained radiologist with neuroradiology subspecialty certification. Median errors relative to the most senior reviewer were calculated to determine model accuracy on test images. Intraclass correlation coefficients (ICCs) were used to assess interrater reliability.

RESULTS SpinePose exhibited the following median (interquartile range) parameter errors: SVA 2.2 mm (2.3 mm) (p = 0.93), PT 1.3° (1.2°) (p = 0.48), SS 1.7° (2.2°) (p = 0.64), PI 2.2° (2.1°) (p = 0.24), LL 2.6° (4.0°) (p = 0.89), T1PA 1.1° (0.9°) (p = 0.42), and L1PA 1.4° (1.6°) (p = 0.49). Model predictions also exhibited excellent reliability at all parameters (ICC 0.91–1.0).

CONCLUSIONS SpinePose accurately predicted spinopelvic parameters with excellent reliability comparable to that of fellowship-trained spine surgeons and neuroradiologists. Utilization of predictive AI tools in spinal imaging can substantially aid in patient selection and surgical planning.

Novel artificial intelligence algorithm: an accurate and independent measure of spinopelvic parameters

J Neurosurg Spine 37:893–901, 2022

The analysis of sagittal alignment by measuring spinopelvic parameters has been widely adopted among spine surgeons globally, and sagittal imbalance is a well-documented cause of poor quality of life. These measurements are time-consuming but necessary to make, which creates a growing need for an automated analysis tool that measures spinopelvic parameters with speed, precision, and reproducibility without relying on user input. This study introduces and evaluates an algorithm based on artificial intelligence (AI) that fully automatically measures spinopelvic parameters.

METHODS Two hundred lateral lumbar radiographs (pre- and postoperative images from 100 patients undergoing lumbar fusion) were retrospectively analyzed by board-certified spine surgeons who digitally measured lumbar lordosis, pelvic incidence, pelvic tilt, and sacral slope. The novel AI algorithm was also used to measure the same parameters. To evaluate the agreement between human and AI-automated measurements, the mean error (95% CI, SD) was calculated and interrater reliability was assessed using the 2-way random single-measure intraclass correlation coefficient (ICC). ICC values larger than 0.75 were considered excellent.

RESULTS The AI algorithm determined all parameters in 98% of preoperative and in 95% of postoperative images with excellent ICC values (preoperative range 0.85–0.92, postoperative range 0.81–0.87). The mean errors were smallest for pelvic incidence both pre- and postoperatively (preoperatively −0.5° [95% CI −1.5° to 0.6°] and postoperatively 0.0° [95% CI −1.1° to 1.2°]) and largest preoperatively for sacral slope (−2.2° [95% CI −3.0° to −1.5°]) and postoperatively for lumbar lordosis (3.8° [95% CI 2.5° to 5.0°]).

CONCLUSIONS Advancements in AI translate to the arena of medical imaging analysis. This method of measuring spinopelvic parameters on spine radiographs has excellent reliability comparable to expert human raters. This application allows users to accurately obtain critical spinopelvic measurements automatically, which can be applied to clinical practice. This solution can assist physicians by saving time in routine work and by avoiding error-prone manual measurements.

Artificial intelligence in predicting early‑onset adjacent segment degeneration following anterior cervical discectomy and fusion

European Spine Journal (2022) 31:2104–2114

Anterior cervical discectomy and fusion (ACDF) is a common surgical treatment for degenerative disease in the cervical spine. However, resultant biomechanical alterations may predispose to early-onset adjacent segment degeneration (EO-ASD), which may become symptomatic and require reoperation. This study aimed to develop and validate a machine learning (ML) model to predict EO-ASD following ACDF.

Methods Retrospective review of prospectively collected data of patients undergoing ACDF at a quaternary referral medical center was performed. Patients > 18 years of age with > 6 months of follow-up and complete pre- and postoperative X-ray and MRI imaging were included. An ML-based algorithm was developed to predict EO-ASD based on preoperative demographic, clinical, and radiographic parameters, and model performance was evaluated according to discrimination and overall performance.

Results In total, 366 ACDF patients were included (50.8% male, mean age 51.4 ± 11.1 years). Over 18.7 ± 20.9 months of follow-up, 97 (26.5%) patients developed EO-ASD. The model demonstrated good discrimination and overall performance according to precision (EO-ASD: 0.70, non-ASD: 0.88), recall (EO-ASD: 0.73, non-ASD: 0.87), accuracy (0.82), F1-score (0.79), Brier score (0.203), and AUC (0.794), with C4/C5 posterior disc bulge, C4/C5 anterior disc bulge, C6 posterior superior osteophyte, presence of osteophytes, and C6/C7 anterior disc bulge identified as the most important predictive features.

Conclusions Through an ML approach, the model identified risk factors and predicted development of EO-ASD following ACDF with good discrimination and overall performance. By addressing the shortcomings of traditional statistics, ML techniques can support discovery, clinical decision-making, and precision-based spine care.

Can artificial intelligence support or even replace physicians in measuring sagittal balance?

European Spine Journal (2022) 31:1943–1951

Sagittal balance (SB) plays an important role in the surgical treatment of spinal disorders. The aim of this research study is to provide a detailed evaluation of a new, fully automated algorithm based on artificial intelligence (AI) for the determination of SB parameters on a large number of patients with and without instrumentation.

Methods Pre- and postoperative sagittal full body radiographs of 170 patients were measured by two human raters, twice by one rater and by the AI algorithm which determined: pelvic incidence, pelvic tilt, sacral slope, L1-S1 lordosis, T4-T12 thoracic kyphosis (TK) and the spino-sacral angle (SSA). To evaluate the agreement between human raters and AI, the mean error (95% confidence interval (CI)), standard deviation and an intra- and inter-rater reliability was conducted using intra-class correlation (ICC) coefficients.

Results ICC values for the assessment of the intra- (range: 0.88–0.97) and inter-rater (0.86–0.97) reliability of human raters are excellent. The algorithm is able to determine all parameters in 95% of all pre- and in 91% of all postoperative images with excellent ICC values (PreOP-range: 0.83–0.91, PostOP: 0.72–0.89). Mean errors are smallest for the SSA (PreOP: −0.1° (95%-CI: −0.9°–0.6°); PostOP: −0.5° (−1.4°–0.4°)) and largest for TK (7.0° (6.1°–7.8°); 7.1° (6.1°–8.1°)).

Conclusion A new, fully automated algorithm that determines SB parameters has excellent reliability and agreement with human raters, particularly on preoperative full spine images. The presented solution will relieve physicians from timeconsuming routine work of measuring SB parameters and allow the analysis of large databases efficiently.

Deep Learning for Outcome Prediction in Neurosurgery

Neurosurgery 90:16–38, 2022

Deep learning (DL) is a powerful machine learning technique that has increasingly been used to predict surgical outcomes. However, the large quantity of data required and lack of model interpretability represent substantial barriers to the validity and reproducibility of DL models.

The objective of this study was to systematically review the characteristics of DL studies involving neurosurgical outcome prediction and to assess their bias and reporting quality.

Literature search using the PubMed, Scopus, and Embase databases identified 1949 records of which 35 studies were included. Of these, 32 (91%) developed and validated a DL model while 3 (9%) validated a pre-existing model. The most commonly represented subspecialty areas were oncology (16 of 35, 46%), spine (8 of 35, 23%), and vascular (6 of 35, 17%). Risk of bias was low in 18 studies (51%), unclear in 5 (14%), and high in 12 (34%), most commonly because of data quality deficiencies.

Adherence to transparent reporting of a multivariable prediction model for individual prognosis or diagnosis reporting standards was low, with a median of 12 transparent reporting of a multivariable prediction model for individual prognosis or diagnosis items (39%) per study not reported. Model transparency was severely limited because code was provided in only 3 studies (9%) and final models in 2 (6%).

With the exception of public databases, no study data sets were readily available. No studies described DL models as ready for clinical use. The use of DL for neurosurgical outcome prediction remains nascent. Lack of appropriate data sets poses a major concern for bias. Although studies have demonstrated promising results, greater transparency in model development and reporting is needed to facilitate reproducibility and validation.

Machine Learning for the Prediction of Molecular Markers in Glioma on Magnetic Resonance Imaging: A Systematic Review and Meta-Analysis

Neurosurgery 89:31–44, 2021

Molecular characterization of glioma has implications for prognosis, treatment planning, and prediction of treatment response. Current histopathology is limited by intratumoral heterogeneity and variability in detection methods. Advances in computational techniques have led to interest in mining quantitative imaging features to noninvasively detect genetic mutations.

OBJECTIVE: To evaluate the diagnostic accuracy of machine learning (ML) models in molecular subtyping gliomas on preoperative magnetic resonance imaging (MRI).

METHODS: A systematic search was performed following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analysis) guidelines to identify studies up to April 1, 2020. Methodological quality of studies was assessed using the Quality Assessment for Diagnostic Accuracy Studies (QUADAS)-2. Diagnostic performance estimates were obtained using a bivariate model and heterogeneity was explored using metaregression.

RESULTS: Forty-four original articles were included. The pooled sensitivity and specificity for predicting isocitrate dehydrogenase (IDH) mutation in training datasets were 0.88 (95% CI 0.83-0.91) and 0.86 (95% CI 0.79-0.91), respectively, and 0.83 to 0.85 in validation sets. Use of data augmentation and MRI sequence type were weakly associated with heterogeneity. Both O6-methylguanine-DNA methyltransferase (MGMT) gene promoter methylation and 1p/19q codeletion could be predicted with a pooled sensitivity and specificity between 0.76 and 0.83 in training datasets.

CONCLUSION: ML application to preoperative MRI demonstrated promising results for predicting IDHmutation, MGMT methylation, and 1p/19q codeletion in glioma. Optimized ML models could lead to a noninvasive, objective tool that captures molecular information important for clinical decisionmaking. Future studies should use multicenter data, external validation and investigate clinical feasibility of ML models.

An Online Calculator for the Prediction of Survival in Glioblastoma Patients Using Classical Statistics and Machine Learning

Neurosurgery, Volume 86, Issue 2, February 2020, Pages E184–E192

Although survival statistics in patients with glioblastoma multiforme (GBM) are well-defined at the group level, predicting individual patient survival remains challenging because of significant variation within strata.

OBJECTIVE: To compare statistical and machine learning algorithms in their ability to predict survival in GBM patients and deploy the best performing model as an online survival calculator.

METHODS: Patients undergoing an operation for a histopathologically confirmed GBM were extracted from the Surveillance Epidemiology and End Results (SEER) database (2005-2015) and split into a training and hold-out test set in an 80/20 ratio. Fifteen statistical and machine learning algorithms were trained based on 13 demographic, socioeconomic, clinical, and radiographic features to predict overall survival, 1-yr survival status, and compute personalized survival curves.

RESULTS: In total, 20 821 patients met our inclusion criteria. The accelerated failure time model demonstrated superior performance in terms of discrimination (concordance index = 0.70), calibration, interpretability, predictive applicability, and computational efficiency compared to Cox proportional hazards regression and other machine learning algorithms. This model was deployed through a free, publicly available software interface (https://cnoc-bwh.shinyapps.io/gbmsurvivalpredictor/).

CONCLUSION: The development and deployment of survival prediction tools require a multimodal assessment rather than a single metric comparison. This study provides a framework for the development of prediction tools in cancer patients, as well as an online survival calculator for patients with GBM. Future efforts should improve the interpretability, predictive applicability, and computational efficiency of existing machine learning algorithms, increase the granularity of population-based registries, and externally validate the proposed prediction tool.