Predicting Intracranial Pressure Levels: A Deep Learning Approach Using Computed Tomography Brain Scans

Neurosurgery 98:256–268, 2026

This clinical study evaluates deep learning models that predict whether intracranial pressure (ICP) exceeds 15 mm Hg from brain CT scans, integrating demographic and Glasgow Coma Scale data into image inputs. Four 3D architectures—including MobileNetV2 3D and DenseNet201 3D—were trained on 578 paired CT–ICP cases with preprocessing, augmentation, and explainability via class activation maps.

Results show MobileNetV2 3D achieved the best generalization (AUC 0.883, recall 81.8%), with demographic embedding improving performance; limitations include single-center data, class imbalance, and lack of external validation, and authors recommend multicenter expansion and refined region-specific feature extraction before clinical deployment.

Intracranial Pressure (ICP) Risk: Elevated ICP is a critical, potentially fatal condition requiring rapid diagnosis and intervention, but current gold-standard invasive monitoring methods carry risks and are not always feasible in emergency settings.

Noninvasive ICP Assessment Challenge: Existing noninvasive methods (e.g., CT-based qualitative markers) lack sufficient accuracy and reliability for routine emergency use, highlighting the need for improved approaches.

Deep Learning Solution: Four deep learning models were trained on a custom dataset of 578 paired brain CT scans, demographic information, and Glasgow Coma Scale (GCS) scores to classify whether ICP exceeds 15 mm Hg, addressing the gap in noninvasive, rapid ICP estimation.

Data Integration Innovation: Demographic and GCS data were embedded and merged with CT imaging, creating a multimodal input that improved model performance compared to imaging-only approaches.

Best Model Performance: The MobileNetV2 3D model with demographic data achieved the highest test AUC of 88.3% and recall of 81.8%, outperforming other architectures and showing promise for high-sensitivity emergency applications.

Explainability: Class Activation Maps (CAMs) were used to visualize which regions of the brain CT scans influenced model predictions, enhancing transparency and interpretability of the AI system.

Limitations: The study’s main limitations include a relatively small, single-center dataset with class imbalance, lack of external/multicenter validation, and potential inconsistencies due to timing mismatches between CT and ICP measurements.

Clinical Impact & Future Directions: This AI approach could reduce reliance on invasive monitoring and accelerate ICP triage in neurocritical care; further multicenter studies, prospective validation, and expansion to multiclass classification are needed for clinical deployment.

Artificial intelligence–based deep learning model for evaluating procedural consistency in microvascular anastomosis

J Neurosurg 144:1–10, 2026

This study presents an LSTM-based deep learning model that objectively evaluates microvascular anastomosis performance by predicting hand-motion trajectories from MediaPipe-derived hand landmarks. It quantifies consistency using Kullback-Leibler divergence and validates complementary metrics—economy and flow of motion—comparing two expert neurosurgeons (repeat sessions) and one trainee in simulated end-to-side anastomoses.

Results show low KL divergence for experts versus higher divergence for the trainee, reflecting greater consistency and efficiency. The authors discuss methodology, model architecture choices, limitations in generalizability, and potential integration into microsurgical training workflows for objective skill assessment.

Deep Learning Model: An LSTM-based neural network was developed to objectively assess consistency and precision in microvascular anastomosis by predicting and comparing suturing hand movements using video-based hand landmark tracking, eliminating the need for physical sensors.

Hand Tracking Technology: The model utilized MediaPipe Hand Landmarker, a CNN-based system that detects 21 hand landmarks from standard video, enabling detailed, sensor-free motion analysis during microsurgical simulation.

Performance Metrics: Three primary metrics were used: Kullback-Leibler (KL) divergence for consistency, economy of motion (mean Euclidean distance of hand movement), and flow of motion (median time per suture), providing quantitative, objective evaluation of surgical skill.

Experimental Setup: Two expert neurosurgeons performed microanastomosis simulations (interrupted and continuous suturing) in two sessions one year apart, and a trainee performed the same task for comparison; all sessions were recorded and analyzed using the AI pipeline.

Results and Interpretation: Experts showed low KL divergence (high consistency) and efficient, rhythmic motion, while the trainee had higher KL divergence, longer suture intervals, and more variable motion, reflecting less developed skill.

Model Application: The approach enables rapid, automated assessment of multiple trainees using standard video equipment, supporting objective tracking of skill progression and facilitating feedback in training environments.

Model Rationale: LSTM architecture was chosen for its ability to model long-term temporal dependencies in sequential hand movement data, making it suitable for predicting surgical motion patterns over extended timeframes.

Limitations and Future Directions: Current findings are based on a small sample of experts and one trainee in a simulated environment; broader validation, metric standardization (especially for KL divergence), and extension to real operative settings are needed for generalizability.

Deep learning–based segmentation of the trigeminal nerve and surrounding vasculature in trigeminal neuralgia

J Neurosurg 143:83–91, 2025

This study developed and validated deep learning U-Net models for automated 3D segmentation of the trigeminal nerve and surrounding vasculature in MRI of trigeminal neuralgia patients, enabling objective quantification of neurovascular conflict features and potentially improving preoperative evaluation and treatment planning.

• Deep learning (U-Net) models were used to segment the trigeminal nerve and surrounding vasculature in patients with trigeminal neuralgia using high-resolution CISS MRI.

• Six U-Net variants with different encoder backbones were tested; SE-ResNet50 performed best overall (Dice score = 0.775, IoU = 0.681).

• The models quantified anatomical features such as the surface area of neurovascular contact and distance to the contact point, showing no significant difference from manual segmentations.

• The best model achieved 100% sensitivity and specificity in detecting neurovascular conflict in the testing set.

• Automated 3D segmentation allows for objective, quantitative evaluation, improving on subjective and time-intensive manual methods.

• Limitations include inability to distinguish vessel type (artery vs. vein) and data from a single institution; future work should address these.

• The method may help standardize neurovascular conflict assessment and improve treatment selection for trigeminal neuralgia.

Deep learning algorithm for fully automated measurement of sagittal balance in adult spinal deformity

European Spine Journal (2024) 33:4119–4124

Deep learning (DL) algorithms can be used for automated analysis of medical imaging. The aim of this study was to assess the accuracy of an innovative, fully automated DL algorithm for analysis of sagittal balance in adult spinal deformity (ASD).

Material and methods Sagittal balance (sacral slope, pelvic tilt, pelvic incidence, lumbar lordosis and sagittal vertical axis) was evaluated in 141 preoperative and postoperative radiographs of patients with ASD. The DL, landmark-based measurements, were compared with the ground truth values from validated manual measurements.

Results The DL algorithm showed an excellent consistency with the ground truth measurements. The intra-class correlation coefficient between the DL and ground truth measurements was 0.71–0.99 for preoperative and 0.72–0.96 for postoperative measurements. The DL detection rate was 91.5% and 84% for preoperative and postoperative images, respectively.

Conclusion This is the first study evaluating a complete automated DL algorithm for analysis of sagittal balance with high accuracy for all evaluated parameters. The excellent accuracy in the challenging pathology of ASD with long construct instrumentation demonstrates the eligibility and possibility for implementation in clinical routine.

Generation and applications of synthetic computed tomography images for neurosurgical planning

J Neurosurg 141:742–751, 2024

CT and MRI are synergistic in the information provided for neurosurgical planning. While obtaining both types of images lends unique data from each, doing so adds to cost and exposes patients to additional ionizing radiation after MRI has been performed. Cross-modal synthesis of high-resolution CT images from MRI sequences offers an appealing solution. The authors therefore sought to develop a deep learning conditional generative adversarial network (cGAN) which performs this synthesis.

METHODS Preoperative paired CT and contrast-enhanced MR images were collected for patients with meningioma, pituitary tumor, vestibular schwannoma, and cerebrovascular disease. CT and MR images were denoised, field corrected, and coregistered. MR images were fed to a cGAN that exported a “synthetic” CT scan. The accuracy of synthetic CT images was assessed objectively using the quantitative similarity metrics as well as by clinical features such as sella and internal auditory canal (IAC) dimensions and mastoid/clinoid/sphenoid aeration.

RESULTS A total of 92,981 paired CT/MR images obtained in 80 patients were used for training/testing, and 10,068 paired images from 10 patients were used for external validation. Synthetic CT images reconstructed the bony skull base and convexity with relatively high accuracy. Measurements of the sella and IAC showed a median relative error between synthetic CT scans and ground truth images of 6%, with greater variability in IAC reconstruction compared with the sella. Aerations in the mastoid, clinoid, and sphenoid regions were generally captured, although there was heterogeneity in finer air cell septations. Performance varied based on pathology studied, with the highest limitation observed in evaluating meningiomas with intratumoral calcifications or calvarial invasion.

CONCLUSIONS The generation of high-resolution CT scans from MR images through cGAN offers promise for a wide range of applications in cranial and spinal neurosurgery, especially as an adjunct for preoperative evaluation. Optimizing cGAN performance on specific anatomical regions may increase its clinical viability.

An externally validated deep learning model for the accurate segmentation of the lumbar paravertebral muscles

European Spine Journal (2022) 31:2156–2164

Imaging studies about the relevance of muscles in spinal disorders, and sarcopenia in general, require the segmentation of the muscles in the images which is very labour-intensive if performed manually and poses a practical limit to the number of investigated subjects. This study aimed at developing a deep learning-based tool able to fully automatically perform an accurate segmentation of the lumbar muscles in axial MRI scans, and at validating the new tool on an external dataset.

Methods A set of 60 axial MRI images of the lumbar spine was retrospectively collected from a clinical database. Psoas major, quadratus lumborum, erector spinae, and multifidus were manually segmented in all available slices. The dataset was used to train and validate a deep neural network able to segment muscles automatically. Subsequently, the network was externally validated on images purposely acquired from 22 healthy volunteers.

Results The median Jaccard index for the individual muscles calculated for the 22 subjects of the external validation set ranged between 0.862 and 0.935, demonstrating a generally excellent performance of the network, although occasional failures were noted. Cross-sectional area and fat fraction of the muscles were in agreement with published data.

Conclusions The externally validated deep neural network was able to perform the segmentation of the paravertebral muscles in an accurate and fully automated manner, although it is not without limitations. The model is therefore a suitable research tool to perform large-scale studies in the field of spinal disorders and sarcopenia, overcoming the limitations of non-automated methods.

Can artificial intelligence support or even replace physicians in measuring sagittal balance?

European Spine Journal (2022) 31:1943–1951

Sagittal balance (SB) plays an important role in the surgical treatment of spinal disorders. The aim of this research study is to provide a detailed evaluation of a new, fully automated algorithm based on artificial intelligence (AI) for the determination of SB parameters on a large number of patients with and without instrumentation.

Methods Pre- and postoperative sagittal full body radiographs of 170 patients were measured by two human raters, twice by one rater and by the AI algorithm which determined: pelvic incidence, pelvic tilt, sacral slope, L1-S1 lordosis, T4-T12 thoracic kyphosis (TK) and the spino-sacral angle (SSA). To evaluate the agreement between human raters and AI, the mean error (95% confidence interval (CI)), standard deviation and an intra- and inter-rater reliability was conducted using intra-class correlation (ICC) coefficients.

Results ICC values for the assessment of the intra- (range: 0.88–0.97) and inter-rater (0.86–0.97) reliability of human raters are excellent. The algorithm is able to determine all parameters in 95% of all pre- and in 91% of all postoperative images with excellent ICC values (PreOP-range: 0.83–0.91, PostOP: 0.72–0.89). Mean errors are smallest for the SSA (PreOP: −0.1° (95%-CI: −0.9°–0.6°); PostOP: −0.5° (−1.4°–0.4°)) and largest for TK (7.0° (6.1°–7.8°); 7.1° (6.1°–8.1°)).

Conclusion A new, fully automated algorithm that determines SB parameters has excellent reliability and agreement with human raters, particularly on preoperative full spine images. The presented solution will relieve physicians from timeconsuming routine work of measuring SB parameters and allow the analysis of large databases efficiently.

Deep Learning for Outcome Prediction in Neurosurgery

Neurosurgery 90:16–38, 2022

Deep learning (DL) is a powerful machine learning technique that has increasingly been used to predict surgical outcomes. However, the large quantity of data required and lack of model interpretability represent substantial barriers to the validity and reproducibility of DL models.

The objective of this study was to systematically review the characteristics of DL studies involving neurosurgical outcome prediction and to assess their bias and reporting quality.

Literature search using the PubMed, Scopus, and Embase databases identified 1949 records of which 35 studies were included. Of these, 32 (91%) developed and validated a DL model while 3 (9%) validated a pre-existing model. The most commonly represented subspecialty areas were oncology (16 of 35, 46%), spine (8 of 35, 23%), and vascular (6 of 35, 17%). Risk of bias was low in 18 studies (51%), unclear in 5 (14%), and high in 12 (34%), most commonly because of data quality deficiencies.

Adherence to transparent reporting of a multivariable prediction model for individual prognosis or diagnosis reporting standards was low, with a median of 12 transparent reporting of a multivariable prediction model for individual prognosis or diagnosis items (39%) per study not reported. Model transparency was severely limited because code was provided in only 3 studies (9%) and final models in 2 (6%).

With the exception of public databases, no study data sets were readily available. No studies described DL models as ready for clinical use. The use of DL for neurosurgical outcome prediction remains nascent. Lack of appropriate data sets poses a major concern for bias. Although studies have demonstrated promising results, greater transparency in model development and reporting is needed to facilitate reproducibility and validation.