Neurosurgery 97:1072–1082, 2025
This systematic review and meta-analysis evaluates machine learning (ML) applications for predicting intracranial aneurysm rupture, comparing 124 ML models across 36 retrospective studies (22,462 patients) with the PHASES score. Results show ML—especially deep learning and SVM—achieves higher AUC and specificity than PHASES, with hemodynamic inputs improving test-set specificity but not external validation.
The authors highlight methodological heterogeneity, risks of bias, and overfitting concerns from retrospective single‑center data, urging prospective, standardized studies and external validation before clinical integration of ML rupture‑risk tools.
Machine Learning (ML) Models: ML techniques, including deep learning (DL), support vector machines (SVM), and regression models, show higher specificity and overall diagnostic accuracy than the traditional PHASES score for predicting intracranial aneurysm rupture risk, with comparable sensitivity.
• Deep Learning Performance: DL models achieved the highest sensitivity (up to 0.87), specificity (up to 0.86), and area under the curve (AUC-ROC up to 0.92) among all ML families, indicating strong discriminative ability in rupture risk prediction.
• PHASES Score Limitations: The PHASES score, though widely used, demonstrates lower specificity (0.51) and modest overall discriminative ability (AUC-ROC 0.66), and does not incorporate important risk factors like aneurysm morphology or family history.
• Hemodynamic Parameters: Incorporating hemodynamic variables (e.g., wall shear stress, flow patterns) into ML models improves specificity and accuracy in test sets, but benefits are less pronounced in external validation, possibly due to sample size and generalizability issues.
• Retrospective Data and Overfitting: All included ML models were trained on retrospective, post-rupture data, raising concerns about overfitting and the applicability of these models to pre-rupture clinical decision-making.
• Generalizability Concerns: ML models often perform less well on external validation data due to biases in patient selection, single-center data, and differences in imaging or clinical protocols, while the PHASES score maintains more consistent performance across settings.
• Need for Prospective Validation: There is a critical need for prospective studies and standardized protocols to confirm the clinical utility and reliability of ML-based rupture risk prediction models before integration into routine practice.
• Clinical Implications: ML approaches, especially DL and SVM, have the potential to enhance individualized risk stratification and reduce overtreatment, but methodological challenges and validation in diverse populations remain essential for safe clinical adoption.



You must be logged in to post a comment.