J Neurosurg Spine 44:847–857, 2026
This clinical article presents a fully automated, interpretable three-stage deep learning pipeline for detecting and grading lumbar spinal stenosis (LSS) using axial T2-weighted MRI. The framework integrates region classification, YOLO-based ROI detection, and CNN-based severity grading, validated on internal (640 patients, 17,440 slices) and external (515 patients, 8,000 slices) datasets with high accuracy and explainability via Grad-CAM.
The study details dataset curation, model architectures (ResNet-18, RegNetX-400MF, EfficientNet-B0, YOLOv5/8), training protocols, evaluation metrics, and clinical implementation pathways, highlighting strengths, limitations (single-rater labels, class imbalance, 2D slice analysis), and future directions toward volumetric and multi-expert validation.
Objective Standardize and automate lumbar spinal stenosis (LSS) identification, classification, and grading from axial T2-weighted lumbar MRI to reduce diagnostic variability.
Pipeline Three-stage cascade: (1) classify slices into sacral/lumbar/thoracic regions, (2) detect and crop anatomical ROIs, (3) grade LSS as binary or multiclass severity.
Datasets Internal training set: 640 patients with 17,440 retained axial T2 slices; external validation set: 8000 preprocessed, neurosurgeon-graded axial slices from an open-access dataset (515 patients).
Grading scheme Labels follow Schizas central canal stenosis grades A–D; a binary version groups A+B as nonstenotic and C+D as clinically significant stenosis.
Models Lightweight CNN backbones (ResNet-18, RegNetX-400MF, EfficientNet-B0) used for stages 1 and 3; YOLOv5/YOLOv8 used for ROI detection.
Validation approach Patient-level splits with 10-fold cross-validation to reduce overfitting and data leakage; an independent internal test set (62 patients, 1679 slices) reserved for final evaluation.
Performance Achieved 97.87% accuracy for binary LSS grading and 95.52% accuracy for multiclass grading, outperforming prior models in this setting.
Interpretability & clinical aim Grad-CAM heat maps highlight regions influencing predictions to support trust and potential workflow integration as an interpretable decision-support tool.

You must be logged in to post a comment.