Abstract
-
Objective
To develop and externally validate a dual-mechanism deep learning (DL) model that integrates vertebral segmentation and lesion detection for automated evaluation of lumbar degeneration and structured report generation on plain radiographs.
-
Methods
In this retrospective study, 5,964 patients who underwent standing anteroposterior and lateral lumbar radiographs at a single institution and 600 patients from a public dataset (BUU-Spine) were included. Vertebral corners from T11–L5 (and S1 on lateral views) and 7 degenerative findings (scoliosis, straightened/preserved lordosis, spondylolisthesis, disc space narrowing, osteophytes, vertebral compression, and abdominal aortic calcification) were annotated by 3 spine surgeons. Two independently trained, parallel networks were developed, including a ResNet-based segmentation network and a YOLOv8-based detection network. A rule-based integration strategy reconciled both outputs and generated structured diagnostic reports. Segmentation accuracy, quantitative measurement agreement, diagnostic performance, and clinical acceptability of reports were evaluated.
-
Results
Intra- and interobserver landmark distances within 3 mm reached 96% and >95%, respectively. On the internal test set, the percentage of correct keypoints within 3 mm was 95.7%–98.6%, with intraclass correlation coefficients of 0.84–0.89 and Pearson correlation coefficient (r) of 0.90–0.94 for key radiographic parameters. The segmentation- and detection-based models achieved precision of 92.2%–96.9% and 91.7%–95.5%, and recall of 91.6%–94.8% and 93.3%–95.2%, respectively. Under the dual-positive condition, the integrated model yielded the highest precision (93.8%–97.3%), whereas the any-positive condition achieved the highest recall (94.1%–97.6%). Of 596 automatically generated structured reports, 557 (93.4%) were deemed clinically acceptable.
-
Conclusion
The proposed dual-mechanism DL framework enables accurate, multilesion assessment of lumbar degeneration and generation of clinically acceptable structured reports from plain radiographs, supporting workflow optimization in lumbar spine imaging.
-
Keywords: Lumbar spine, Deep learning, Radiography, Vertebral segmentation, Artificial intelligence
INTRODUCTION
Low back pain is a major global public health problem [
1,
2]. Lumbar degeneration is a primary potential pain generator, including spinal instability, intervertebral disc degeneration, facet joint arthritis, and paraspinal muscle atrophy [
2]. Lumbar plain radiography is a convenient and cost-effective imaging modality that is extensively used in clinical practice for low back pain evaluation [
3,
4]. It reveals several degenerative features, including scoliosis, spondylolisthesis, osteophyte formation, and disc space narrowing. Furthermore, key radiographic parameters, including spinopelvic alignment measurements, are extracted from lumbar radiographs, providing valuable insights for comprehensive patient assessment and treatment planning [
5-
7].
The interpretation of lumbar radiographs is not particularly complex; however, the large number of examinations still demands considerable clinical workload [
8,
9]. In recent years, deep learning (DL) technology has achieved rapid development. DL employs a neural network for data processing and pattern recognition, thereby enabling automation of medical imaging diagnosis and analysis [
10,
11]. DL has been applied to lumbar radiographs employing 2 primary approaches. (1) Segmentation-based methods delineate vertebral structures at the pixel level, enabling rapid extraction of quantitative spinal parameters based on the localization of key anatomical landmarks [
12-
15]. (2) Detection-based models localize regions of interest and assign diagnostic probabilities to pathological features, allowing automated disease identification and classification [
16-
18]. However, the clinical application of DL models remains limited compared with the promising results that were reported in previous studies. Several limitations exist in the DL systems’ assistance in lumbar radiograph interpretation, such as in the diagnosis of degenerative conditions, including scoliosis and spondylolisthesis, depend on the spatial relationship between vertebrae due to the unique anatomical structure of the lumbar spine. Therefore, accurate automated diagnosis requires the integration of both vertebral segmentation and detection-based diagnosis. Reportedly, no existing studies have simultaneously leveraged both the techniques. Furthermore, most previous studies have focused on detecting a single lumbar pathology, which limits the comprehensiveness of the model’s diagnostic capabilities [
13,
17,
19,
20].
Therefore, this study proposes a dual-mechanism strategy that integrates outputs from both segmentation-based and detection-based networks to enhance the diagnostic performance for lumbar degenerative diseases. This model is developed for deployment in primary healthcare settings, helping in screening common lumbar degenerative changes and generating structured diagnostic reports.
Fig. 1 illustrates an overview of the proposed DL model.
MATERIALS AND METHODS
1. Data Collection
This retrospective study employed the DL method in the assessment of lumbar plain radiographs from the institutional and public datasets. This study adhered to the Declaration of Helsinki and was approved by the institutional review board of Beijing Chaoyang Hospital, Capital Medical University (2024-KE-385), and the requirement of informed consent was waived due to its retrospective nature. Private information of all patients was anonymized before image acquisition.
The imaging and clinical data of 5,964 patients who underwent standing anteroposterior (AP) and lateral lumbar radiographs at Beijing Chaoyang Hospital, Capital Medical University were collected. Specifically, the dataset comprises 4,034 consecutive patients imaged from October 2022 to August 2024. To improve the sample size and class balance for each lesion detection task, this study included an additional 1,210 patients with lumbar scoliosis and 720 patients with lumbar spondylolisthesis from January 2019 to September 2022. Only standard AP and lateral lumbar radiographs were included. Oblique, flexion-extension, or other special projections were excluded. For patients with multiple radiographs, the clearest AP and lateral views meeting predefined quality criteria (adequate coverage, minimal rotation, and sufficient visualization of vertebral margins) were selected. Inclusion criteria were (1) age of >18 years and (2) complete visualization of vertebral bodies from T11 to L5 on both AP and lateral views. Exclusion criteria were (1) poor image quality that precluded clear delineation of vertebral bodies; (2) severe lumbar spine trauma; (3) lumbar surgery history; (4) lumbosacral transitional vertebrae, congenital spinal deformities, tuberculosis, infection, or tumors; and (5) conditions that could not be reliably diagnosed on plain radiographs. For external validation, cases numbered 1–1,033 from the BUU-Spine [
21] public dataset were reviewed. After applying the same inclusion and exclusion criteria as the local dataset, 600 AP–lateral pairs met all requirements and were included as the external test set.
2. Image Acquisition and Labeling
The lumbar plain radiographs were retrieved from the PACS (Picture Archiving and Communication System) in Digital Imaging and Communications in Medicine (DICOM) format. Images were not preprocessed before annotation. The DICOM files were then converted to JPG format for annotation. All manual labeling was conducted using the Labelme (ver. 5.2.0, Anaconda, USA). The BUU-Spine dataset provides existing annotations for the external dataset; however, all images were reannotated to maintain consistency with the labeling standards employed in our study.
The images were annotated with bounding boxes for lesion detection and vertebral segmentation masks. Bounding boxes were used to label the presence of lumbar scoliosis on AP views for lesion detection annotation. Bounding boxes were applied to annotate the spondylolisthesis, disc space narrowing, osteophyte formation, vertebral compression, abdominal aortic calcification, and preserved/straightened lumbar lordosis on lateral views. Rectangular bounding boxes were applied to cover the radiographic region corresponding as completely as possible to each specific degenerative finding. The diagnostic criteria for each lesion were (1) lumbar scoliosis (defined as a Cobb angle of >10°); (2) lumbar lordosis (defined by the angle between the superior endplates of L1 and S1, where an angle of >30° and <30° indicated preserved and straightened lumbar lordosis, respectively [
22]); (3) spondylolisthesis (defined as anterior or posterior displacement of one vertebral body relative to the adjacent inferior vertebra; a displacement of ≥3 mm was used as the diagnostic threshold [
23]); (4) disc space narrowing (defined as intervertebral disc height loss of >33% compared with the average disc height of normal segments [
24]); (5) osteophyte formation (defined as a bony outgrowth of >3 mm beyond the vertebral margin [
25]); (6) vertebral compression (defined as a wedging deformity or height loss of the vertebral body); (7) abdominal aortic calcification (defined as linear or curvilinear calcifications anterior to the vertebral bodies, consistent with the expected course of the abdominal aorta).
The 4 corners of each vertebral body from T11 to L5 were annotated on both AP and lateral views. Polygonal annotations were used to mark the 4 anatomical corner points of each vertebral body. Sacral segmentation was also performed in lateral radiographs by marking the anterior-superior and posteriorsuperior corners of S1 and delineating the sacral contour as accurately as possible. The annotated boundary was positioned to align with the midline of the actual vertebral edge when the anterior and posterior margins of a vertebral body did not overlap on the radiograph. Quantitative measurements were performed after vertebral segmentation to identify the presence of lumbar scoliosis, lordosis alteration, spondylolisthesis, and disc space narrowing based on the aforementioned criteria.
Fig. 2 illustrates all vertebral segmentation and lesion detection annotations.
Three spinal surgeons (S1, S2, and S3 with 25, 15, and 10 years of experience, respectively) performed the annotation process. Surgeons were blinded to each other’s annotations and they independently annotated all images. To assess the intraobserver reliability, the test set was reannotated by S1 after 1 month. The average of the annotating results of the 3 surgeons was considered the reference standard for vertebral segmentation. Annotations from all surgeons were compared for lesion detection. The bounding box annotated by S1 was used as the reference standard for cases with consistent diagnoses. The final reference was identified through consensus discussion for cases with disagreement. The annotated data were then used to train and validate the model.
3. Development of the DL Algorithm
In this study, we proposed a DL-based framework for automatic vertebral corner detection from lumbar spine x-ray images. The overall architecture comprised 3 main components: (1) an encoder for feature extraction, (2) a decoder for feature reconstruction, and (3) a corner prediction module composed of heatmap and offset regressions. We developed a detection-based framework for identifying multiple lumbar spine pathologies directly from x-ray images, in addition to vertebral corner localization. The proposed model is established upon the YOLOv8 architecture, which integrates efficient feature extraction, multiscale feature fusion, and multitask prediction. The detection model was developed using all internal data, whereas the segmentation model was constructed with 3,303 paired images from the same dataset. The dataset was randomly divided into training, validation, and test sets in a ratio of 7:2:1, respectively.
Furthermore, to translate the outputs of the dual-mechanism model into clinically interpretable diagnostic reports, a rule-based integration strategy was developed. Specifically, when the segmentation and detection models demonstrated consistent diagnostic results, the system directly output a definite positive or negative conclusion. A possible positive conclusion is generated if the 2 model outputs are inconsistent but the segmentation-based parameters are close to the diagnostic threshold, considering that quantitative measurements from the segmentation model may contain minor errors. The automatic analysis is considered unreliable when the diagnostic results of the 2 models differ substantially, and a manual review by radiologists is recommended. The
Supplementary Material provides a detailed description of the model architecture and the structured report generation strategy.
4. Statistical Analysis
1) Reliability of the annotations
The percentages within 1-, 2-, 3-, 4-, and 5-mm landmark-to-landmark distance thresholds were computed to assess the inter-and intraobserver reliability of vertebral segmentation annotation. The pairwise precision and recall were calculated to reflect the inter- and intraobserver reliability of lesion detection annotation.
2) Model performance of vertebral segmentation
The percentage of correct key points (PCK) was used to investigate the spatial accuracy of all predicted anatomical landmarks. The intraclass correlation coefficient (ICC), Pearson correlation coefficient (r), mean difference, standard deviation, and mean absolute error between the model estimates and the reference standard were then calculated on the test set for the quantitative measurements derived from the predicted landmarks.
3) Model performance of automated diagnosis
Precision, recall, and the area under the receiver operating characteristic (ROC) curve were calculated to assess diagnostic performance concerning the diagnostic outputs generated by the segmentation-based model, detection-based model, and the dual-mechanism model on the test set. Further, a senior spine surgeon who was not involved in the model development process independently evaluated all diagnostic reports generated by the model. Reports without apparent missed or incorrect diagnoses were considered clinically acceptable, and the clinical acceptability rate was calculated accordingly.
RESULTS
1. Patient Information
The internal dataset comprised paired AP and lateral lumbar radiographs from 5,964 patients. The mean age of the patients was 55.91±7.22 years, and the mean body mass index was 24.21± 2.57. The external dataset contained 600 patients with a mean age of 55.39±10.86 years.
Table 1 presents the demographic data distribution.
2. Reliability of Annotations
The percentage of intraobserver landmark distance within the 3-mm threshold reached 96%, and the percentage of interobserver landmark distances within the 3-mm threshold exceeded 95%, indicating relatively good consistency after predefining the segmentation criteria (
Supplementary Table 1).
Supplementary Table 2 provides the precision and recall of intra- and interobserver reliability for lesion bounding box annotation in the internal test set. Generally, the lesion annotations demonstrated satisfactory consistency. Upon discussion and adjudication, 1,711 lumbar scoliosis, 1,311 lumbar spondylolisthesis, 858 vertebral compressions, 2,571 osteophytes, 1,093 intervertebral space narrowings, and 888 aortic calcifications were annotated as the ground truth.
3. Performance of Vertebral Segmentation
The PCKs within 3 mm ranged from 95.7% to 97.5% in the AP view and from 95.7% to 98.6% in the lateral view in the internal test set (
Supplementary Table 3). The PCKs within 3 mm ranged from 92.2% to 96.4% in the AP view and from 92.7% to 96.3% in the lateral view in the external test set (
Supplementary Table 4).
Fig. 3 illustrates particular examples of the model predicting landmarks. Additionally, the performance of parameter measurement based on model segmentation was compared with the reference standard (
Tables 2 and
3). The ICC between the model predictions and the reference standard in the internal dataset is 0.84–0.89, and the Pearson correlation coefficient (r) ranged from 0.90 to 0.94, indicating good measurement consistency.
4. Model Performance of Automated Diagnosis
Table 4 summarizes diagnostic performance for each model in the internal test set. The segmentation-based model achieved precision of 92.2%–96.9% and recall of 91.6%–94.8%, whereas the detection-based model was comparable (precision 91.7%–95.5%; recall 93.3%–95.2%).
Fig. 4 illustrates the ROC curves of the detection-based model for diagnosing degenerative changes. Two positive determination conditions were defined for the integration of the dual-mechanism model: dual-positive, in which both models detected a positive finding, and any-positive, in which a case was considered positive if either model indicated a positive result. The dual-positive condition model delivered the highest precision for 4 lesions (93.8%–97.3%), whereas the any-positive condition model yielded the highest recall (94.1%–97.6%).
Table 5 and
Fig. 5 show the diagnostic performance for each model in the external test set.
Furthermore, 596 structured diagnostic reports were generated using the previously described strategy. After review by a senior physician, 557 reports contained no evident omissions or misdiagnoses and were considered acceptable, leading to an overall acceptance rate of 93.4%. Among the 39 reports with issues, 25 involved clear missed findings and 17 contained incorrect diagnoses; 3 reports contained both types of errors. The supplementary file provides case illustrations of the diagnostic reports.
DISCUSSION
In this study, we proposed a dual-mechanism DL model for assessing lumbar degeneration, which requires consideration of intervertebral positional relationships. Although previous studies have applied DL technique to lumbar radiograph analysis [
26,
27], most relied on a single segmentation-based pipeline. In contrast, our study adopts 2 independent and parallel networks that integrate segmentation-based quantification with detection-based lesion recognition. This helps mitigate the inherent limitations and failure modes of individual models, enabling more reliable multidisease diagnosis. In addition, the structured diagnostic report generated by our framework provides a more direct and clinically interpretable interface than isolated numerical outputs. This design is intended to mimic the real-world radiological workflow, thus enhances the practical applicability of automated lumbar radiograph analysis in routine clinical workflows. The test results revealed that the proposed model achieved a precision of 93.8%–97.3% under the dual-positive condition (indicating a lower false-positive rate), and a recall of 94.1%–97.6% under the any-positive condition (denoting a lower false-negative rate). Moreover, structured diagnostic reports enabled reinforcement of dual-positive results and cautious presentation of any-positive results, thereby theoretically allowing the model to achieve both high precision and recall simultaneously. External testing further revealed promising outcomes, underscoring the model’s potential for clinical application.
Although the interpretation of lumbar spine radiographs is not inherently complex, substantial global demand, particularly in screening settings within primary healthcare institutions and the high imaging volume in general hospitals, makes missed diagnoses, misdiagnoses, and fatigue-related errors still common [
28]. Therefore, the development of artificial intelligence techniques for automated image assessment holds considerable promise. For example, our model could be integrated into existing clinical imaging systems in the future, enabling simultaneous generation of lesion annotations, quantitative measurements, and preliminary diagnostic reports directly from radiographic images. This approach can improve the efficiency of radiologists and promote standardization in radiographic interpretation. In addition, it allows spine surgeons to more rapidly identify prominent as well as atypical imaging findings, thereby facilitating a more effective correlation between radiographic findings and patient symptoms. These considerations highlight that the primary value of the proposed model lies in serving as a high-efficiency screening and quantitative assistance tool rather than a complete replacement for clinical decision-making. Furthermore, the model can function as an essential foundational component for future artificial intelligence-based clinical decision support systems that integrate multidimensional clinical information.
Segmentation-based assessment of lumbar radiographs is a classical and widely adopted approach, as most diagnoses of lumbar degeneration depend on the identification and measurement of key anatomical landmarks. Recently, several studies have focused on improving vertebral segmentation, with proposed strategies, including the incorporation of attention mechanisms, multiscale feature fusion, cascaded network architectures, and level-set-based postprocessing for refinement [
15,
26,
29-
31]. In the present study, the proposed segmentation model incorporates a ResNet-based encoder and a decoder with skip connections. The ResNet encoder provides hierarchical feature abstraction, whereas the decoder effectively restores spatial information and transfers fine-grained details from early encoder layers to the final predictions. ResNet-34 was chosen to balance model capacity and generalization. Shallower networks such as ResNet-18 may lack sufficient receptive field to capture the global context of vertebral alignment across multiple levels, whereas deeper backbones increase the risk of overfitting without consistent performance gains on datasets of this scale [
32]. In our experiments, ResNet-34 provided a stable compromise between anatomical representation ability and training robustness. The corner prediction module jointly optimizes heatmap generation and offsets regression, thereby permitting coarse-to-fine localization that effectively mitigates quantization errors that are commonly observed in such dense prediction tasks. In comparison to previous studies, the proposed model achieved a higher PCK within the 1-mm threshold [
13,
33], indicating a satisfactory segmentation precision level.
However, the PCK did not reach 100% even under the 5-mm threshold, indicating that notable segmentation deviations may persist in certain cases. This can be attributed to several inherent limitations of lumbar radiographs that constrain precise vertebral segmentation. First, the clarity of the vertebral boundary is susceptible to surrounding thoracoabdominal structures (including intestinal gas) and suboptimal image quality [
26]. Second, the anterior and posterior vertebral edges do not overlap in many images due to lumbar lordosis or scoliosis. Existing DL models generally lack prior knowledge of vertebral 3-dimensional structures; thus, their boundary determination may be less reliable than that of experienced radiologists. In present study, the segmentation accuracy of the L5 vertebra was lower compared with other vertebrae, partly because of its location at the lower end of the lumbar lordosis and adjacency to the sacrum with complex surrounding structures [
34]. Third, the vertebral morphology and positional relationships may be markedly altered in certain cases of severe degeneration, where segmentation models often fail to produce usable results. Further, the ICC for segmentation-based quantitative measurements between the model estimates and the reference standard remained <0.90. This is primarily because even minor deviations in corner localization cause obvious errors in angle and relative distance calculations.
Compared with segmentation-based methods, detection-based approaches have been less commonly applied in previous studies on lumbar radiographs, possibly because their outputs are less interpretable and more challenging to integrate with clinical knowledge frameworks, thereby imposing higher costs for diagnostic grading and classification. However, segmentation errors and failures remain inevitable due to the aforementioned reasons. Therefore, detection-based diagnosis is considered a useful validation and complement to segmentation-based approaches. Several studies have applied detection-based methods to diagnose pathologies, including lumbar spondylolisthesis and aortic calcification [
17,
18,
35,
36]. Unlike segmentation, object detection focuses more on local imaging features of lesions rather than the overall vertebral structure, making it relatively less sensitive to image quality and radiographic projection angles. In this study, a lightweight and computationally efficient detection network is essential, considering the cross-validation characteristics of the dual-mechanism framework and the practical constraints of clinical deployment. Therefore, a modified architecture based on YOLOv8 was adopted. YOLOv8 uses a single-stage and anchor-free design that enables streamlined end-to-end image processing with relatively low parameter counts and computational complexity [
37,
38]. Furthermore, YOLOv8 has the advantage of stable multiscale feature learning capability [
39], which is particularly suitable for lumbar degenerative findings with heterogeneous sizes and appearances. However, future studies may explore newer architectures to further optimize computational efficiency in specific deployment scenarios.
In this study, the dual-mechanism model was employed to assess 4 types of lesions. A marked improvement in diagnostic performance was observed for scoliosis and straightened lumbar lordosis compared with the single-mechanism models. This finding indicates that our strategy is particularly effective for global pathological features, where segmentation-based quantitative measurements and detection-based morphological assessments provide cross-validation, leading to more reliable outputs. In contrast, the improvement for spondylolisthesis and disc space narrowing was less pronounced. This may be because these conditions are characterized by local imaging changes, and the mild cases in the dataset increase the difficulty of model training. Specifically, disc space narrowing is challenging to quantify because the diagnosis depends on relative comparisons rather than on absolute numerical thresholds. This ambiguity poses a fundamental challenge for developing automated models. To further improve the diagnostic capability for these lesions, future research is warranted to focus on strategies, including local refinement of segmentation, 2-stage detection frameworks, and reinforcement of training with hard example mining to improve performance on challenging cases.
Currently, the integration of outputs from the dual-mechanism model is based on a logic-gating strategy. This approach enables effective cross-validation while tolerating minor segmentation errors, providing a cost-effective solution with high interpretability. Decision-level fusion [
40], which theoretically enables more flexible weighting of each model to make the diagnostic results closer to the gold standard, is an alternative option. However, this approach requires a large and perfectly annotated dataset to reliably learn optimal fusion weights and compromises decision transparency. Therefore, our rule-based strategy was considered more appropriate for the current study. The use of large language models (LLMs) represents another promising direction. LLMs can use their advanced reasoning capabilities to emulate clinicians’ decision-making process and generate optimized diagnostic reports [
41]. However, the training and deployment costs associated with developing an LLM with sufficient clinical expertise are substantial, indicating that this approach should be carefully assessed [
42].
This study has several limitations. First, we have expanded the sample size as much as possible; however, the current dataset remains insufficient to cover all types and degrees of lumbar degeneration. Larger-scale, multicenter, and more balanced datasets are required in the future to enhance model performance. Second, the optimal architectures for segmentation and detection networks remain to be identified, and continued optimization of both networks would enhance the overall performance of the dual-mechanism model. Third, the proposed model can only provide preliminary diagnoses according to lumbar radiographs and is unable to incorporate multidimensional information, such as patient history and clinical symptoms, for comprehensive assessment. Fourth, the cost-effectiveness and ethical considerations in clinical application remain unexplored.
CONCLUSION
This dual-mechanism DL model, combining vertebral segmentation with lesion detection, enables automated multilesion assessment of lumbar degeneration on plain radiographs. It demonstrated high diagnostic performance, strong agreement with expert measurements, and a high rate of clinically acceptable structured reports. This approach may help standardize reporting, assist nonspecialists, and improve efficiency in the radiographic evaluation of patients with lumbar degenerative disease, although further multicenter validation is warranted before routine clinical adoption.
Supplementary Materials
Supplementary Fig. 3.
Model discrepancies in lumbar spondylolisthesis diagnosis (for clarity of presentation, only landmarks and bounding boxes relevant to the illustrated discrepancies are shown).
ns-2551672-836-Supplementary-Fig-3.pdf
Supplementary Fig. 4.
Model discrepancies in disc space narrowing diagnosis (for clarity of presentation, only landmarks and bounding boxes relevant to the illustrated discrepancies are shown).
ns-2551672-836-Supplementary-Fig-4.pdf
NOTES
-
Conflict of Interest
The authors have nothing to disclose.
-
Funding/Support
This study received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.
-
Acknowledgments
We thank the following individual for his valuable contributions to this study: Jian Li, MD, Department of Orthopedics, Beijing Chaoyang Hospital, Capital Medical University, Beijing, China.
-
Author Contribution
Conceptualization: ZM, AW, XL, TW, YZ, LZ; Data curation: ZM, AW, XL, DL; Formal analysis: ZM, AW, RC, YX; Funding acquisition: LZ; Methodology: ZM, XL, TW, ML; Project administration: XL, SY, LZ; Visualization: ZM, RC, NF, PD, YZ; Writing – original draft: ZM, AW, RC, YX, PD; Writing – review & editing: ZM, AW, XL, TW, ML, NF, SY, DL, YZ, LZ.
Fig. 1.Overview of the proposed dual-mechanism deep learning pipeline.
Fig. 2.Illustration of the image annotation workflow. Vertebral segmentation in the anteroposterior (AP) (A) and lateral (B) views, lumbar scoliosis (C), no scoliosis in the AP view (D), preserved lumbar lordosis (E), straightened lumbar lordosis (F), abdominal aortic calcification (G), lumbar spondylolisthesis (H), disc space narrowing (I), osteophyte formation (J), vertebral compression (K), multiple lesions (L) were annotated with different label names in an image. The display color was used for visual distinction only and carried no categorical information.
Fig. 3.Predicted positions of landmarks (red points) for representative images compared with the reference standard (blue points).
Fig. 4.Receiver operating characteristic curves (AUCs) of the detection-based model for diagnosing multiple degenerative changes in the internal test set.
Fig. 5.Receiver operating characteristic curves (AUCs) of the detection-based model for diagnosing multiple degenerative changes in the external test set.
Table 1.Demographic information of included patients
Table 1.
|
Characteristic |
Training set (n = 4,175) |
Validation set (n = 1,193) |
Test set (n = 596) |
External test set (n = 600) |
|
Age (yr) |
54.99 ± 6.99 |
59.03 ± 7.21 |
56.14 ± 7.09 |
55.39 ± 10.86 |
|
Sex, male:female |
2,196:1,979 |
553:640 |
317:279 |
275:325 |
|
BMI (kg/m²) |
24.01 ± 2.43 |
25.60 ± 2.55 |
22.87 ± 2.38 |
NA |
Table 2.Comparison between the model and the reference standard for the quantitative measurements in the internal test set
Table 2.
|
Parameters |
ICC (95% CI) |
r |
MD |
SD |
MAE |
|
Cobb angle (°) |
0.87 (0.85–0.90) |
0.91 |
0.7 |
1.2 |
1.1 |
|
Lumbar lordosis (°) |
0.89 (0.88–0.91) |
0.94 |
0.4 |
0.7 |
0.8 |
|
Percentage of spondylolisthesis (%) |
0.85 (0.83–0.87) |
0.90 |
1.5 |
0.8 |
1.8 |
|
Disc height index (%) |
0.84 (0.82–0.86) |
0.90 |
0.9 |
1.4 |
1.5 |
Table 3.Comparison between the model and the reference standard for the quantitative measurements in the external test set
Table 3.
|
Parameters |
ICC (95% CI) |
r |
MD |
SD |
MAE |
|
Cobb angle (°) |
0.83 (0.82–0.88) |
0.89 |
1.0 |
1.3 |
1.3 |
|
Lumbar lordosis (°) |
0.86 (0.84–0.90) |
0.91 |
0.9 |
0.9 |
1.1 |
|
Percentage of spondylolisthesis (%) |
0.83 (0.81–0.86) |
0.89 |
1.7 |
1.0 |
2.0 |
|
Disc height index (%) |
0.81 (0.78-0.85) |
0.88 |
1.1 |
1.7 |
2.0 |
Table 4.Diagnostic performance in the internal test set
Table 4.
|
Variable |
Segmentation-based
|
Detection-based
|
Dual-mechanism (dual-positive)
|
Dual-mechanism (any-positive)
|
|
Precision |
Recall |
Precision |
Recall |
Precision |
Recall |
Precision |
Recall |
|
Lumbar scoliosis |
96.7% |
94.4% |
95.5% |
95.2% |
97.1% |
94.1% |
95.1% |
97.6% |
|
Straightened lumbar lordosis |
96.9% |
94.8% |
95.5% |
95.1% |
97.3% |
94.4% |
95.2% |
96.2% |
|
Lumbar spondylolisthesis |
92.2% |
91.6% |
93.3% |
94.1% |
93.8% |
91.2% |
91.9% |
94.5% |
|
Disc space narrowing |
93.7% |
92.1% |
92.5% |
93.3% |
94.1% |
91.6% |
91.7% |
94.1% |
|
Osteophytes |
- |
- |
94.4% |
93.5% |
- |
- |
- |
- |
|
Vertebral compression |
- |
- |
95.2% |
94.2% |
- |
- |
- |
- |
|
Abdominal aortic calcification |
- |
- |
91.7% |
93.3% |
- |
- |
- |
- |
Table 5.Diagnostic performance in the external test set
Table 5.
|
Variable |
Segmentation-based
|
Detection-based
|
Dual-mechanism (dual-positive)
|
Dual-mechanism (any-positive)
|
|
Precision |
Recall |
Precision |
Recall |
Precision |
Recall |
Precision |
Recall |
|
Lumbar scoliosis |
93.3% |
92.6% |
93.6% |
91.7% |
94.2% |
91.6% |
92.8% |
93.1% |
|
Straightened lumbar lordosis |
93.7% |
92.9% |
92.3% |
93.9% |
94.1% |
92.7% |
92.0% |
94.3% |
|
Lumbar spondylolisthesis |
91.1% |
90.2% |
90.5% |
92.0% |
91.9% |
89.9% |
90.0% |
92.7% |
|
Disc space narrowing |
90.3% |
91.3% |
90.8% |
91.6% |
91.3% |
90.8% |
89.7% |
92.2% |
|
Osteophytes |
- |
- |
92.1% |
91.8% |
- |
- |
- |
- |
|
Vertebral compression |
- |
- |
93.1% |
91.7% |
- |
- |
- |
- |
|
Abdominal aortic calcification |
- |
- |
89.8% |
90.9% |
- |
- |
- |
- |
REFERENCES
- 1. Chou R. Low Back Pain. Ann Intern Med 2021;174:ITC113-28.
- 2. Knezevic NN, Candido KD, Vlaeyen JW, et al. Low back pain. Lancet 2021;398:78-92.
- 3. Jarvik JG, Hollingworth W, Martin B, et al. Rapid magnetic resonance imaging vs radiographs for patients with low back pain: a randomized controlled trial. JAMA 2003;289:2810-8.
- 4. Njeze NR, Ezeofor SN, Agwu-Umahi OR. Plain radiographs of lumbar spine in patients with low back pain. Arch Osteoporos 2018;13:104.
- 5. Cha E, Park JH. Spinopelvic alignment as a risk factor for poor balance function in low back pain patients. Global Spine J 2023;13:2193-200.
- 6. Chun SW, Lim CY, Kim K, et al. The relationships between low back pain and lumbar lordosis: a systematic review and meta-analysis. Spine J 2017;17:1180-91.
- 7. Matsumoto T, Okuda S, Maeno T, et al. Spinopelvic sagittal imbalance as a risk factor for adjacent-segment disease after single-segment posterior lumbar interbody fusion. J Neurosurg Spine 2017;26:435-40.
- 8. Tan A, Zhou J, Kuo YF, et al. Variation among primary care physicians in the use of imaging for older patients with acute low back pain. J Gen Intern Med 2016;31:156-63.
- 9. Vader JP, Terraz O, Perret L, et al. Use of and irradiation from plain lumbar spine radiography in Switzerland. Swiss Med Wkly 2004;134:419-22.
- 10. Chen X, Wang X, Zhang K, et al. Recent advances and clinical applications of deep learning in medical image analysis. Med Image Anal 2022;79:102444.
- 11. Jiang H, Diao Z, Shi T, et al. A review of deep learning-based multiple-lesion recognition from medical images: classification, detection and segmentation. Comput Biol Med 2023;157:106726.
- 12. Schwartz JT, Cho BH, Tang P, et al. Deep learning automates measurement of spinopelvic parameters on lateral lumbar radiographs. Spine (Phila Pa 1976) 2021;46:E671-8.
- 13. Wu Y, Chen X, Dong F, et al. Performance evaluation of a deep learning-based cascaded HRNet model for automatic measurement of X-ray imaging parameters of lumbar sagittal curvature. Eur Spine J 2024;33:4104-18.
- 14. Yao H, Zhang Z, Cheng G, et al. Automatic measurement of anatomical parameters of the lumbar vertebral body and the intervertebral disc on radiographs by deep learning. Quant Imaging Med Surg 2024;14:5877-90.
- 15. Yuh WT, Khil EK, Yoon YS, et al. Deep learning-assisted quantitative measurement of thoracolumbar fracture features on lateral radiographs. Neurospine 2024;21:30-43.
- 16. Suzuki H, Kokabu T, Yamada K, et al. Deep learning-based detection of lumbar spinal canal stenosis using convolutional neural networks. Spine J 2024;24:2086-101.
- 17. Xu C, Liu X, Bao B, et al. Two-stage deep learning model for diagnosis of lumbar spondylolisthesis based on lateral x-ray images. World Neurosurg 2024;186:e652-61.
- 18. Zhang J, Lin H, Wang H, et al. Deep learning system assisted detection and localization of lumbar spondylolisthesis. Front Bioeng Biotechnol 2023;11:1194009.
- 19. Wang K, Wang X, Xi Z, et al. Automatic segmentation and quantification of abdominal aortic calcification in lateral lumbar radiographs based on deep-learning-based algorithms. Bioengineering (Basel) 2023;10:1164.
- 20. Yeşilmen N, Danacı Ç, Baydoğan MP, et al. Enhanced vision transformer with custom attention mechanism for automated idiopathic scoliosis classification. J Imaging Inform Med 2025 Jun 2. doi: 10.1007/s10278-025-01564-w. [Epub].
- 21. Klinwichit P, Yookwan W, Limchareon S, et al. BUU-LSPINE: a Thai Open Lumbar Spine Dataset for spondylolisthesis detection. Appl Sci 2023;13:8646.
- 22. Been E, Kalichman L. Lumbar lordosis. Spine J 2014;14:87-97.
- 23. Warashina H, Kato M, Kitamura S, et al. The progression of osteoarthritis of the hip increases degenerative lumbar spondylolisthesis and causes the change of spinopelvic alignment. J Orthop 2019;16:275-9.
- 24. Wilke HJ, Rohlmann F, Neidlinger-Wilke C, et al. Validity and interobserver agreement of a new radiographic grading system for intervertebral disc degeneration: Part I. Lumbar spine. Eur Spine J 2006;15:720-30.
- 25. Mimura M, Panjabi MM, Oxland TR, et al. Disc degeneration affects the multidirectional flexibility of the lumbar spine. Spine (Phila Pa 1976) 1994;19:1371-80.
- 26. Kim KC, Cho HC, Jang TJ, et al. Automatic detection and segmentation of lumbar vertebrae from X-ray images for compression fracture evaluation. Comput Methods Programs Biomed 2021;200:105833.
- 27. Song SY, Seo MS, Kim CW, et al. AI-driven segmentation and automated analysis of the whole sagittal spine from x-ray images for spinopelvic parameter evaluation. Bioengineering (Basel) 2023;10:1229.
- 28. Zhang L, Wen X, Li JW, et al. Diagnostic error and bias in the department of radiology: a pictorial essay. Insights Imaging 2023;14:163.
- 29. Kim DH, Jeong JG, Kim YJ, et al. Automated vertebral segmentation and measurement of vertebral compression ratio based on deep learning in x-ray images. J Digit Imaging 2021;34:853-61.
- 30. Shi W, Xu T, Yang H, et al. Attention gate based dual-pathway network for vertebra segmentation of x-ray spine images. IEEE J Biomed Health Inform 2022;26:3976-87.
- 31. Zhang B, Chen K, Yuan H, et al. Automatic Lenke classification of adolescent idiopathic scoliosis with deep learning. JOR Spine 2024;7:e1327.
- 32. Wang S, Tong X, Cheng Q, et al. Fully automated deep learning system for osteoporosis screening using chest computed tomography images. Quant Imaging Med Surg 2024;14:2816-27.
- 33. Zhou S, Yao H, Ma C, et al. Artificial intelligence X-ray measurement technology of anatomical parameters related to lumbosacral stability. Eur J Radiol 2022;146:110071.
- 34. Jang JS, Kim JI, Ku B, et al. Reliability analysis of vertebral landmark labelling on lumbar spine x-ray images. Diagnostics (Basel) 2023;13:1411.
- 35. Paik S, Park J, Hong JY, et al. Deep learning application of vertebral compression fracture detection using mask R-CNN. Sci Rep 2024;14:16308.
- 36. Voss A, Suoranta S, Nissinen T, et al. Deep learning ensemble for abdominal aortic calcification scoring from lumbar spine X-ray and DXA images. Comput Biol Med 2025;197:110961.
- 37. Jeong SW, Ahmad S, Kim JS, et al. A lightweight YOLOv8-based model for gastric cancer detection. Comput Biol Med 2025;196:110689.
- 38. A Hasib U, Md Abu R, Yang J, et al. YOLOv8 framework for COVID-19 and pneumonia detection using synthetic image augmentation. Digit Health 2025;11:20552076251341092.
- 39. Cai Z, Zhou K, Liao Z. A systematic review of YOLO-based object detection in medical imaging: advances, challenges, and future directions. CMC 2025;85:2255-303.
- 40. Li J, Qiu T, Wen C, et al. Robust face recognition using the deep C2D-CNN model based on decision-level fusion. Sensors (Basel) 2018;18:2080.
- 41. Zhang L, Liu M, Wang L, et al. Constructing a large language model to generate impressions from findings in radiology reports. Radiology 2024;312:e240885.
- 42. Koohi-Moghadam M, Bae KT. Generative AI in medical imaging: applications, challenges, and ethics. J Med Syst 2023;47:94.