Development and validation of an interpretable ultrasound radiomics model for benign and malignant classification of breast lesions: a multicenter large-sample study.
Objectives To develop and validate a combined ultrasound-based radiomics-clinical model for differentiating benign and malignant breast lesions. Materials and methods A total of 3142 patients from eight hospitals between February 2012 and September 2024 were included in this multicenter retrospective development and validation study, with an additional single-center prospective test cohort. Lesions were manually segmented, and radiomics features were automatically extracted to construct five machine learning models. The best-performing radiomics model was combined with clinical features to build a combined model. Model performance and its impact on Breast Imaging Reporting and Data System (BI-RADS)-based biopsy decisions were evaluated. Results Logistic regression (LR) showed the best radiomics performance, with area under the curves (AUCs) of 0.83, 0.82, 0.81, and 0.82 across the training, internal test, external test, and prospective test sets. The clinical model achieved AUCs of 0.87, 0.85, 0.87, and 0.86, whereas the combined model achieved AUCs of 0.92, 0.90, 0.92, and 0.93, significantly outperforming both single-modality models (all p Conclusion The interpretable ultrasound-based radiomics model enables reliable, noninvasive breast lesion diagnosis and may reduce unnecessary biopsies. Critical relevance statement This work developed an interpretable radiomics-clinical combined model in a multicenter retrospective development and validation study, with additional testing in a single-center prospective cohort, and may support breast lesion risk stratification and biopsy decision-making after further prospective clinical utility evaluation. Key points Conventional ultrasound diagnosis of breast cancer shows limited specificity. A multicenter radiomics-clinical combined model showed improved diagnostic performance, with additional validation in a prospective test cohort.
Introduction
Breast cancer has become the most common malignancy and the leading cause of cancer-related mortality in women globally [1]. Early detection and accurate diagnosis are crucial for guiding treatment and improving outcomes [2,3]. Common imaging modalities for breast lesion evaluation include mammography, ultrasound, and magnetic resonance imaging (MRI) [4]. Mammography is effective for early detection but has limited sensitivity in women with dense breasts [5]. MRI offers high resolution but is restricted by cost and accessibility [6].
Ultrasound, as a real-time and radiation-free imaging modality, is widely used for the initial screening and diagnosis of breast lesions [7,8]. However, the interpretation of ultrasound images is highly dependent on the experience of radiologists, introducing a degree of subjectivity that may lead to inconsistencies in diagnosis. To address this issue, the American College of Radiology introduced the fifth edition of the Breast Imaging Reporting and Data System (BI-RADS) in 2013 [9]. Although BI-RADS has improved objectivity, distinguishing atypical benign findings from malignant lesions remains difficult due to overlapping imaging features. In particular, when breast lesions are evaluated using B-mode ultrasound (BMUS) alone, category 4 lesions show high false-positive rates and lack detailed sub-classification criteria [10–13], leading to frequent reliance on biopsy. Although incorporating mammography, color Doppler, and relevant clinical information may further improve lesion assessment, biopsy remains common and increases patient burden, highlighting the need to improve diagnostic accuracy while reducing unnecessary interventions.
Radiomics is a new analytical method that allows for the extraction of numerous quantitative features from medical images, helping to reveal subtle pathological information hidden within the imaging data [14,15]. Radiomics has been widely applied in breast disease, especially when combined with ultrasound imaging [16–18]. Previous studies have leveraged machine learning algorithms to develop models for accurate classification of breast lesions [19–22]. However, most existing studies only focused on single machine learning methods and lacked prospective, multicenter, large-scale validation, limiting the generalizability and clinical applicability. Moreover, few studies have specifically focused on reducing unnecessary biopsies in clinical practice, and comprehensive evaluations of the clinical utility of these models remain limited.
Therefore, this study aims to develop and validate an integrated radiomics model combining ultrasound-based features with clinical-ultrasound characteristics. The model’s performance will be tested across multicenter and prospective datasets, with particular emphasis on its impact on BI-RADS classification and biopsy recommendations. To improve interpretability, Shapley additive explanations (SHAP) will be used to clarify the model’s decision-making process.
Materials and methods
This study received approval from the Institutional Ethics Committee, with separate approval numbers for the retrospective (No. PJ2023-07-11) and prospective (No. PJ2024-02-12) components. The informed consent was waived for the retrospective analysis due to its design, whereas all participants enrolled in the prospective phase provided written informed consent.
Research subjects
Patients with breast lesions from eight tertiary hospitals were enrolled in this study, including a retrospective cohort between February 2012 and March 9, 2024, and a prospective cohort at Hospital 1 between March 10 and September 2024. The inclusion criteria were as follows: (1) age ≥ 18 years; (2) breast lesions identified by ultrasound; (3) breast ultrasound performed within 2 weeks before biopsy or surgery. For the retrospective cohort, pathological confirmation of the lesion was required, whereas for the prospective cohort, patients were enrolled when biopsy or surgical treatment was planned and were included in the final analysis only if pathology was obtained. The exclusion criteria were: (1) prior biopsy of the target lesion before ultrasound examination, or a history of chemotherapy, radiotherapy, or other malignancies; (2) incomplete clinical or pathological data; (3) incomplete or poor-quality ultrasound images. In the prospective cohort, patients who did not ultimately undergo biopsy or surgery were also excluded. The patient enrollment process is illustrated in Supplementary Fig.S1.
In total, this study included data from 3142 eligible patients with breast lesions from eight hospitals for model training and testing. Among them, 2412 patients from three hospitals in the retrospective cohort were randomly divided into a training set (1688 patients) and an internal testing set (724 patients) in a 7:3 ratio. Additionally, 463 patients from the remaining five hospitals were included in a combined external testing set. Furthermore, 267 patients from Hospital 1 in the prospective cohort were included as the prospective testing set.
Clinic-pathologic data and ultrasound image collection
Ultrasound images were acquired using high-frequency linear-array transducers on the following systems: Resona 5S/6S/7S/7T/8, DC-8, and Nuewa R9 (Mindray), VOLUSON E8 and LOGIQ E9 (GE), iU22 and EPIQ 5/7/7 C (Philips), WS80A and RS80A (Samsung), ACUSON S2000 and ACUSON Sequoia (Siemens), AixPlorer (SuperSonic Imaging), MyLab 9 (Esaote), and ARIETTA 70 (Hitachi). During image acquisition, depth, gain, and focal zone were optimized to ensure adequate image quality, and the maximum lesion diameter was recorded. Representative images capturing key features of each breast mass were stored in the picture archiving and communication system (PACS) for subsequent analysis and validation.
Two senior radiologists with over 15 years of experience in breast ultrasound independently reviewed all ultrasound images without knowledge of pathology, assessing lesion features according to the fifth edition of BI-RADS. In the external and prospective test sets, they further classified lesions into BI-RADS categories 2, 3, 4a, 4b, 4c, and 5. Discrepancies were resolved by consensus. For patients with multiple ultrasound examinations, only the most recent scan within 2 weeks before biopsy/surgery was analyzed. The mean interval between the ultrasound examination and biopsy/surgery was 4.1 ± 2.9 days (range 0–13 days). The analysis was performed per patient, with one lesion included per patient. For multifocal lesions, only the largest lesion and its corresponding pathology result were included in the analysis.
Image segmentation and feature extraction
The lesion segmentation process was initially performed by a radiologist A ( > 7 years of experience) using ITK-SNAP 3.8.0 software (http://www.itksnap.org). The region of interest (ROI) was manually delineated along the lesion boundary on the stored static ultrasound image showing the largest cross-section of each lesion. Radiomics features were then extracted automatically in Python 3.10.6 using the PyRadiomics package (version 3.1.0) [23]. Before feature extraction, image normalization (normalizeScale = 25) and resampling (ResampledPixelSpacing = [1, 1, 1]) were performed [24].
To evaluate reliability, 50 patients were randomly selected from the training set for re-segmentation by radiologist A and another radiologist B. Intraclass correlation coefficients (ICCs) were calculated, and features with ICC ≥ 0.80 were retained.
Feature selection and machine learning classifier selection
Feature selection aimed to retain representative, stable, and clinically relevant features, as detailed in Supplementary MaterialS1. Based on the selected radiomics features, five commonly used machine learning classifiers were evaluated to identify the optimal predictive model. These classifiers included logistic regression (LR), Random Forest (RF), support vector machine (SVM), multi-layer perceptron (MLP), and K-nearest neighbor (KNN). Model performance was assessed by the area under the receiver operating characteristic (ROC) curve (AUC), accuracy, sensitivity, and specificity. The classifier with the best overall performance was chosen as the final radiomics model, and a radiomics score was calculated for each patient.
Model development and performance evaluation
In the training set, univariate LR was used to identify clinical ultrasound risk factors for lesion malignancy. Variables withp< 0.1 were entered into multivariate LR. A stepwise backward selection approach was applied, and variables withp< 0.05 were selected to construct the clinical model. Additionally, a combined model was developed by integrating the radiomics score with clinical risk factors.
Decision curve analysis (DCA) was conducted to visualize the net clinical benefit of different models in guiding clinical decision-making. Net reclassification improvement (NRI) and integrated discrimination improvement (IDI) were calculated to quantify the incremental value of the combined model. Model calibration was evaluated using calibration curves and the Hosmer-Lemeshow test. Furthermore, subgroup analyses were performed according to lesion size and patient age.
Clinical utility of the combined model in BI-RADS reclassification
Lesions in the external and prospective test sets were categorized by senior radiologists following BI-RADS recommendations: categories 2–3 as non-biopsy and categories 4–5 as biopsy. The combined model was applied to reclassify equivocal lesions: BI-RADS 4a lesions were downgraded to 3 if predicted low risk, and BI-RADS 3 lesions were upgraded to 4a if predicted high risk. Potential changes in biopsy recommendations after model-assisted BI-RADS reclassification and the malignancy proportion of BI-RADS 4a lesions were analyzed.
Model interpretation and visualization
The SHAP method was incorporated into this study to enhance the interpretability of both the radiomics model and the combined model. Rooted in game theory, SHAP quantifies the contribution of each feature to the final prediction, providing a deeper understanding of the model’s decision-making process. Specifically, SHAP not only highlights the impact of individual features on different samples but also reveals interactions between features. By employing SHAP, physicians can better review and validate the model’s predictions, further ensuring its reliability and fairness in practical applications.
Statistical analysis
Statistical analyses were performed using R software (version 4.2.2), MedCalc software (version 20.100), and SPSS software (version 24.0). The AUC values were compared using the DeLong test. Radiologists’ BI-RADS assessments were analyzed under three thresholds: (A) 4a, (B) 4b, and (C) 4c as cutoffs for malignancy. The McNemar test was used to compare the accuracy, sensitivity, and specificity between the combined model and radiologist assessments. A two-sidedp-value < 0.05 was considered statistically significant.
Results
Study population and baseline characteristics
A total of 3142 patients with breast lesions were included, with a mean age of 49.1 ± 13.0 years (range, 18–93 years), comprising 1121 (35.6%) benign and 2021 (64.3%) malignant lesions. The mean ages of patients in the training, internal test, external test, and prospective test sets were 48.9 ± 13.0 years (range, 18–92 years), 48.9 ± 12.8 years (range, 18–85 years), 48.9 ± 13.4 (range, 18–83 years), and 51.0 ± 12.1 years (range, 23–93 years), respectively. Baseline clinical-ultrasound characteristics are summarized in Table1.
Table: Clinical pathological characteristics of patients in the training set, internal test set, external test set, and prospective test set
Radiomics feature selection and development of the radiomics model
A total of 851 radiomics features were extracted from each ROI. After reproducibility screening, redundancy filtering, and least absolute shrinkage and selection operator (LASSO)-based selection, 12 representative features were retained to build the radiomics model. The complete list of selected features is provided in Supplementary MaterialS2. Five machine learning classifiers were then trained with these features.
Although the RF model achieved an AUC of 1.00 in the training set, its performance decreased substantially in the internal, external, and prospective test sets (AUC = 0.74, 0.71, and 0.70), suggesting overfitting and limited generalizability. In contrast, the LR model maintained stable performance with AUCs of 0.83, 0.82, 0.81, and 0.82 across the four datasets. Importantly, the LR model significantly outperformed the other classifiers in all three test sets (allp< 0.001). Based on its robustness and generalizability, LR was selected as the final radiomics model, and a radiomics score was generated for each patient. The detailed performance metrics of all classifiers are provided in Table2. Supplementary Fig.S2illustrates the discriminative performance of the LR model.
Table: Diagnostic performance evaluation of five machine learning models in each dataset
To enhance interpretability, SHAP analysis was applied to the LR model (Fig.1), providing both global feature importance (bar and beeswarm plots) and case-level explanations (waterfall plots).

Feature importance and SHAP value visualization of the LR model.ABar chart of average absolute SHAP values, ranking features by importance.BBeeswarm plot showing SHAP value distributions, with color indicating feature value (blue = low, red = high).C,DWaterfall plots for two patients illustrate feature contributions: red bars indicate positive, blue bars negative impacts. PatientCshows an overall negative prediction, while patientDshows an overall positive prediction
Construction of the clinical model
Univariate LR of clinical-ultrasound features identified age, tumor size, orientation, margin, shape, echotexture, and posterior features as significant predictive factors for malignancy. Multivariate analysis further retained age, tumor size, orientation, margin, and shape as independent predictors. These five variables were used to construct the clinical model. The clinical model achieved an AUC of 0.87 in the training set, 0.85 in the internal test set, 0.87 in the external test set, and 0.86 in the prospective test set. Detailed regression results are shown in Supplementary TableS2.
Construction and performance evaluation of the combined model
A combined model was constructed by integrating the radiomics score with the five independent clinical-ultrasound predictors (age, tumor size, orientation, margin, and shape). In multivariate analysis, all six variables remained significant predictors of malignancy. The combined model achieved AUCs of 0.92, 0.90, 0.92, and 0.93 in the training, internal, external, and prospective test sets, respectively. Across all datasets, the combined model significantly outperformed both the radiomics-only and clinical-only models (allp< 0.05, DeLong test). The ROC curves comparing model performance are shown in Fig.2, with detailed metrics summarized in Supplementary TableS3. The combined model demonstrated robust and balanced diagnostic performance across key patient subgroups, including patients with smaller lesions and younger age, achieving consistently high AUCs ranging from 0.87 to 0.94, along with balanced sensitivity and specificity across subgroups (Supplementary TableS4and Fig.3). This early presentation highlights the model’s reliability across clinically relevant subgroups.

ROC curves of the radiomics model, clinical model, and combined model in different datasets.Atraining set;Binternal test set;Cexternal test set;Dprospective test set

ROC curves of the combined model in different patient subgroups.Atraining set;Binternal test set;Cexternal test set;Dprospective test set
The DCA showed that the combined model provided a higher net clinical benefit than either the radiomics or clinical model alone across a wide range of threshold probabilities (0.05–0.95) in all datasets (Fig.4). Furthermore, both NRI and IDI analyses confirmed that incorporating the radiomics score significantly improved the discriminative ability of the clinical model (allp< 0.05; Supplementary TableS5). In addition, analysis of predicted risk probabilities demonstrated a clear distinction between benign and malignant breast lesions in all datasets, with malignant lesions showing markedly higher predicted probabilities (Supplementary Fig.S3). Model calibration was satisfactory, as demonstrated by calibration curves and Hosmer–Lemeshow tests across all datasets (Supplementary Fig.S4). SHAP visualizations provided global and individual-level interpretability by ranking feature importance and illustrating case-specific predictions (Fig.5).

DCA curves of the radiomics model, clinical model, and combined model in different datasets.Atraining set;Binternal test set;Cexternal test set;Dprospective test set

Feature importance and SHAP visualization of the combined model.ABar chart ranking features by mean absolute SHAP values.BSwarm plot showing SHAP value distributions.C,DCase-specific waterfall plots with corresponding ultrasound images illustrate feature contributions. For patientC(f(x) = 3.66, invasive cancer), radiomics score, margin, shape, and tumor size contribute positively, while orientation is negative. For patientD(f(x) = −0.36, adenosis), margin and shape increase prediction, whereas age, radiomics score, tumor size, and orientation decrease it
Model performance compared with radiologists
In both the external and prospective test sets, we compared the diagnostic performance of the combined model with BI-RADS assessments made by experienced radiologists using different malignancy cutoff thresholds. When BI-RADS 4a was used as the cutoff, radiologist assessment achieved very high sensitivity but extremely low specificity. In contrast, the combined model significantly improved overall accuracy and specificity in both datasets, although its sensitivity was lower than that of BI-RADS 4. Using BI-RADS 4b as the cutoff, the combined model showed comparable overall accuracy to radiologist assessment, with significantly higher specificity but lower sensitivity. When BI-RADS 4c was applied, the overall diagnostic performance of the combined model and radiologist assessment was largely comparable, with no consistent significant differences in accuracy across datasets (Supplementary TableS6).
Clinical utility of the combined model in BI-RADS reclassification
Using BI-RADS-based recommendations made by experienced radiologists, unnecessary biopsy rates were 29.67% (119/401) in the external test set and 24.21% (61/252) in the prospective test set. When biopsy recommendations were generated solely according to the combined model, the unnecessary biopsy rates were 10.22% (28/274) in the external test set and 10.47% (20/191) in the prospective test set, which were significantly lower than those based on BI-RADS recommendations (bothp< 0.001).
Furthermore, the combined model was used to adjust BI-RADS 3 and 4a categories. After model-based reclassification, unnecessary biopsy rates decreased from 29.67% to 18.84% in the external test set (absolute reduction, 10.83%) and from 24.21% to 14.41% in the prospective test set (absolute reduction, 9.80%), without a significant reduction in diagnostic sensitivity. Notably, the malignancy rate among BI-RADS 4a lesions increased markedly after reclassification, from 9.23% to 44.44% in the external test set and from 16.67% to 50.00% in the prospective test set (bothp< 0.05). Detailed results are summarized in Table3.
Table: Performance of the combined model and BI-RADS classification in recommending biopsy for breast lesions
Discussion
This study extracted radiomics features from ultrasound images of breast lesions to evaluate the performance of multiple machine learning models in distinguishing benign from malignant lesions. By integrating clinical ultrasound features, a combined predictive model was developed and validated across multiple datasets and patient subgroups. The model also showed potential value in BI-RADS reclassification and biopsy decision-making.
Early and accurate diagnosis of breast cancer is essential for improving prognosis [25]. However, conventional ultrasound diagnosis depends heavily on physician expertise, and the BI-RADS system, while standardized, remains limited in diagnostic certainty. In particular, BI-RADS category 4 often presents overlapping imaging features between benign and malignant lesions, leading to high false-positive rates [26]. Radiomics provides a more objective and quantitative approach by extracting high-dimensional features beyond visual interpretation. This study extracted radiomics features that included shape descriptors, first-order statistical features, and various texture-based high-order features. Among them, wavelet-transformed gray-level and texture features were the most common. For example, “wavelet.LLL_glszm_ZonePercentage” describes the uniformity of textures within the lesion, and “wavelet.LLL_glrlm_RunLengthNonUniformityNormalized” captures gray-level consistency. These features characterize the microstructural heterogeneity of breast tumors [27], information that is difficult to capture with conventional ultrasound [28].
After completing feature selection, we constructed predictive models using five machine learning algorithms. The LR model demonstrated consistent performance across all datasets. The AUCs were 0.83, 0.82, 0.81, and 0.82 in the training, internal test, external test, and prospective test sets, respectively. In all three test sets, the LR model significantly outperformed the other algorithms. This may be attributed to the linear combination of feature weights in LR, which helps prevent overfitting and is suitable for radiomics datasets with fewer but informative features. These findings are in line with previous research by Ye et al [29]. Therefore, the LR model was selected as the final radiomics model. Furthermore, SHAP analysis was applied to improve interpretability, clarifying both the overall and individual contributions of radiomics features to model predictions.
In addition to radiomics features, traditional clinical ultrasound indicators also play an important role in predicting breast lesion malignancy. In this study, age, tumor size, orientation, margin, and shape were identified as independent predictors, consistent with previous studies [30–32]. Notably, orientation, margin, and shape are standard ultrasound descriptors from the BI-RADS lexicon, whereas age and tumor size are objective clinical/measurement variables that may provide complementary information beyond imaging appearance. We did not use the final BI-RADS assessment category as an input feature because it represents a composite, reader-dependent conclusion and was evaluated separately for biopsy-decision analysis and comparison with our models. The predictive value of age may be related to hormone-driven structural changes in breast tissue, especially with aging and menopausal status [33,34]. Although the clinical model achieved good performance, it was less effective than the radiomics model in capturing subtle imaging details, particularly for complex or ambiguous lesions.
To leverage the complementary strengths of both approaches, we constructed a combined model that incorporated radiomics scores with clinical and ultrasound factors. This model achieved consistently high AUCs across all datasets (0.92, 0.90, 0.92, and 0.93), significantly outperforming either model alone (allp< 0.01). NRI and IDI analyses confirmed that radiomics scores added discriminative value beyond clinical features. Furthermore, across most threshold probability ranges across all test sets, the combined model showed higher net benefit in decision-curve analysis than the radiomics and clinical models alone. These results highlight that radiomics captures complementary information, particularly wavelet-based texture features, which are not accessible through traditional clinical indicators [35–37]. Importantly, SHAP value analysis clarified the contribution of each variable to the model’s output. Clinicians can examine SHAP plots to observe how values of specific features influence the prediction for each individual case. This type of quantitative interpretability enhances clinical acceptance of the model and supports its future implementation in practice.
The combined model also performed well across patient subgroups, achieving balanced sensitivity and specificity in smaller lesions and younger patients, which is important for early detection and diagnosis of breast cancer. Prior studies have suggested that combining radiomics with conventional ultrasound enhances early diagnostic accuracy [17,29,35–37], but most were limited by small cohorts and insufficient validation. By leveraging multicenter data and prospective testing, this study ensured broader generalizability and robustness, with consistent performance across different hospitals and ultrasound systems.
Compared with senior radiologists, the combined model achieved significantly higher accuracy than BI-RADS at 4a as the cutoff, while comparable accuracy was observed at higher thresholds (4b or 4c), consistent with previous findings [38]. Importantly, model-assisted reclassification of BI-RADS categories suggested the potential to reduce biopsy recommendations for benign lesions and increased the proportion of malignant lesions among BI-RADS 4a cases. These findings suggest that the combined model may support biopsy decision-making and lesion risk stratification, pending further prospective clinical utility studies. However, both the clinical and combined models incorporated reader-derived ultrasound descriptors assessed by experienced radiologists (including orientation, margin, and shape); therefore, comparisons with BI-RADS assessment should not be interpreted as purely automated human-versus-machine comparisons.
This study has several limitations. First, the study population was highly selected because only lesions with pathological confirmation were included, resulting in a relatively high prevalence of malignancy and a cohort that is not fully representative of routine diagnostic ultrasound practice. This cohort composition may have influenced the apparent model performance, decision-curve analysis findings, and biopsy reclassification estimates. Therefore, these results should be interpreted with caution, as they may not directly translate to lower-prevalence real-world settings. Second, although data from multiple centers were included, variations in ultrasound equipment and protocols across centers may have influenced image quality and model performance, suggesting the need for standardized acquisition in future studies. Finally, the radiomics features were manually extracted. In the future, deep learning-based automatic feature extraction methods may be considered to reduce manual intervention during feature selection and improve the efficiency and automation of the modeling process.
Conclusions
In summary, this study developed and validated a combined model integrating radiomics and clinical features for breast cancer diagnosis. Multicenter and prospective validation confirmed its accuracy, robustness, and interpretability. By showing improved diagnostic performance and interpretable prediction patterns, the model may support breast lesion risk stratification and biopsy decision-making, pending further prospective clinical utility studies.
Acknowledgements
The authors thank all radiologists of the participating hospitals for assisting with the collection of the imaging data used in this study.
Author contributions
D.Z.: conceptualization, data curation, formal analysis, writing—original draft, writing—review & editing, WL: data curation, formal analysis, software, validation, X.Q.: data curation, formal analysis, W.Z.: data curation, formal analysis, X.Z.: formal analysis, methodology, writing—review & editing, Y.L.: data curation, L.W.: data curation, J.W.: data curation, J.L.W.: data curation, J.Z.: data curation, L.Z.: data curation, C.Z.: funding acquisition, project administration, resources, supervision, writing—review & editing. All authors reviewed the manuscript and approved its final version for publication.
Funding
This work was supported by the Natural Science Foundation of Anhui Province (Grant number: 2308085MH278) and the Health Research Program of Anhui Province (Grant number: AHWJ2023A10017).
Data availability
The datasets used and analyzed during the current study are available from the corresponding author on reasonable request.
Declarations
Ethics approval and consent to participate
This study received approval from the Institutional Ethics Committee, with separate approval numbers for the retrospective (No. PJ2023-07-11) and prospective (No. PJ2024-02-12) components.
Consent for publication
The informed consent was waived for the retrospective analysis due to its design, whereas all participants enrolled in the prospective phase provided written informed consent.
Competing interests
The authors have no competing interests to declare.
Supplementary information
The online version contains supplementary material available at 10.1186/s13244-026-02344-y.
Associated Data
Data Availability Statement
The datasets used and analyzed during the current study are available from the corresponding author on reasonable request.
References
- Filho AM, Laversanne M, Ferlay J et al (2024) The GLOBOCAN 2022 cancer estimates: Data sources, methods, and a snapshot of the cancer burden worldwide. Int J Cancer 156:1336–1346 doi.org/10.1002/ijc.35278
- Freeman K, Geppert J, Stinton C et al (2021) Use of artificial intelligence for image analysis in breast cancer screening programmes: systematic review of test accuracy. BMJ 374:n1872 doi.org/10.1136/bmj.n1872
- Pashayan N, Antoniou AC, Ivanus U et al (2020) Personalized early detection and prevention of breast cancer: ENVISION consensus statement. Nat Rev Clin Oncol 17:687–705 doi.org/10.1038/s41571-020-0388-9
- Mann RM, Hooley R, Barr RG, Moy L (2020) Novel approaches to screening for breast cancer. Radiology 297:266–285 doi.org/10.1148/radiol.2020200172
- Shen Y, Shamout FE, Oliver JR et al (2021) Artificial intelligence system reduces false-positive findings in the interpretation of breast ultrasound exams. Nat Commun 12:5645 doi.org/10.1038/s41467-021-26023-2
- Mann RM, Kuhl CK, Moy L (2019) Contrast-enhanced MRI for breast cancer screening. J Magn Reson imaging 50:377–390 doi.org/10.1002/jmri.26654
- Yang L, Wang S, Zhang L et al (2020) Performance of ultrasonography screening for breast cancer: a systematic review and meta-analysis. BMC Cancer 20:499 doi.org/10.1186/s12885-020-06992-1
- Gu Y, Tian JW, Ran HT et al (2022) The utility of the fifth edition of the BI-RADS ultrasound lexicon in category 4 breast lesions: a prospective multicenter study in China. Acad Radiol 29:S26–s34 doi.org/10.1016/j.acra.2020.06.027
- Spak DA, Plaxco JS, Santiago L, Dryden MJ, Dogan BE (2017) BI-RADS(®) fifth edition: A summary of changes. Diagn Intervent Imaging 98:179–190 doi.org/10.1016/j.diii.2017.01.001
- He P, Cui LG, Chen W, Yang RL (2019) Subcategorization of ultrasonographic BI-RADS category 4: assessment of diagnostic accuracy in diagnosing breast lesions and influence of clinical factors on positive predictive value. Ultrasound Med Biol 45:1253–1258 doi.org/10.1016/j.ultrasmedbio.2018.12.008
- Mercado CL (2014) BI-RADS update. Radiol Clin North Am 52:481–487 doi.org/10.1016/j.rcl.2014.02.008
- Stavros AT, Freitas AG, deMello GGN et al (2017) Ultrasound positive predictive values by BI-RADS categories 3-5 for solid masses: an independent reader study. Eur Radiol 27:4307–4315 doi.org/10.1007/s00330-017-4835-7
- Elverici E, Barça AN, Aktaş H et al (2015) Nonpalpable BI-RADS 4 breast lesions: sonographic findings and pathology correlation. Diagn Intervent Radiol 21:189–194 doi.org/10.5152/dir.2014.14103
- Bera K, Braman N, Gupta A, Velcheti V, Madabhushi A (2022) Predicting cancer outcomes with radiomics and artificial intelligence in radiology. Nat Rev Clin Oncol 19:132–146 doi.org/10.1038/s41571-021-00560-7
- Romeo V, Cuocolo R, Apolito R et al (2021) Clinical value of radiomics and machine learning in breast ultrasound: a multicenter study for differential diagnosis of benign and malignant lesions. Eur Radiol 31:9511–9519 doi.org/10.1007/s00330-021-08009-2
- Qi YJ, Su GH, You C et al (2024) Radiomics in breast cancer: Current advances and future directions. Cell Rep Med 5:101719 doi.org/10.1016/j.xcrm.2024.101719
- Li X, Zhang L, Ding M (2024) Ultrasound-based radiomics for the differential diagnosis of breast masses: a systematic review and meta-analysis. J Clin ultrasound 52:778–788 doi.org/10.1002/jcu.23690
- Gu J, Jiang T (2022) Ultrasound radiomics in personalized breast management: current status and future prospects. Front Oncol 12:963612 doi.org/10.3389/fonc.2022.963612
- Wang H, Zha H, Du Y, Li C, Zhang J, Ye X (2023) An integrated radiomics nomogram based on conventional ultrasound improves discriminability between fibroadenoma and pure mucinous carcinoma in breast. Front Oncol 13:1170729 doi.org/10.3389/fonc.2023.1170729
- Yang L, Ma Z (2023) Nomogram based on super-resolution ultrasound images outperforms in predicting benign and malignant breast lesions. Breast Cancer 15:867–878 doi.org/10.2147/BCTT.S435510
- Ma Q, Lu X, Qin X et al (2023) A sonogram radiomics model for differentiating granulomatous lobular mastitis from invasive breast cancer: a multicenter study. La Radiol Med 128:1206–1216 doi.org/10.1007/s11547-023-01694-7
- Shi S, An X, Li Y (2023) Ultrasound radiomics-based logistic regression model to differentiate between benign and malignant breast nodules. J Ultrasound Med 42:869–879 doi.org/10.1002/jum.16078
- van Griethuysen JJM, Fedorov A, Parmar C et al (2017) Computational radiomics system to decode the radiographic phenotype. Cancer Res 77:e104–e107 doi.org/10.1158/0008-5472.CAN-17-0339
- Zwanenburg A, Vallières M, Abdalah MA et al (2020) The Image biomarker standardization initiative: standardized quantitative radiomics for high-throughput image-based phenotyping. Radiology 295:328–338 doi.org/10.1148/radiol.2020191145
- Hong R, Xu B (2022) Breast cancer: an up-to-date review and future perspectives. Cancer Commun 42:913–936 doi.org/10.1002/cac2.12358
- Lazarus E, Mainiero MB, Schepps B, Koelliker SL, Livingston LS (2006) BI-RADS lexicon for US and mammography: interobserver variability and positive predictive value. Radiology 239:385–391 doi.org/10.1148/radiol.2392042127
- Mayerhoefer ME, Materka A, Langs G et al (2020) Introduction to radiomics. J Nucl Med 61:488–495 doi.org/10.2967/jnumed.118.222893
- Gillies RJ, Kinahan PE, Hricak H (2016) Radiomics: Images are more than pictures, they are data. Radiology 278:563–577 doi.org/10.1148/radiol.2015151169
- Ye J, Chen Y, Pan J et al (2024) US-based radiomics analysis of different machine learning models for differentiating benign and malignant BI-RADS 4A breast lesions. Acad Radiol 32:67–78 doi.org/10.1016/j.acra.2024.08.024
- Wu P, Jiang Y, Xing H et al (2023) Multimodality deep learning radiomics nomogram for preoperative prediction of malignancy of breast cancer: a multicenter study. Phys Med Biol 68. 10.1088/1361-6560/acec2d. doi.org/10.1088/1361-6560/acec2d
- Guo R, Lu G, Qin B, Fei B (2018) Ultrasound imaging technologies for breast cancer detection and management: a review. Ultrasound Med Biol 44:37–70 doi.org/10.1016/j.ultrasmedbio.2017.09.012
- Niu Q, Zhao L, Wang R et al (2024) Predictive value of contrast-enhanced ultrasonography and ultrasound elastography for management of BI-RADS category 4 nonpalpable breast masses. Eur J Radiol 173:111391 doi.org/10.1016/j.ejrad.2024.111391
- Sherman ME, de Bel T, Heckman MG et al (2022) Serum hormone levels and normal breast histology among premenopausal women. Breast Cancer Res Treat 194:149–158 doi.org/10.1007/s10549-022-06600-9
- Michaels E, Worthington RO, Rusiecki J (2023) Breast cancer: risk assessment, screening, and primary prevention. Med Clin North Am 107:271–284 doi.org/10.1016/j.mcna.2022.10.007
- Su HZ, Hong LC, Su YM, Chen XS, Zhang ZB, Zhang XD (2024) A nomogram based on conventional ultrasound radiomics for differentiating between radial scar and invasive ductal carcinoma of the breast. Ultrasound Q 40:e00685 doi.org/10.1097/RUQ.0000000000000685
- Luo WQ, Huang QX, Huang XW, Hu HT, Zeng FQ, Wang W (2019) Predicting Breast Cancer in Breast Imaging Reporting and Data System (BI-RADS) ultrasound category 4 or 5 lesions: a nomogram combining radiomics and BI-RADS. Sci Rep 9:11921 doi.org/10.1038/s41598-019-48488-4
- Hong ZL, Chen S, Peng XR, Li JW, Yang JC, Wu SS (2022) Nomograms for prediction of breast cancer in breast imaging reporting and data system (BI-RADS) ultrasound category 4 or 5 lesions: a single-center retrospective study based on radiomics features. Front Oncol 12:894476 doi.org/10.3389/fonc.2022.894476
- Wei Q, Yan YJ, Wu GG et al (2022) The diagnostic performance of ultrasound computer-aided diagnosis system for distinguishing breast masses: a prospective multicenter study. Eur Radiol 32:4046–4055 doi.org/10.1007/s00330-021-08452-1
Republished from the open web under CC-BY. Authors: Zhang D, Lu WW, Qin XC, Zhou W, Zhang XY, Luo YH, Wu LS, Wu J, Wang JL, Zhao JJ, Zhang L, Zhang CX. Read the original.