Medicine

A perioperative multi-modal fusion and deep learning-based prognostic system for upper tract urothelial carcinoma: a multi-institutional study.

Peng X, Li Y, Shi W, Xiao B, Xiao X, Yue X, Xv Q, Jiang Q, He W, Xv Y, Xiao M. Published June 25, 2026 CC-BY

Objectives Precise perioperative risk stratification for upper tract urothelial carcinoma (UTUC) is essential. We developed a multimodal prognostic model integrating perioperative clinical data, radiomics, and deep learning (DL) features from baseline CT urography to improve survival prediction and guide adjuvant management. Materials and methods We retrospectively enrolled 623 patients from six institutions, divided into training, internal validation, and independent external validation sets. Four single-modal models (clinical, radiomics, 2D DL, and 2.5D DL) were developed, and an integrated combined model was constructed by fusing their prognostic scores. Performance was evaluated using the C-index, area under the curve (AUC), calibration curves, and decision curve analysis (DCA). Results The combined model consistently outperformed all single-modal models across all cohorts. C-indices reached 0.758 (95% CI: 0.712-0.804), 0.725 (95% CI: 0.651-0.798), and 0.704 (95% CI: 0.631-0.777) in the training, internal validation, and external validation sets, respectively, numerically surpassing the best single-modal models. Notably, our 2.5D DL model (C-index: 0.705) demonstrated a consistent incremental improvement over the 2D DL model (C-index: 0.681) in capturing prognostic information. In external validation, the combined model achieved a 3-year AUC of 0.766. DCA indicated the comprehensive model exhibited excellent calibration and provided the highest net benefits. Conclusion This multimodal system, featuring a robust 2.5D DL strategy, improves overall survival prediction in UTUC. It offers a valuable tool for accurate perioperative risk stratification immediately after radical nephroureterectomy, demonstrating particularly reliable value for 3-year intermediate-term clinical decision-making. Critical relevance statement This multimodal system advances clinical radiology by fusing perioperative clinical data, radiomics, and DL features from CTU images, enhancing risk stratification accuracy to guide postoperative adjuvant management for UTUC. Key points Current prognostic models for UTUC lack accuracy, creating an unmet clinical need for precise perioperative risk stratification to guide adjuvant management. A multimodal prognostic model fusing clinical, radiomic, and DL features from baseline CT urography consistently outperformed single-modality models in predicting overall survival.

Introduction

Upper urinary tract urothelial carcinoma (UTUC) is a relatively rare but aggressive malignancy originating in the renal pelvis and ureteral urothelium [1,2]. Although it accounts for only 5%–10% of all urothelial carcinomas, UTUC is characterized by a high recurrence rate and rapid progression [1,35]. While radical nephroureterectomy (RNU) remains the standard treatment for localized UTUC, accurate risk stratification is crucial for guiding perioperative management, particularly regarding the administration of adjuvant chemotherapy and the intensity of follow-up surveillance [1,6]. However, current prognostic systems, which primarily rely on conventional clinicopathological features such as tumor stage and grade, are limited in their predictive ability and fail to fully capture the complexity and heterogeneity of the disease [79]. Furthermore, current preoperative staging based solely on CT (cT) is often inaccurate compared to pathological staging (pT) [1,6], which may introduce substantial noise and limit the reliability of clinical decision-making. Consequently, there is an urgent clinical need for a more comprehensive and accurate prognostic model.

Recently, artificial intelligence (AI), including radiomics and deep learning (DL), has opened up promising new avenues for medical prognosis by decoding tumor heterogeneity [1012]. Radiomics extracts high-throughput quantitative features from standard medical images, while DL can automatically learn high-dimensional feature patterns directly from original data [1316]. These technologies allow for the identification of profound image features closely associated with clinical prognosis that are often difficult to discern through visual inspection [1719]. The deep integration of radiomics and DL offers a potent technical tool for constructing a more precise and quantifiable prognostic model for UTUC, which is hypothesized to overcome the current impasse in clinical risk assessment.

Despite these advancements, existing UTUC prognostic research often relies on single-modal approaches. It must be explicitly clarified that applying multiple algorithms—such as both radiomics and DL—exclusively to CT images still fundamentally constitutes a single-modality approach. Such single-source imaging strategies inherently lack the capacity to capture the complete clinical landscape. We posit that a true multimodal framework, which integrates these high-dimensional imaging phenotypes with complementary clinical and pathological data, can provide biological information that a single source may not capture [20,21]. The synergistic integration of multimodal data is essential for a comprehensive understanding of tumor behavior. Notable progress has been made in specific oncological fields by integrating multimodal data [22,23]. For example, in their research on pediatric low-grade gliomas (pLGGs), Mahootiha et al [24] constructed a comprehensive multimodal DL model that significantly improved the accuracy of predicting postoperative recurrence by combining traditional clinical indicators with DL features automatically extracted from preoperative MRI scans. Therefore, the primary goal of this study is to construct a robust multimodal fusion model.

However, the performance of such a multimodal model depends heavily on the quality of feature extraction from the imaging modality. In this study, we constructed a combined prognostic model by integrating perioperative clinical data, radiomic features, and DL signatures, which markedly enhanced the accuracy of overall survival (OS) prediction in UTUC. To enhance the imaging component within this fusion framework, we introduce an innovative 2.5D multi-view DL strategy. This strategy effectively captures multiplanar information and reveals tumor heterogeneity more comprehensively than traditional 2D slice analysis. Through a large-scale, multicenter cohort study, we rigorously validated the model’s robustness and generalizability, fully demonstrating its potential for perioperative clinical application.

Methods

Patients and study design

This study was ethically approved (CY2024-308-02) with informed consent waived. In this study, perioperative risk stratification is explicitly defined as the risk assessment performed immediately following RNU. At this clinical time point, the model integrates baseline preoperative imaging features with the newly available postoperative pathological staging to provide a comprehensive evaluation for subsequent adjuvant management. The overall workflow of this study is illustrated in Fig.1. This multicenter retrospective study was approved by the Ethics Committee of our institutions, and the requirement for informed consent was waived; it was conducted in accordance with the Declaration of Helsinki. This study follows the CLEAR checklist guidelines [25]. A total of 623 patients with pathologically confirmed upper tract urothelial carcinoma (UTUC) who underwent RNU between December 2012 and December 2024 were included. Inclusion criteria were as follows: (1) histologically confirmed UTUC and (2) availability of preoperative computed tomography urography (CTU) data. Patients were excluded if they had poor image quality, received prior oncologic treatment before CTU, had a coexisting primary malignancy, or lacked complete clinicopathological or follow-up data. The detailed patient selection process is illustrated in Fig. S1. Patients from Centers 1 and 2 (n= 473) were randomly divided into a training set (n= 331) and an internal validation set (n= 142) at a 7:3 ratio. For independent validation, an external validation set was composed of 150 patients from four other centers (Centers 3, 4, 5, and 6). The primary endpoint of this study was OS, defined as the time from the initial diagnosis to death from any cause. Patients with limited follow-up duration (e.g., those enrolled in 2024) were included in the risk set, and their survival data were validly treated as censored at their last follow-up date.Fig. 1Overall workflow schematic. The comprehensive workflow for predicting the OS of patients with UTUC using CT images is shown.AImage acquisition and processing involved collecting baseline CTU scans, segmenting ROIs using ITK-SNAP software.BClinical features were collected and a clinical model was constructed using Cox regression analysis.CThe radiomics model was built by extracting first-order, shape, and texture features, followed by feature selection using Lasso-Cox regression analysis.DThe DL model and DL25D model were constructed using a pre-trained ResNet-50 network for feature extraction from 2D and 2.5D CT inputs, respectively, followed by Cox regression analysis.EThe combined model was established by integrating the clinical, radiomics, and DL model scores using Cox regression. The performance of all models was evaluated using ROC curves, calibration curves, and DCA.FThe prognostic value of the final Combined Model was assessed by stratifying patients into high- and low-risk groups based on a cutoff value. OS, overall survival; UTUC, upper tract urothelial carcinoma; CTU, computed tomography urography; ROIs, regions of interest; DL, deep learning; ROC, receiver operating characteristic; DCA, decision curve analysis.

Image acquisition and preprocessing

Preoperative CTU images were retrieved in digital imaging and communications in medicine (DICOM) format from the Picture Archiving and Communication System for each institution. Acquisition protocols from the six centers are summarized in Table S1. To ensure consistency across institutions, the corticomedullary phase images were consistently selected for analysis, as this phase provided optimal tumor-to-background contrast. To minimize inter-scanner variability, all images underwent a standardized preprocessing pipeline: (1) Voxel spacing normalization: All CTU images were resampled to an isotropic voxel spacing of 1 × 1 × 1 mm³ using cubic interpolation. (2) Intensity normalization and contrast adjustment: The window level and width were standardized to 50 HU and 200 HU (range: −50 to 150 HU). (3) Region of interest (ROIs) segmentation: The Regions of Interest (ROIs) were manually delineated on the axial CT slices. To ensure segmentation precision, all delineations were uniformly performed on the corticomedullary phase images. This segmentation process was conducted by two experienced radiologists using ITK-SNAP software. Discrepancies were resolved by a senior radiologist (20 years of experience). To quantitatively assess inter-observer reproducibility, 30 cases were randomly selected and re-segmented independently by the second radiologist to calculate the intraclass correlation coefficient (ICC, two-way mixed-effects model, single rater, absolute agreement).

Clinical and radiomics model construction

Clinical variables, including age, sex, tumor size, pT stage, grade, lymphovascular invasion (LVI), and hydronephrosis, were extracted from medical records. Continuous variables were tested for normality (Shapiro–Wilk test) and compared using thet-test or Mann–WhitneyU-test; categorical variables were compared using the Chi-square (χ2) test. In our analysis, the pathologic T stage (pT) was further simplified into a binary variable, pT_binary, where patients with stages T0/Tx/T1 were classified as non-muscle invasive (< T2), and those with stages T2, T3, and T4 were classified as muscle-invasive (≥ T2). This binary classification was chosen because the distinction between organ-confined and muscle-invasive disease is a critical clinical determinant for adjuvant chemotherapy and lymph node dissection indications [1].

Radiomic features were extracted from segmented ROIs using PyRadiomics (v3.0.1) following the Imaging Biomarker Standardization Initiative (IBSI) guidelines. Prior to feature reduction,Z-score normalization was applied to standardize the distribution of the extracted radiomics features. After voxel resampling and intensity discretization (64 gray levels), seven feature classes were computed: first-order, shape, GLCM, GLRLM, GLSZM, GLDZM, and NGTDM. Initially, 1288 radiomics features were extracted from the defined ROIs. The four-step feature reduction process yielded the following feature counts: (1) reproducibility filtering retained 864 features with ICC ≥ 0.85; (2) correlation filtering removed highly redundant variables, reducing the pool to 187 features; (3) univariate Cox analysis identified 28 features significantly associated with OS (p< 0.05); and (4) Finally, LASSO-Cox regression with 10-fold cross-validation selected an optimal subset of 15 features to construct the final radiomic signature (Fig. S2).

DL model construction

The DL models were developed using a transfer learning approach.

Data preparation

To prepare the data, we adopted two distinct input strategies. For the DL (2D) model, we used the max-axial slice of the tumor as the representative input. For the DL25D (2.5D) model, a multi-view ROI extraction strategy was adopted, where the max-axial slice, max-sagittal slice, and max-coronal slice were extracted from each 3D tumor volume and stacked to form a multi-channel input (Fig.1D). To ensure spatial alignment and preserve high-resolution intratumoral heterogeneity across different network architectures, the extracted ROIs were cropped and resized to 448×448 pixels using bicubic interpolation. Image intensities were subsequently Z-score normalized to standard deviation units. This 2.5D approach was designed to integrate global cross-sectional multiplanar features and simultaneously mitigate the memory limitations and severe overfitting risks associated with full 3D processing or adjacent-slice sequence modeling on clinical datasets.

Model architecture and transfer learning

We tested four well-known architectures—ResNet50, Vgg19, DenseNet121, and Inception_v3—using weights that had already been trained on ImageNet. The initial convolutional layer was reconfigured to receive multi-channel inputs, and the late-stage features were fed into a fully connected layer with L2 regularization, followed by a single-node output for log-risk prediction. The network was trained using the Cox partial likelihood loss to model non-linear relationships between imaging-derived features and survival outcomes [26]. Detailed hyperparameter configurations, including the learning rate scheduler, optimization strategies, and regularization techniques, are provided inSupplementary materials.

Model interpretability

To offer commentary on the model’s decision-making process, we used gradient-weighted class activation mapping (Grad-CAM) to generate heatmaps [27]. These visualizations highlight the specific tumor regions that were most influential for the model’s prognostic predictions.

Statistical analysis and model evaluation

All data analyses were executed using Python (v3.7.12) on the OnekeyAI platform. Our DL frameworks were developed on PyTorch (v1.11.0), optimized for enhanced performance through CUDA (v11.3.1) and cuDNN (v8.2.1).

The prognostic scores derived from each optimized single-modality model were entered into a multivariate Cox regression with stepwise selection to construct the final combined prognostic model. For survival analysis, Kaplan–Meier (KM) curves were generated, and samples were stratified based on the predicted hazard ratios (HRs), with the significance of group separation assessed via the log-rank test. The discrimination performance of the models was evaluated using the Concordance index (C-index) and time-dependent area under the curve (AUC). To robustly assess internal and external validity, 95% confidence intervals (CIs) for the metrics were calculated using bootstrap resampling with 1000 iterations. Calibration was evaluated using calibration curves and quantified by the Integrated Brier Score, reflecting the agreement between predicted and observed survival probabilities. Furthermore, decision curve analysis (DCA) was utilized to evaluate the clinical net benefit of the models across various threshold probabilities.

Results

Patient characteristics and cohort distribution

As shown in Table1, a total of 623 patients with UTUC were enrolled from six centers and divided into a training cohort (n= 331), an internal validation cohort (n= 142), and an external validation cohort (n= 150). The mean age of all enrolled patients was 68.21 years (± 10.06). The majority of patients were male (58.9%) and had a tumor size greater than or equal to 2 cm (74.0%). The baseline clinical and pathological characteristics were statistically similar across the three cohorts, with no significant differences found in key variables such as age, gender, tumor size, laterality, location, grade, margin, LVI, and hydronephrosis (p> 0.05). However, significant differences were observed in pT stage (p= 0.038) and pT_binary (p= 0.006). The median follow-up duration for the entire cohort was 24.03 months (IQR: 13.68–32.28).Table 1Baseline characteristics of UTUC patientsCharacteristicAll patientsTraining setInternal validation setExternal validation setp(n= 623)(n= 331)(n= 142)(n= 150)Age68.212 ± 10.05867.644 ± 10.20468.634 ± 9.79469.067 ± 9.9660.303Gender0.782Female256 (41.1%)138 (41.7%)60 (42.3%)58 (38.7%)Male367 (58.9%)193 (58.3%)82 (57.7%)92 (61.3%)Size (cm)3.142 ± 2.0803.092 ± 1.9222.995 ± 1.7993.393 ± 2.5950.214Size_binary0.443< 2 cm162 (26.0%)93 (28.1%)33 (23.2%)36 (24.0%)≥ 2 cm461 (74.0%)238 (71.9%)109 (76.8%)114 (76.0%)Laterality0.35Left342 (54.9%)189 (57.1%)78 (54.9%)75 (50.0%)Right281 (45.1%)142 (42.9%)64 (45.1%)75 (50.0%)Location0.122Pelvis284 (45.6%)144 (43.5%)75 (52.8%)65 (43.3%)Ureter303 (48.6%)168 (50.8%)63 (44.4%)72 (48.0%)Both36 (5.8%)19 (5.7%)4 (2.8%)13 (8.7%)Grade0.687High496 (79.6%)262 (79.2%)111 (78.2%)123 (82.0%)Low127 (20.4%)69 (20.8%)31 (21.8%)27 (18.0%)pT stage0.038Tx/T083 (13.2%)52 (15.7%)20 (14.1%)11 (7.3%)T1182 (29.2%)101 (30.5%)45 (31.7%)36 (24.0%)T2176 (28.3%)92 (27.8%)41 (28.9%)43 (28.7%)T3163 (26.2%)78 (23.6%)33 (23.2%)52 (34.7%)T419 (3.1%)8 (2.4%)3 (2.1%)8 (5.3%)pT_binary0.006< T2265 (42.5%)153 (46.2%)65 (45.8%)47 (31.3%)≥ T2358 (57.5%)178 (53.8%)77 (54.2%)103 (68.7%)Margin0.337Negative598 (96.0%)315 (95.2%)136 (95.8%)147 (98.0%)Positive25 (4.0%)16 (4.8%)6 (4.2%)3 (2.0%)LVI0.504No557 (89.4%)296 (89.4%)130 (91.5%)131 (87.3%)Yes66 (10.6%)35 (10.6%)12 (8.5%)19 (12.7%)Hydronephrosis0.577No277 (44.5%)145 (43.8%)60 (42.3%)72 (48.0%)Yes346 (55.5%)186 (56.2%)82 (57.7%)78 (52.0%)Median follow-up, months24.03 (13.68 - 32.28)23.18 (13.26–29.98)17.63 (13.54–30.81)29.90 (15.59–36.08)0.086

Performance of clinical and radiomics models

We first evaluated the predictive power of two single-modal models. The clinical model, constructed using multivariate Cox regression (Fig.1B), identified pT_binary and tumor size as significant prognostic factors (Fig.2A). This model achieved a C-index of 0.648 (95% CI: 0.596–0.699) in the training cohort and was validated with C-indexes of 0.613 (95% CI: 0.533–0.694) and 0.604 (95% CI: 0.525–0.682) in the internal and external validation sets, respectively (Fig.2B–Dand Table2).Fig. 2Construction and prognostic value of the clinical model.AUnivariate and multivariate Cox regression analysis was performed to identify significant clinical features for the clinical model. The forest plot shows the HR and 95% CI for each variable.B–DKM survival curves for the Clinical Model are shown in the training set (B), internal validation set (C), and external validation set (D). HR, hazard ratios; CI, confidence intervals; KM, Kaplan–MeierTable 2The performance of models in predicting OS in UTUC patientsCohortsModelsC-index1-year AUC3-year AUC5-year AUCTraining setClinical model0.648 (0.596–0.699)0.518 (0.392–0.643)0.699 (0.612–0.786)0.743 (0.641–0.844)Radiomics model0.664 (0.614–0.715)0.542 (0.406–0.679)0.750 (0.667–0.832)0.798 (0.693–0.904)DL model0.681 (0.631–0.731)0.617 (0.493–0.741)0.755 (0.674–0.837)0.792 (0.674–0.909)DL25D model0.705 (0.656–0.754)0.678 (0.562–0.793)0.782 (0.706–0.858)0.843 (0.750–0.937)Combined model0.758 (0.712–0.804)0.688 (0.594–0.783)0.842 (0.776–0.907)0.887 (0.803–0.971)Internal validation setClinical model0.613 (0.533–0.694)0.528 (0.312–0.743)0.645 (0.499–0.792)0.689 (0.504–0.873)Radiomics model0.656 (0.578–0.734)0.726 (0.541–0.911)0.732 (0.598–0.866)0.690 (0.544–0.837)DL model0.656 (0.578–0.734)0.764 (0.581–0.946)0.623 (0.475–0.771)0.595 (0.418–0.773)DL25D model0.668 (0.591–0.746)0.809 (0.705–0.913)0.738 (0.603–0.873)0.756 (0.618–0.894)Combined model0.725 (0.651–0.798)0.830 (0.738–0.921)0.755 (0.627–0.883)0.782 (0.647–0.917)External validation setClinical model0.604 (0.525–0.682)0.664 (0.400–0.927)0.703 (0.578–0.829)0.678 (0.506–0.850)Radiomics model0.657 (0.581–0.733)0.785 (0.591–0.979)0.731 (0.613–0.849)0.711 (0.497–0.925)DL model0.651 (0.575–0.727)0.863 (0.663–1.000)0.630 (0.499–0.761)0.597 (0.387–0.806)DL25D model0.658 (0.582–0.734)0.704 (0.385–1.000)0.664 (0.535–0.793)0.702 (0.512–0.891)Combined model0.704 (0.631–0.777)0.818 (0.620–1.000)0.766 (0.653–0.880)0.760 (0.575–0.945)

During the construction of the radiomics model, we adopted a standardized workflow that included the extraction of features from three sets of dimensions: first-order, shape, and texture features (Figure S2). Following feature selection using LASSO-Cox regression, the KM survival curves of the final model demonstrated the predictive performance, with C-indexes of 0.664 (95% CI: 0.614–0.715), 0.656 (95% CI: 0.578–0.734), and 0.657 (95% CI: 0.581–0.733) in the three study cohorts, respectively (Fig.3and Table2).Fig. 3Construction and prognostic value of the radiomics model.AA forest plot shows the HR and 95% CI of the selected radiomics features, which were identified using Lasso-Cox regression analysis.B–DKM survival curves for the Radiomics Model are shown in the training set (B), internal validation set (C), and external validation set (D). HR, hazard ratios; CI, confidence intervals; KM, Kaplan–Meier

Optimization and performance of DL models

To identify the most effective DL architecture, we evaluated the prognostic performance of four popular networks on both 2D and 2.5D inputs (Fig.4A). As detailed in Table S2, the 2D ResNet-50 model achieved a C-index of 0.681 in the training cohort, whereas the 2.5D ResNet-50 model achieved a modestly higher C-index of 0.705. Given its enriched spatial feature representation, the ResNet-50 architecture was selected for both 2D and 2.5D feature extraction.Fig. 4Workflow of DL models and Grad-CAM visualization.AThe framework for building the DL Model and DL25D Model is illustrated. Both models were based on a pre-trained ResNet-50 network, with the DL Model using 2D image inputs and the DL25D Model using 2.5D image inputs to capture multi-plane information. The models were fine-tuned to extract prognostic features (DL-score) for risk stratification.BRepresentative images and Grad-CAM heatmaps are shown for both 2D and 2.5D models in high-risk and low-risk patients. DL, deep learning; Grad-CAM, gradient-weighted class activation mapping

The prognostic performance of the models based on these extracted features was then evaluated. As shown in Table2, the DL model (based on 2D features) demonstrated C-indexes of 0.681 (95% CI: 0.631–0.731), 0.656 (95% CI: 0.578–0.734), and 0.651 (95% CI: 0.575–0.727) in the training, internal, and external validation sets, respectively (Fig. S3A). The DL25D model (based on 2.5D features) demonstrated C-indexes of 0.705 (95% CI: 0.656–0.754), 0.668 (95% CI: 0.591–0.746), and 0.658 (95% CI: 0.582–0.734) in the three cohorts, respectively (Fig. S3B). The decision-making process of both DL models was further visualized using Grad-CAM, which highlighted the tumor regions most crucial for their predictions (Fig.4B).

Construction and performance of the combined model

The prognostic scores from the clinical, radiomics, DL, and DL25D models were integrated to construct a comprehensive combined model, which served as a final prognostic system presented as a nomogram (Fig.5A). This combined model consistently exhibited higher discrimination compared to all single-modal models across all cohorts. As presented in Table2, the combined model achieved the highest C-index values of 0.758 (95% CI: 0.712–0.804), 0.725 (95% CI: 0.651–0.798), and 0.704 (95% CI: 0.631–0.777) in the training set, internal validation set, and external validation set, respectively. The robust prognostic capacity of the model was further corroborated by KM survival analysis, which disclosed significant disparities in OS between the high-risk group and the low-risk group across all three cohorts (Fig.5B–D).Fig. 5Development and prognostic validation of the combined model.AA nomogram for the combined model is presented, which integrates clinical (e.g., Size, pT_binary), radiomics, and DL features to predict 1-, 3-, and 5-year OS.B–DKM survival curves for the combined model are shown in the training set (B), internal validation set (C), and external validation set (D). DL, deep learning; KM, Kaplan–Meier

To further evaluate the prognostic robustness of the combined model across different tumor stages, a stratified analysis was performed based on the pathological T (pT) stage. Patients in each sub-cohort were stratified into high- and low-risk groups using the global median risk score derived from the training cohort. As shown in Figure S6, the combined model successfully discriminated survival outcomes in both the non-muscle invasive (< T2) and muscle-invasive (≥ T2) subgroups within the training and internal validation cohorts (all log-rankp< 0.05). In the independent external validation cohort, the model maintained significant risk stratification for the muscle-invasive (≥ T2) patients (log-rankp= 0.011). For the external early-stage (< T2) patients, although the survival difference was marginally significant (p= 0.066), likely due to the limited sample size and low event rate, a consistent trend of risk separation was observed. These results suggest that the combined model provides consistent prognostic information independent of the baseline pT stage.

Comprehensive performance evaluation

Lastly, ROC curve analysis, calibration analysis, and DCA were employed to comprehensively assess the performance of all models at 1, 3, and 5 years. As shown in Figs.6, S4, and S5, the combined model exhibited the highest AUC values, suggesting superior discriminatory ability. For example, in the external validation set, the 3-year AUC value of the combined model attained 0.766 (95% CI: 0.653–0.880), numerically surpassing that of the clinical model (0.703), radiomics model (0.731), DL model (0.630), and DL25D model (0.664). The calibration curve of the combined model indicated a strong association between the predicted survival probability and the actual survival probability (e.g., Brier score = 0.187 at 3—year in the training set). The findings of the DCA revealed that the combined model consistently offered the highest net benefits compared to the single-modality models across a wide spectrum of probability thresholds, highlighting its clinical utility for adjuvant decision-making.Fig. 6Model performance evaluation in the training set. This figure shows the performance of the Clinical, Radiomics, DL, DL25D, and Combined models within the training cohort. The performance of each model was evaluated at 1-, 3-, and 5-year follow-up intervals.AROC curves are presented for all models. The AUC and 95% CI for each model are listed, demonstrating their discriminative ability.BCalibration curves are shown to assess the agreement between the predicted and actual survival probabilities. The dashed line represents a perfectly calibrated model.CDCA is presented to evaluate the clinical utility of each model. ROC, receiver operating characteristic; AUC, area under the ROC curve; CI, confidence interval; DCA, decision curve analysis

Discussion

In this study, we developed and validated a combined prognostic model for patients with UTUC by integrating clinical, radiomics, and DL features. Our findings demonstrate that this combined model holds a substantial advantage over single-modal models, exhibiting superior discrimination ability, calibration performance, and clinical utility across the training, internal validation, and independent external validation sets. This prediction system provides a robust tool for personalized perioperative risk stratification immediately after RNU. By offering highly reliable 3-year intermediate-term survival predictions, it effectively guides adjuvant management strategies for UTUC patients.

Our findings provide compelling evidence that integrating multiple data modalities is an effective strategy for enhancing prognostic prediction. As shown in Table2, the combined model achieved the highest C-index across all cohorts, indicating its robust performance. This conclusion is consistent with a growing consensus within the oncology community that multimodal models, which incorporate both clinical and imaging data, are more effective in predicting patient prognosis [28,29]. This enhanced performance is underpinned by the complementary value between different data types. Clinical models provide macroscopic insights into patient status and tumor characteristics, such as pT stage and tumor size, which are established prognostic indicators [1,30]. In contrast, radiomics features offer a quantitative lens to uncover tumor heterogeneity and provide profound insights into tumor biology and the microenvironment, which are difficult to discern through visual inspection alone [11,13]. Previous studies have confirmed that specific radiomics features can serve as noninvasive biomarkers to predict gene mutations and treatment response [31,32]. Additionally, DL features can automatically capture complex and subtle patterns in raw image data, including spatial associations—details that effectively complement standard radiomics profiling [33,34]. By treating this integration as a cohesive analytical framework, our comprehensive combined model constructs a more complete representation of tumor characteristics.

While the overarching goal of this study is multimodal fusion, our core technical advancement lies in employing a 2.5D DL strategy to maximize the quality of features extracted from the imaging modality. Given the limited sample sizes typical in clinical research, full 3D DL models are prone to overfitting and require a substantial number of training samples for robust generalization. Our 2.5D model offers a practical middle ground, synthesizing maximum axial, sagittal, and coronal information to provide vital multiplanar context while reducing computational costs. As demonstrated in Table2, the C-index of our 2.5D ResNet-50 model demonstrated a modest but consistent incremental improvement over the 2D model across all three cohorts. This synthesis allows the model to characterize tumor morphology, boundaries, and internal structures more precisely. Previous research has also shown the benefits of employing 2.5D or 3D convolutional neural networks for tasks like lung nodule classification and brain tumor segmentation, confirming their ability to enhance performance by leveraging volumetric data [3537]. Moreover, model interpretability is crucial for clinical applications. The visualization results from Grad-CAM underscore that tumor margins and internal heterogeneity are key determinants of predictive risk, aligning well with established pathological prognostic factors [25,38].

The most significant contribution of our study is the rigorous validation of our model across multi-institutional cohorts and varying disease severities. As shown in Table1, significant differences in pT stage were observed across the cohorts—a common challenge in multi-institutional research [39,40]. To explicitly address this imbalance and justify the incremental value of our imaging features beyond basic clinical staging, we conducted a rigorous subgroup analysis based on pT stage. Using a strict global median risk cutoff derived from the training cohort, the combined model successfully stratified patients into distinct prognostic groups within both the non-muscle invasive (< T2) and muscle-invasive (≥ T2) sub-cohorts (Fig. S6). Notably, in the unseen external validation cohort, the model maintained its robust stratification ability for muscle-invasive patients. For the external early-stage (< T2) subgroup, although the survival difference was marginally significant (p= 0.066) due to the limited sample size and the inherently low event rate, a strong clinical trend of risk separation was still evident. DCA depicted in Figs.6, S4, and S5further emphasize the clinical utility of our model. In the context of UTUC perioperative management, adopting the model’s recommendation (“treat”) implies implementing more intensive postoperative surveillance or considering adjuvant therapies for high-risk patients, whereas “treat none” implies standard routine follow-up. The DCA data indicate that our combined model offers a higher net benefit than default strategies across a broad spectrum of probability thresholds, aiding physicians in developing tailored follow-up plans.

Despite these promising results, our study has several limitations. First, while our study included a large multicenter cohort, all data were retrospectively collected. CT urography protocols (e.g., scanning phases, slice thicknesses) varied across institutions, and although images were resampled, inherent variations and potential interpolation artifacts may still influence feature extraction. Second, due to the retrospective nature of this real-world multicenter study, detailed postoperative treatment information (e.g., adjuvant chemotherapy) and clinical N stage were inconsistently documented across the majority of participating institutions and thus could not be incorporated into the current model. Consequently, we cannot completely exclude the possibility that the predicted survival outcomes might be partially confounded by variations in subsequent clinical management, alongside the inherent biological risk of the tumor. Third, our inclusion period extended to December 2024, meaning recently enrolled patients had limited follow-up, which directly impacted the stability of our long-term predictions. While our model demonstrated robust overall predictive ability, its calibration varied across time points. The slight deviation observed in the 1-year calibration curve is primarily attributable to the inherently low short-term postoperative mortality rate of UTUC; this extreme scarcity of positive events causes the model to lean towards predicting excessively high survival probabilities. Conversely, the instability in 5-year predictions is a direct consequence of the heavy right-censoring mentioned above. Because we did not impose a strict minimum follow-up duration in order to maximize the sample size for DL, this extreme tail-end sample attrition at the 5-year mark leads to wider CIs and less reliable long-term predictions. Fourth, regarding the 2.5D strategy, selecting only the maximum-area slices from three orthogonal planes captures global macroscopic morphology but may not fully encapsulate the continuous longitudinal spread of UTUC compared to slab-based (adjacent slices) or full 3D models. Finally, our model was designed to predict OS. Integrating other endpoints, such as cancer-specific survival (CSS) or recurrence-free survival (RFS), could provide a more comprehensive evaluation of patient prognosis in future prospective studies.

Supplementary information

ELECTRONIC SUPPLEMENTARY MATERIAL

References

  1. Masson-LecomteABirtleAPradereBEuropean Association of Urology guidelines on upper urinary tract urothelial carcinoma: summary of the 2025 updateEur Urol202587697716Masson-Lecomte A, Birtle A, Pradere B et al (2025) European Association of Urology guidelines on upper urinary tract urothelial carcinoma: summary of the 2025 update. Eur Urol 87: 697–716 doi.org/10.1016/j.eururo.2025.02.023
  2. AminAKhalatbariFChengLGenetic profiling of upper tract urothelial carcinoma: a necessity for precision medicineExpert Rev Mol Diagn202525695708Amin A, Khalatbari F, Cheng L (2025) Genetic profiling of upper tract urothelial carcinoma: a necessity for precision medicine. Expert Rev Mol Diagn 25: 695–708 doi.org/10.1080/14737159.2025.2549834
  3. QuYYaoZXuNPlasma proteomic profiling discovers molecular features associated with upper tract urothelial carcinomaCell Rep Med20234101166Qu Y, Yao Z, Xu N et al (2023) Plasma proteomic profiling discovers molecular features associated with upper tract urothelial carcinoma. Cell Rep Med 4: 101166 doi.org/10.1016/j.xcrm.2023.101166
  4. FujiiYSatoYSuzukiHMolecular classification and diagnostics of upper urinary tract urothelial carcinomaCancer Cell202139793809. e798Fujii Y, Sato Y, Suzuki H et al (2021) Molecular classification and diagnostics of upper urinary tract urothelial carcinoma. Cancer Cell 39: 793–809. e798 doi.org/10.1016/j.ccell.2021.05.008
  5. AudenetFIsharwalSChaEKClonal relatedness and mutational differences between upper tract and bladder urothelial carcinomaClin Cancer Res201925967976Audenet F, Isharwal S, Cha EK et al (2019) Clonal relatedness and mutational differences between upper tract and bladder urothelial carcinoma. Clin Cancer Res 25: 967–976 doi.org/10.1158/1078-0432.CCR-18-2039
  6. Bitaraf M, Ghafoori Yazdi M, Amini E (2023) Upper tract urothelial carcinoma (UTUC) diagnosis and risk stratification: a comprehensive review. Cancers (Basel) 15: 4987 doi.org/10.3390/cancers15204987
  7. MuNJylhäCAxelssonTSydénFBrehmerMThamEPatient-specific targeted analysis of circulating tumour DNA in plasma is feasible and may be a potential biomarker in UTUCWorld J Urol20234134213427Mu N, Jylhä C, Axelsson T, Sydén F, Brehmer M, Tham E (2023) Patient-specific targeted analysis of circulating tumour DNA in plasma is feasible and may be a potential biomarker in UTUC. World J Urol 41: 3421–3427 doi.org/10.1007/s00345-023-04583-w
  8. Collà RuvoloCNoceraLStolzenbachLFIncidence and survival rates of contemporary patients with invasive upper tract urothelial carcinomaEur Urol Oncol20214792801Collà Ruvolo C, Nocera L, Stolzenbach LF et al (2021) Incidence and survival rates of contemporary patients with invasive upper tract urothelial carcinoma. Eur Urol Oncol 4: 792–801 doi.org/10.1016/j.euo.2020.11.005
  9. MargulisVShariatSFMatinSFOutcomes of radical nephroureterectomy: a series from the upper tract urothelial carcinoma collaborationCancer200911512241233Margulis V, Shariat SF, Matin SF et al (2009) Outcomes of radical nephroureterectomy: a series from the upper tract urothelial carcinoma collaboration. Cancer 115: 1224–1233 doi.org/10.1002/cncr.24135
  10. KeylJKeylPMontavonGDecoding pan-cancer treatment outcomes using multimodal real-world data and explainable artificial intelligenceNat Cancer20256307322Keyl J, Keyl P, Montavon G et al (2025) Decoding pan-cancer treatment outcomes using multimodal real-world data and explainable artificial intelligence. Nat Cancer 6: 307–322 doi.org/10.1038/s43018-024-00891-1
  11. BeraKBramanNGuptaAVelchetiVMadabhushiAPredicting cancer outcomes with radiomics and artificial intelligence in radiologyNat Rev Clin Oncol202219132146Bera K, Braman N, Gupta A, Velcheti V, Madabhushi A (2022) Predicting cancer outcomes with radiomics and artificial intelligence in radiology. Nat Rev Clin Oncol 19: 132–146 doi.org/10.1038/s41571-021-00560-7
  12. ScapicchioCGabelloniMBarucciACioniDSabaLNeriEA deep look into radiomicsRadiol Med202112612961311Scapicchio C, Gabelloni M, Barucci A, Cioni D, Saba L, Neri E (2021) A deep look into radiomics. Radiol Med 126: 1296–1311 doi.org/10.1007/s11547-021-01389-x
  13. VithayathilMKokuDCampaniCMachine learning based radiomic models outperform clinical biomarkers in predicting outcomes after immunotherapy for hepatocellular carcinomaJ Hepatol202583959970Vithayathil M, Koku D, Campani C et al (2025) Machine learning based radiomic models outperform clinical biomarkers in predicting outcomes after immunotherapy for hepatocellular carcinoma. J Hepatol 83: 959–970 doi.org/10.1016/j.jhep.2025.04.017
  14. Long S, Li M, Chen J et al (2024) Spatial patterns and MRI-based radiomic prediction of high peritumoral tertiary lymphoid structure density in hepatocellular carcinoma: a multicenter study. J Immunother Cancer 12: e009879 doi.org/10.1136/jitc-2024-009879
  15. Pallumeera M, Giang JC, Singh R, Pracha NS, Makary MS (2025) Evolving and novel applications of artificial intelligence in cancer imaging. Cancers (Basel) 17: 1510 doi.org/10.3390/cancers17091510
  16. TranKAKondrashovaOBradleyAWilliamsEDPearsonJVWaddellNDeep learning in cancer diagnosis, prognosis and treatment selectionGenome Med202113152Tran KA, Kondrashova O, Bradley A, Williams ED, Pearson JV, Waddell N (2021) Deep learning in cancer diagnosis, prognosis and treatment selection. Genome Med 13: 152 doi.org/10.1186/s13073-021-00968-x
  17. Zhang Z, Luo T, Yan M et al (2025) Voxel-level radiomics and deep learning for predicting pathologic complete response in esophageal squamous cell carcinoma after neoadjuvant immunotherapy and chemotherapy. J Immunother Cancer 13: e011149 doi.org/10.1136/jitc-2024-011149
  18. JiangYZhangZYuanQPredicting peritoneal recurrence and disease-free survival from CT images in gastric cancer with multitask deep learning: a retrospective studyLancet Digit Health20224e340e350Jiang Y, Zhang Z, Yuan Q et al (2022) Predicting peritoneal recurrence and disease-free survival from CT images in gastric cancer with multitask deep learning: a retrospective study. Lancet Digit Health 4: e340–e350 doi.org/10.1016/S2589-7500(22)00040-1
  19. KouJPengJYLvWBA serial MRI-based deep learning model to predict survival in patients with locoregionally advanced nasopharyngeal carcinomaRadiol Artif Intell20257e230544Kou J, Peng JY, Lv WB et al (2025) A serial MRI-based deep learning model to predict survival in patients with locoregionally advanced nasopharyngeal carcinoma. Radiol Artif Intell 7: e230544 doi.org/10.1148/ryai.230544
  20. YuAMengXHouZComparative outcomes and prognosis in patients wit ureteral upper tract urothelial carcinoma undergoing segmental ureterectomy versus radical nephroureterectomy: a multicenter cohort studyInt J Surg202511188188836Yu A, Meng X, Hou Z et al (2025) Comparative outcomes and prognosis in patients wit ureteral upper tract urothelial carcinoma undergoing segmental ureterectomy versus radical nephroureterectomy: a multicenter cohort study. Int J Surg 111: 8818–8836 doi.org/10.1097/JS9.0000000000003292
  21. Alqahtani A, Bhattacharjee S, Almopti A, Li C, Nabi G (2024) Radiomics-based computed tomography urogram approach for the prediction of survival and recurrence in upper urinary tract urothelial carcinoma. Cancers (Basel) 16: 3119 doi.org/10.3390/cancers16183119
  22. ChenZChenYSunYPredicting gastric cancer response to anti-HER2 therapy or anti-HER2 combined immunotherapy based on multi-modal dataSignal Transduct Target Ther20249222Chen Z, Chen Y, Sun Y et al (2024) Predicting gastric cancer response to anti-HER2 therapy or anti-HER2 combined immunotherapy based on multi-modal data. Signal Transduct Target Ther 9: 222 doi.org/10.1038/s41392-024-01932-y
  23. ZhouSSunDMaoWDeep radiomics-based fusion model for prediction of bevacizumab treatment response and outcome in patients with colorectal cancer liver metastases: a multicentre cohort studyEClinicalMedicine202365102271Zhou S, Sun D, Mao W et al (2023) Deep radiomics-based fusion model for prediction of bevacizumab treatment response and outcome in patients with colorectal cancer liver metastases: a multicentre cohort study. EClinicalMedicine 65: 102271 doi.org/10.1016/j.eclinm.2023.102271
  24. MahootihaMTakDYeZMultimodal deep learning improves recurrence risk prediction in pediatric low-grade gliomasNeuro Oncol202527277290Mahootiha M, Tak D, Ye Z et al (2025) Multimodal deep learning improves recurrence risk prediction in pediatric low-grade gliomas. Neuro Oncol 27: 277–290 doi.org/10.1093/neuonc/noae173
  25. KocakBBaesslerBBakasSCheckList for EvaluAtion of Radiomics research (CLEAR): a step-by-step reporting guideline for authors and reviewers endorsed by ESR and EuSoMIIInsights Imaging20231475Kocak B, Baessler B, Bakas S et al (2023) CheckList for EvaluAtion of Radiomics research (CLEAR): a step-by-step reporting guideline for authors and reviewers endorsed by ESR and EuSoMII. Insights Imaging 14: 75 doi.org/10.1186/s13244-023-01415-8
  26. KatzmanJLShahamUCloningerABatesJJiangTKlugerYDeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural networkBMC Med Res Methodol20181824Katzman JL, Shaham U, Cloninger A, Bates J, Jiang T, Kluger Y (2018) DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network. BMC Med Res Methodol 18: 24 doi.org/10.1186/s12874-018-0482-1
  27. Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D (2017) Grad-CAM: visual explanations from deep networks via gradient-based localization. In: 2017 IEEE international conference on computer vision (ICCV). IEEE, pp 618–626
  28. LiZQinYLiaoXComparison of clinical, radiomics, deep learning, and fusion models for predicting early recurrence in locally advanced rectal cancer based on multiparametric MRI: a multicenter studyEur J Radiol2025189112173Li Z, Qin Y, Liao X et al (2025) Comparison of clinical, radiomics, deep learning, and fusion models for predicting early recurrence in locally advanced rectal cancer based on multiparametric MRI: a multicenter study. Eur J Radiol 189: 112173 doi.org/10.1016/j.ejrad.2025.112173
  29. WangWLiangHZhangZComparing three-dimensional and two-dimensional deep-learning, radiomics, and fusion models for predicting occult lymph node metastasis in laryngeal squamous cell carcinoma based on CT imaging: a multicentre, retrospective, diagnostic studyEClinicalMedicine202467102385Wang W, Liang H, Zhang Z et al (2024) Comparing three-dimensional and two-dimensional deep-learning, radiomics, and fusion models for predicting occult lymph node metastasis in laryngeal squamous cell carcinoma based on CT imaging: a multicentre, retrospective, diagnostic study. EClinicalMedicine 67: 102385 doi.org/10.1016/j.eclinm.2023.102385
  30. LeowJJOrsolaAChangSLBellmuntJA contemporary review of management and prognostic factors of upper tract urothelial carcinomaCancer Treat Rev201541310319Leow JJ, Orsola A, Chang SL, Bellmunt J (2015) A contemporary review of management and prognostic factors of upper tract urothelial carcinoma. Cancer Treat Rev 41: 310–319 doi.org/10.1016/j.ctrv.2015.02.006
  31. LiuYKimJBalagurunathanYRadiomic features are associated with EGFR mutation status in lung adenocarcinomasClin Lung Cancer201617441448. e446Liu Y, Kim J, Balagurunathan Y et al (2016) Radiomic features are associated with EGFR mutation status in lung adenocarcinomas. Clin Lung Cancer 17: 441–448. e446 doi.org/10.1016/j.cllc.2016.02.001
  32. Hua Y, Sun Z, Xiao Y et al (2024) Pretreatment CT-based machine learning radiomics model predicts response in unresectable hepatocellular carcinoma treated with lenvatinib plus PD-1 inhibitors and interventional therapy. J Immunother Cancer 12: e008953 doi.org/10.1136/jitc-2024-008953
  33. HuangJGaoJLiJGaoSChengFWangYTransformer-based deep learning model for predicting osteoporosis in patients with cervical cancer undergoing external-beam radiotherapyExpert Syst Appl2025273126716Huang J, Gao J, Li J, Gao S, Cheng F, Wang Y (2025) Transformer-based deep learning model for predicting osteoporosis in patients with cervical cancer undergoing external-beam radiotherapy. Expert Syst Appl 273: 126716 doi.org/10.1016/j.eswa.2025.126716
  34. ChangBGengZMeiJApplication of multimodal deep learning and multi-instance learning fusion techniques in predicting STN-DBS outcomes for Parkinson’s disease patientsNeurotherapeutics202421e00471Chang B, Geng Z, Mei J et al (2024) Application of multimodal deep learning and multi-instance learning fusion techniques in predicting STN-DBS outcomes for Parkinson’s disease patients. Neurotherapeutics 21: e00471 doi.org/10.1016/j.neurot.2024.e00471
  35. MehtaKJainAMangalagiriJMenonSNguyenPChapmanDRLung nodule classification using biomarkers, volumetric radiomics, and 3D CNNsJ Digit Imaging202134647666Mehta K, Jain A, Mangalagiri J, Menon S, Nguyen P, Chapman DR (2021) Lung nodule classification using biomarkers, volumetric radiomics, and 3D CNNs. J Digit Imaging 34: 647–666 doi.org/10.1007/s10278-020-00417-y
  36. van der VoortSRIncekaraFWijnengaMMJCombined molecular subtyping, grading, and segmentation of glioma using multi-task deep learningNeuro Oncol202325279289van der Voort SR, Incekara F, Wijnenga MMJ et al (2023) Combined molecular subtyping, grading, and segmentation of glioma using multi-task deep learning. Neuro Oncol 25: 279–289 doi.org/10.1093/neuonc/noac166
  37. OttesenJAYiDTongE2. 5D and 3D segmentation of brain metastases with deep learning on multinational MRI dataFront Neuroinform2022161056068Ottesen JA, Yi D, Tong E et al (2022) 2. 5D and 3D segmentation of brain metastases with deep learning on multinational MRI data. Front Neuroinform 16: 1056068 doi.org/10.3389/fninf.2022.1056068
  38. WeiZXvYLiuHA CT-based deep learning model predicts overall survival in patients with muscle invasive bladder cancer after radical cystectomy: a multicenter retrospective cohort studyInt J Surg202411029222932Wei Z, Xv Y, Liu H et al (2024) A CT-based deep learning model predicts overall survival in patients with muscle invasive bladder cancer after radical cystectomy: a multicenter retrospective cohort study. Int J Surg 110: 2922–2932 doi.org/10.1097/JS9.0000000000001194
  39. FoerschSLang-SchwarzCEcksteinMpT3 colorectal cancer revisited: a multicentric study on the histological depth of invasion in more than 1000 pT3 carcinomas-proposal for a new pT3a/pT3b subclassificationBr J Cancer202212712701278Foersch S, Lang-Schwarz C, Eckstein M et al (2022) pT3 colorectal cancer revisited: a multicentric study on the histological depth of invasion in more than 1000 pT3 carcinomas-proposal for a new pT3a/pT3b subclassification. Br J Cancer 127: 1270–1278 doi.org/10.1038/s41416-022-01889-1
  40. MolinaTJRodenACSzolkowskaMInternational reproducibility study of thymic epithelial tumors staging: pT stage is an issue. proposals for improvement. A RYTHMIC/ITMIG studyLung Cancer2024189107479Molina TJ, Roden AC, Szolkowska M et al (2024) International reproducibility study of thymic epithelial tumors staging: pT stage is an issue. proposals for improvement. A RYTHMIC/ITMIG study. Lung Cancer 189: 107479 doi.org/10.1016/j.lungcan.2024.107479

Republished from the open web under CC-BY. Authors: Peng X, Li Y, Shi W, Xiao B, Xiao X, Yue X, Xv Q, Jiang Q, He W, Xv Y, Xiao M. Read the original.

0 comments

Sign in to join the discussion