Prognostic models for large cell neuroendocrine lung carcinoma: a machine learning and regression approach
Original Article

Prognostic models for large cell neuroendocrine lung carcinoma: a machine learning and regression approach

Xian Gong1,2,3#, Maojie Pan1,2,3,4#, Yuxing Lin1,2,3,5#, Xiaoxuan Ye6, Jiekun Qian1,2,3, Guoliang Liao1,2,3, Jianting Du1,2,3, Bin Zheng1,2,3, Chun Chen1,2,3, Zhang Yang1,2,3

1Department of Thoracic Surgery, Fujian Medical University Union Hospital, Fuzhou, China; 2Fujian Key Laboratory of Cardiothoracic Surgery (Fujian Medical University), Fuzhou, China; 3Clinical Research Center for Thoracic Tumors of Fujian Province, Fuzhou, China; 4Department of Thoracic Surgery, Linyi People’s Hospital, Linyi, China; 5Health Management Department, Fuzhou University Affiliated Provincial Hospital, Fuzhou, China; 6State Key Laboratory of Cellular Stress Biology, Fujian Provincial Key Laboratory of Innovative Drug Target Research, School of Pharmaceutical Sciences, Xiamen University, Xiamen, China

Contributions: (I) Conception and design: X Gong, M Pan; (II) Administrative support: C Chen, Z Yang; (III) Provision of study materials or patients: X Gong, Y Lin, B Zheng; (IV) Collection and assembly of data: X Ye, J Qian; (V) Data analysis and interpretation: X Gong, G Liao, J Du; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

#These authors contributed equally to this work as co-first authors.

Correspondence to: Chun Chen, MD; Zhang Yang, MD, PhD. Department of Thoracic Surgery, Fujian Medical University Union Hospital, 29 Xinquan Road, Fuzhou 350001, China; Fujian Key Laboratory of Cardiothoracic Surgery (Fujian Medical University), 29 Xinquan Road, Fuzhou 350001, China; Clinical Research Center for Thoracic Tumors of Fujian Province, 29 Xinquan Road, Fuzhou 350001, China. Email: chenchun0209@fjmu.edu.cn; zhangyang@fjmu.edu.cn.

Background: Large cell neuroendocrine lung carcinoma (LCNEC) is a rare and aggressive subtype of lung cancer with high rates of lymph node metastasis (60–80%) and distant metastasis (40%) at diagnosis. This study aimed to develop and evaluate 5-year survival prognostic models for patients with LCNEC, comparing the traditional Cox proportional hazards regression model with machine learning approaches, including Gradient Boosting, XGboost, Random Survival Forests, Extra Survival Trees, and Neural Networks.

Methods: This retrospective cohort study utilized data from the Surveillance, Epidemiology, and End Results (SEER) database (2000–2021), including 6,062 patients with pathologically confirmed LCNEC. The primary outcome was the 5-year survival probability. The study employed regression and machine learning approaches, with data that was stratified into training and testing sets based on the year of diagnosis, and four stratification variables were analyzed. Internal-external cross-validation assessed the model performance, while decision curve analysis (DCA) evaluated clinical utility.

Results: The Gradient Boosting model showed better discrimination than all others, achieving the best pooled metrics. Harrell’s C-index of 0.799, Brier score of 0.047, Calibration slope of 1.126 and Calibration-in-the-large of 0.155. Our SHAP value analysis identified chemotherapy as one of the most influential predictors of survival outcomes in LCNEC patients, highlighting its potential clinical importance in guiding treatment strategies for this population. DCA confirmed its superior clinical utility.

Conclusions: Gradient Boosting exhibited excellent predictive accuracy and clinical utility, demonstrating its potential for prognostic evaluation for LCNEC patients.

Keywords: Lung cancer; carcinoma; neuroendocrine; machine learning; survival analysis


Submitted Feb 05, 2025. Accepted for publication May 09, 2025. Published online Jul 28, 2025.

doi: 10.21037/tlcr-2025-130


Highlight box

Key findings

• Survival analysis can be utilized for constructing predictive models with machine learning, demonstrating predictive efficacy comparable to that of Cox proportional hazards regression.

• Chemotherapy was identified as the most critical therapeutic factor influencing survival probability in large cell neuroendocrine lung carcinoma (LCNEC) patients, outweighing both surgical intervention and radiotherapy.

What is known and what is new?

• LCNEC represents a rare subtype of pulmonary neuroendocrine tumors, demonstrating significantly poorer prognosis compared to the other subtypes of non-small cell lung cancer.

• Our study demonstrates that the gradient boosting algorithm exhibits robust discriminative power and clinical utility, representing the most effective machine learning approach for survival analysis in LCNEC.

What is the implication, and what should change now?

• The gradient boosting-based predictive model shows promising potential to facilitate clinical prognosis prediction in LCNEC. Particular attention should be directed toward the impact of chemotherapy in LCNEC management.


Introduction

Large cell neuroendocrine lung carcinoma (LCNEC) is a rare and aggressive subtype of pulmonary neuroendocrine tumors, representing a distinct category of high-grade malignancies alongside typical carcinoid (TC), atypical carcinoid (AC), and small cell lung cancer (SCLC) (1). First described in 1991, LCNEC is defined by morphological and immunohistochemical evidence of neuroendocrine differentiation (2). Despite accounting for only 2.1–3.5% of resected lung cancer cases, LCNEC is characterized by poor prognosis, with lymph node metastasis rates of 60% to 80% and distant metastasis of 40%, mirroring the clinical trajectory of SCLC (3,4). A population-based study in the United States reported an age-adjusted incidence rate of 0.3 per 100,000 for LCNEC, with a gradual increase observed between 2000 and 2013 (5). However, current therapeutic strategies remain controversial, and outcomes are particularly poor in patients with metastatic or chemotherapy-resistant disease (6,7). These challenges highlight the urgent need for tailored prognostic tools to improve risk stratification and support personalized treatment decisions. Prognostic modeling is crucial for optimizing clinical decision-making in the context of LCNEC, given its aggressive behavior and poor outcomes (8,9), most existing prognostic models for lung cancer predominantly focus on non-small cell lung cancer (NSCLC) such as adenocarcinoma or squamous cell carcinoma (10-13) tailored to LCNEC remain scarce and often fail to incorporate essential steps like standardization, imputation for missing values, and validation for transferability across diverse subgroups, which are crucial for ensuring fairness and clinical applicability (14-17). Such limitations reduce their effectiveness, especially for underrepresented patient groups.

Machine learning has transformed medical research by enabling the analysis of complex, nonlinear relationships, offering advantages in risk stratification for malignancies (18). Prior studies have demonstrated the utility of machine learning in cancer diagnosis and survival prediction (19-21). However, several challenges remain in applying these methods to survival analysis. For example, many studies overlook critical preprocessing steps such as normalization, scaling, or imputation, leading to potential bias and reduced validity (22,23). Survival predictions are often reduced to binary outcomes, which may not effectively utilize time-to-event data (22,23).

This study addresses these gaps by leveraging the Surveillance, Epidemiology, and End Results (SEER) database to develop a robust, clinically applicable 5-year survival prognostic model for LCNEC. By comparing traditional Cox proportional hazards regression with advanced machine learning methods, including Gradient Boosting, XGBoost, Random Survival Forests, Extra Survival Trees, and Neural Networks, this study aims to identify the most effective approach. Furthermore, we applied rigorous internal-external cross-validation and subgroup analyses to evaluate model transportability and ensure equitable performance across diverse patient populations. These efforts aim to provide a reliable and clinically relevant tool for addressing the challenges associated with LCNEC prognosis. The study adhered to the TRIPOD guidelines, ensuring rigorous methodology and transparent reporting of the machine learning-based prognostic analysis (24). We present this article in accordance with the TRIPOD reporting checklist (available at https://tlcr.amegroups.com/article/view/10.21037/tlcr-2025-130/rc).


Methods

Study approval and design

This retrospective cohort study utilized data from the SEER Program, specifically the SEER*Stat database (incidence—SEER Research Data, 17 registries, Nov 2023 Sub). Data covering diagnoses from 2000 to 2021 were accessed in April 2024, based on the November 2023 submission. As the SEER database contains de-identified information, institutional review board (IRB) approval and patient consent were not required for this study. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.

Sample size calculations

To ensure adequate power, we calculated the required sample size based on 25 potential predictor variables and an estimated annual mortality rate of 60% for LCNEC patients. A minimum of 955 participants, yielding approximately 1,219 events (48.75 events per predictor variable), was necessary to fit the regression models reliably (25) is crucial to highlight that there is no established standard for determining the minimum sample size required for our specific machine learning models. Nonetheless, some studies suggest that certain machine learning techniques might require considerably larger datasets (26,27).

Patient population and outcome

The study cohort comprised patients diagnosed with LCNEC between 2000 and 2021. The outcome for this study was the 5-year survival rate from the date of a diagnosis of LCNEC. We defined the diagnosis of LCNEC as the presence of the histologic type International Classification of Diseases for Oncology, Third Edition (ICD-O-3) code 8013/3 (large cell neuroendocrine carcinoma) combined with a primary site code within the range of 340 to 343 or 349 (340 main bronchus; 341 primary site; 342 middle lobe, bronchus, or lung; 343 lower lobe, bronchus, or lung; and 349 bronchus or lung).

Candidate variables

We initially selected 46 variables from the SEER database, encompassing demographic, clinical, and treatment-related factors. Variables capturing redundant information across different years were consolidated, resulting in a final set of 25 predictors (Table S1). Correlation analysis was used to remove highly correlated variables, minimizing multicollinearity. Correlation analysis in Figure 1 reveals strong correlations between specific variables, such as “Reason no cancer-directed surgery” and “Surg/Rad Seq”. As a result, “Reason no cancer-directed surgery” was excluded from further analysis to avoid redundancy. After a thorough feature selection process, the final dataset included 17 categorical variables and 5 numerical variables, providing a robust foundation for model development (Table S2).

Figure 1 The correlation heatmap for all variables. LN, lymph node; M, metastasis; N, node; T, tumor.

Data preprocessing and missing data

To prepare the data for modeling, we employed label encoding for categorical variables and standardized continuous variables. Missing data were addressed using multiple imputation by chained equations (MICE) (28), generating 20 imputations for all subsequent model fitting and evaluations. The Python package miceforest (version 6.0.3) was used for this process, ensuring robust handling of incomplete data.

Statistical analysis

We evaluated six modeling approaches: Cox proportional hazards regression and five machine learning survival methods (Gradient Boosting, XGBoost, Random Survival Forests, Extra Survival Trees, and Neural Networks). A specialized internal-external cross-validation framework was implemented to assess model performance and transportability across subgroups stratified by race, age, primary site, and stage group (Figure 2). Data were partitioned into training and testing sets based on diagnosis year to mimic real-world scenarios and evaluate temporal generalizability. We stratified the dataset based on four variables: race, age, primary site, and stage group. To address small sample sizes, data for “Asian or Pacific Islander” and “American Indian/Alaska Native” were combined into a single group comprising 161 cases, reducing the risk of suboptimal predictive performance. For the stage group, patients were categorized into four groups: stages 1, 2, 3, and 4. Due to the uneven distribution of age, we divided patients into four age brackets: 0–49, 50–59, 60–69, and 70 years and above. This stratification framework helped balance subgroup sizes, thereby improving the robustness, reliability, and predictive accuracy of our analysis.

Figure 2 Summary of internal-external cross validation framework. LCNEC, large cell neuroendocrine lung carcinoma; SEER, Surveillance, Epidemiology, and End Results.

Hyperparameter tuning for the machine learning models was conducted using Bayesian optimization with the Python package scikit-optimize (version 0.10.2). The Cox model, which does not require internal cross-validation and hyperparameter tuning, was directly fitted using the training dataset (29).

Cox proportional hazards model

A pooled Cox model was initially fitted using all candidate variables for the purpose of variable selection (30). For each imputed dataset, we constructed individual Cox models and pooled the exponentiated coefficients using Rubin’s rules. This approach yielded pooled P values and confidence intervals for each variable. Variables with 95% confidence intervals (CIs) excluding 1 and P values <0.05 were retained for subsequent analyses. Fractional polynomial terms were used to model nonlinear relationships for continuous variables such as age, year of diagnosis, time from diagnosis to treatment, regional nodes examined, and tumor size, with a maximum of three polynomial powers. All Cox models were implemented using the Python package lifelines (version 0.29.0).

Machine learning survival models

Five machine learning survival models were employed in this study, each assuming proportional hazards to estimate survival functions. These models included Gradient Boosting, XGBoost, Random Survival Forests, Extra Survival Trees, and Neural Networks. The machine learning models were developed using the Python packages scikit-survival (version 0.23.0), XGBoost (version 2.1.1), and TFDeepSurv (version 2.1.0). For variable selection, Lasso regression with an L1 regularization coefficient of 0.01 was applied, ensuring that key predictive variables were retained. This step was performed using the scikit-survival package, preparing the data for internal-external cross-validation and enhancing model robustness.

Performance evaluation

Model performance was assessed in three key dimensions: discrimination, calibration, and clinical utility. Discrimination was measured using Harrell’s concordance index (C-index), accounting for censoring. Calibration was evaluated via the Brier score, calibration slope, and calibration-in-the-large. Subgroup-level performance metrics for race, stage group, primary site, and age group were computed during internal-external cross-validation and pooled using random-effects meta-analysis to derive estimates with 95% CIs. Calibration plots were generated to compare observed and predicted survival probabilities across the range of predicted values. Clinical utility was evaluated using decision curve analysis (DCA), comparing the net benefit of the models at various probability thresholds. The baseline strategies were defined as ‘treat-all’ (assuming all patients receive treatment) and ‘treat-none’ (assuming no patients receive treatment). Net benefit was calculated across a range of threshold probabilities (0% to 100%) by varying the threshold along the x-axis, allowing assessment of the model’s clinical utility at different decision points. This comprehensive evaluation framework ensured robust comparisons across models and subgroups (29,31).


Results

Study cohort and incidence rates

The study cohort comprised 6,062 patients diagnosed with LCNEC between 2000 and 2021. Among these, 4,985 deaths were recorded and 1,077 patients were alive at the end of the follow-up period. As shown in Figure 3, the overall 5-year survival rate was 18.64%, with slight differences observed between the training cohort (2000–2010, survival rate: 19.17%) and the testing cohort (2011–2021, survival rate:18.18%), indicating minimal temporal variation in outcomes. At the same time, the differences between the various subgroups of Age, Primary Site, and Stage Group are relatively pronounced.

Figure 3 Kaplan-Meier survival curves for the four subgroups. NOS, not otherwise specified.

Comparison of models

Table 1 compares the predictive performance of gradient boosting, XGBoost, random survival forest, and other models across race, age, primary site, and stage subgroups. The gradient boosting model achieved the best comprehensive performance, with the highest Harrell’s C-index (0.799) and lowest Brier score (0.047). Its calibration slope (1.126) further indicated superior predictive reliability compared to both traditional Cox regression and other machine learning approaches. As shown in Table S3, non-linear fractional polynomial terms were selected for age, year of diagnosis, time from diagnosis to treatment, regional nodes examined, and tumor size for the final Cox models. The optimal hyperparameter for the gradient boosting model was identified using Bayesian optimization. Key parameters included a learning rate of 0.017, a maximum tree depth of 8, 352 weak learners, and subsampling ratio of 0.707. Detailed tuning ranges and optimization methodologies are provided in Table S4. Overall, the Gradient Boosting showed improved performance. Overall, the Gradient Boosting showed improved performance, with pooled estimates from the random-effects meta-analysis consistently outperforming those of the Cox model (Tables S5-S8).

Table 1

Summary performance metrics for all models across subgroups

Stratification variable Performance metric Gradient boosting (95% CI) XGBoost (95% CI) Random survival forest (95% CI) Extra survival trees (95% CI) Neural networks (95% CI) Cox model (95% CI)
Race Harrell’s C index 0.799 (0.780, 0.819) 0.788 (0.762, 0.815) 0.786 (0.772, 0.800) 0.779 (0.776, 0.782) 0.776 (0.743, 0.808) 0.792 (0.788, 0.795)
Brier score 0.047 (0.035, 0.060) 0.049 (0.031, 0.068) 0.048 (0.036, 0.060) 0.049 (0.037, 0.060) 0.052 (0.036, 0.067) 0.053 (0.042, 0.065)
Calibration slope 1.126 (0.938, 1.313) 1.637 (1.505, 1.768) 1.121 (0.815, 1.427) 1.309 (1.019, 1.599) 1.123 (0.828, 1.418) 0.907 (0.816, 0.998)
Calibration-in-the-large 0.155 (0.148, 0.161) 0.185 (0.168, 0.201) 0.133 (0.123, 0.144) 0.129 (0.114, 0.144) 0.148 (0.141, 0.155) 0.122 (0.103, 0.142)
Age (years) Harrell’s C index 0.793 (0.781, 0.806) 0.784 (0.767, 0.801) 0.782 (0.773, 0.791) 0.775 (0.763, 0.787) 0.781 (0.772, 0.789) 0.780 (0.770, 0.790)
Brier score 0.065 (0.054, 0.077) 0.068 (0.051, 0.086) 0.065 (0.056, 0.075) 0.066 (0.056, 0.076) 0.068 (0.056, 0.079) 0.070 (0.062, 0.078)
Calibration slope 1.057 (0.841, 1.272) 1.468 (1.152, 1.784) 1.159 (0.884, 1.433) 1.356 (1.037, 1.676) 1.156 (0.844, 1.469) 0.912 (0.750, 1.074)
Calibration-in-the-large 0.134 (0.104, 0.164) 0.166 (0.132, 0.201) 0.122 (0.080, 0.165) 0.122 (0.073, 0.171) 0.136 (0.092, 0.179) 0.112 (0.070, 0.154)
Primary site Harrell’s C index 0.790 (0.765, 0.816) 0.784 (0.764, 0.803) 0.778 (0.754, 0.802) 0.758 (0.723, 0.793) 0.774 (0.749, 0.800) 0.774 (0.745, 0.803)
Brier score 0.049 (0.032, 0.066) 0.048 (0.031, 0.065) 0.049 (0.033, 0.065) 0.048 (0.032, 0.064) 0.050 (0.034, 0.066) 0.053 (0.035, 0.072)
Calibration slope 1.108 (0.624, 1.592) 1.249 (0.862, 1.637) 1.166 (0.821, 1.511) 1.259 (0.921, 1.597) 0.959 (0.670, 1.247) 0.825 (0.583, 1.067)
Calibration-in-the-large 0.122 (0.088, 0.156) 0.133 (0.084, 0.182) 0.000 (-0.251, 0.251) 0.099 (0.055, 0.143) 0.112 (0.073, 0.151) 0.096 (0.070, 0.122)
Stage Harrell’s C index 0.666 (0.624, 0.708) 0.650 (0.610, 0.690) 0.654 (0.626, 0.682) 0.639 (0.610, 0.668) 0.643 (0.604, 0.682) 0.646 (0.609, 0.684)
Brier score 0.093 (0.039, 0.147) 0.100 (0.046, 0.155) 0.094 (0.045, 0.144) 0.097 (0.045, 0.148) 0.095 (0.042, 0.149) 0.104 (0.049, 0.158)
Calibration slope 0.751 (0.586, 0.917) 0.899 (0.556, 1.241) 0.816 (0.596, 1.036) 1.251 (0.871, 1.632) 0.890 (0.541, 1.238) 0.592 (0.517, 0.668)
Calibration-in-the-large 0.192 (0.084, 0.300) 0.214 (0.049, 0.380) 0.156 (0.038, 0.274) 0.168 (0.031, 0.306) 0.177 (0.057, 0.297) 0.107 (0.028, 0.185)

CI, confidence interval.

Figure 4 illustrates that the calibration curves for Gradient Boosting align well with observed survival probabilities across the entire risk spectrum, though minor over-predictions were noted. Other models exhibited varying degrees of calibration instability (Figures S1-S5). Across subgroups, the Cox model demonstrated relatively stable Harrell’s C-index values (e.g., 0.792 for race, 0.780 for age, 0.774 for primary site, and 0.646 for stage), generally matching or outperforming other machine learning models (XGBoost, Random Survival Forest, Extra Survival Trees, Neural Networks) in most subgroups. However, its calibration slopes often deviated from the ideal value of 1 (e.g., 0.907 for race, 0.912 for age), suggesting suboptimal calibration. Among the machine learning models excluding Gradient Boosting, XGBoost achieved competitive C-index values (e.g., 0.788 for race, 0.784 for age) but exhibited significant calibration instability, with slopes exceeding 1.5 in some subgroups (e.g., 1.637 for race). Similarly, Random Survival Forest and Extra Survival Trees showed moderate C-index performance (e.g., 0.786–0.797 for race and primary site) but inconsistent calibration slopes (e.g., 1.121–1.599), indicating potential overfitting or model misspecification. Neural Networks demonstrated the lowest C-index values (e.g., 0.776–0.781) and variable calibration slopes (0.959–1.469), reflecting limited generalizability. Overall, these findings highlight a trade-off between discrimination (C-index) and calibration across models, with the Cox model offering more stable discrimination but poorer calibration, while machine learning models exhibited greater variability in both metrics. DCA further confirmed that the clinical benefits of Gradient Boosting were the most stable. The XGBoost model demonstrated significant losses in benefit in certain subgroups, and similar issues were observed with the Cox model and the Extra Survival Trees model in some subgroups (Figures S6,S7).

Figure 4 The calibration curves of Gradient Boosting in different subgroups. The correlation between predicted and actual risks across all models is represented using histograms to illustrate the distribution of predicted risks for each model.

Comparison of the stratification variables

In the internal-external cross-validation framework, data from Period 1 (2000–2010), excluding a specific category of a stratification variable, were used for model fitting, while data from Period 2 (2011–2021) for that category were used for validation. If the pooled estimate from the random-effects meta-analysis yielded favorable results, it indicated strong transportability of the corresponding stratification variable. As shown in Figure 5, race demonstrated the highest transportability among the four stratification variables. This was reflected by its pooled Harrell’s C-index for the Gradient Boosting model, which reached 0.799 (95% CI: 0.780–0.819), along with superior calibration metrics. Both age group and primary site also showed strong transportability, with pooled Harrell’s C-index values of 0.793 (95% CI: 0.781–0.806) and 0.790 (95% CI: 0.765–0.818), respectively, although their calibration metrics were slightly less favorable compared to race. In contrast, the stage group exhibited the lowest transportability, with a pooled Harrell’s C-index of 0.666 (95% CI: 0.624–0.708) for the Gradient Boosting model, suggesting reduced predictive consistency in this subgroup.

Figure 5 Results from the internal-external cross-validation of the Gradient Boosting model across different subgroups. The performance metric estimates and 95% CIs for each subgroup, along with the overall summary obtained using random-effects meta-analysis and their 95% CI. CI, confidence interval; E/O, events divided by overall cohort size (E/O), indicating the proportion of the cohort with the outcome event; NOS, not otherwise specified.

SHapley Additive exPlanations (SHAP) value analysis

In this study, we utilized SHAP values to enhance the interpretability of the machine learning model. As shown in Figure 6, chemotherapy appears to emerge as a key predictor associated with 5-year survival outcomes in LCNEC patients, though further validation is required to confirm its relative contribution compared to other clinical variables. Currently, the treatment strategies for LCNEC are not well defined (32). The results of the SHAP value analysis indicate that the impact of chemotherapy on patient survival may exceed that of surgical treatment and radiotherapy.

Figure 6 The SHAP summary plot for the LCNEC machine learning prognostic prediction model. The SHAP values of the final prediction model rank the features according to their contributions to the survival outcome, from highest to lowest. LCNEC, large cell neuroendocrine lung carcinoma; LN, lymph node; M, metastasis; N, node; SHAP, SHapley Additive exPlanation; T, tumor.

Clinical implementation context

The model offers valuable support for multidisciplinary team (MDT) discussions during the post-diagnostic phase, before therapeutic decisions are made. To facilitate clinical use, the computational framework will be made publicly available through a GitHub repository (https://github.com/grox0914/LCNEC-GraBoost-Prognosticator), allowing clinicians to input patient data using standardized protocols. The system will generate individualized 5-year survival probability estimates to assist in developing evidence-based treatment strategies. These prognostic outputs are intended to complement existing clinical guidelines and support personalized decision-making for LCNEC patients.


Discussion

This study is the first to apply an internal-external cross-validation framework for prognostic modeling in LCNEC. By stratifying the data based on race, age, primary site, and stage group, we assessed the transportability and robustness of survival models, addressing critical gaps in existing prognostic tools (33,34). Our findings demonstrated that the Gradient Boosting model showed better discrimination than traditional Cox regression and other machine learning models in terms of discrimination, calibration, and clinical utility. In contrast to conventional approaches that rely on static, dichotomous classification strategies—which often fail to capture the temporal dynamics of disease progression—our methodology represents a paradigm shift. By systematically integrating time-series characteristics from right-censored data into machine learning algorithms, we transition from categorical predictions to time-varying risk probability modeling. This advancement addresses a critical limitation of previous models—the underutilization of temporal prognostic information—and establishes a novel framework for dynamic risk stratification. The Gradient Boosting model demonstrated comparatively higher performance across multiple metrics, including Harrell’s C-index and Brier score, while showing strong alignment in calibration curves. Its clinical utility, evaluated through DCA, further confirmed its potential for real-world application. Importantly, the internal-external cross-validation framework validated the model’s ability to perform consistently across diverse patient subgroups, with good results in racial, age, and primary site subgroups.

An important strength of our study is the evaluation of model fairness and transportability across subgroups, particularly racial and age categories. The Gradient Boosting model achieved equitable performance across these subgroups, minimizing predictive disparities often observed in clinical prediction tools (35-37). Such fairness is critical for ensuring that prognostic models benefit all patient populations and do not exacerbate existing healthcare inequalities.

Machine learning methods, such as Gradient Boosting, offer notable advantages over traditional statistical approaches like the Cox proportional hazards model, particularly in capturing complex, nonlinear relationships between variables (18,38). While previous prognostic models for lung cancer primarily focused on NSCLC subtypes (10-12), this study highlights the feasibility and utility of machine learning for the underrepresented and aggressive LCNEC subtype. During the feature selection phase, we employed two distinct methodologies: (I) stepwise regression with P value thresholding (for the Cox proportional hazards model); and (II) LASSO regression (within the machine learning framework). Due to the limited dimensionality of the original feature set, most variables were retained after screening. Although preliminary evaluations suggested minimal impact on predictive performance, methodological differences between these selection approaches—particularly inherent biases—may introduce confounding effects when comparing models. To address this analytical variability, our ongoing work focuses on developing a standardized feature selection protocol that harmonizes statistical and machine learning methodologies.

While recent studies have explored prognostic models for LCNEC, methodological limitations raise concerns about their generalizability. For instance, Fisch et al. (39) conducted a retrospective analysis of 191 metastatic LCNEC patients to assess treatment-outcome relationships; however, their reliance on conventional Kaplan-Meier survival analysis and Cox regression—without applying formal model validation frameworks such as discrimination or calibration assessments—compromises methodological robustness. Moreover, the lack of stratified analyses across clinically relevant subgroups (e.g., treatment modalities or metastatic patterns) introduces potential selection bias. Similarly, Xia et al. (40) developed a nomogram for advanced LCNEC using Cox regression-based variable selection, but its clinical utility is limited by modest predictive performance (training cohort C-index =0.681, 95% CI: 0.656–0.706; validation cohort C-index =0.663, 95% CI: 0.628–0.698). Xu et al. (22) used Gradient Boosted Decision Trees (GBDT) to predict 2-year survival in 2,897 LCNEC cases, identifying surgery and chemotherapy as key prognostic factors [area under the curve (AUC) =0.831]. However, their binary endpoint design (alive/deceased at 2 years) failed to account for the temporal dynamics of survival data, limiting the model’s utility in informing time-sensitive clinical decisions.

Unlike prior studies that often neglected key steps in data preprocessing, such as multiple imputation for missing data and standardization (41), our study incorporated rigorous preprocessing methods to ensure fairness and reduce bias. This comprehensive approach, coupled with internal-external cross-validation, enhances the reliability and generalizability of our findings. While traditional regression models remain effective for low-dimensional clinical datasets (42), the enhanced clinical utility of Gradient Boosting, as demonstrated in DCA, underscores its relevance in improving prognostic predictions for LCNEC patients.

The SHAP value analysis in this study highlights chemotherapy as a significant predictor of 5-year survival for LCNEC patients. However, these findings should be interpreted with caution, given their reliance on sensitivity analyses and the incomplete adjustment for potential confounding factors. Currently, no standardized treatment protocols exist for LCNEC, and therapeutic strategies are largely extrapolated from those used for SCLC and NSCLC. A previous study has shown that cisplatin-etoposide regimens result in slightly inferior median overall survival compared to SCLC patients receiving the same treatment (43), whereas other research suggests that NSCLC-based protocols may achieve better response rates than SCLC-oriented approaches (44). Although immunotherapy has significantly improved survival outcomes across several lung cancer subtypes, its efficacy in LCNEC remains inconclusive (32). Most available evidence comes from small-sample, retrospective studies, underscoring the need for future multicenter randomized controlled trials to comprehensively evaluate the effectiveness of immunotherapy relative to chemotherapy, surgery, and radiotherapy, as well as to explore biomarker-driven personalized treatment strategies. Additionally, while tumor laterality was identified as a secondary prognostic factor in our analysis, its potential influence on survival outcomes in LCNEC warrants further investigation. Prior research (45) has suggested that left-sided tumors may be associated with poorer outcomes. However, due to limitations in sample size and incomplete documentation in our dataset, we were unable to definitively assess the prognostic impact of laterality. Thus, rigorously designed prospective studies are necessary to validate the clinical significance of tumor laterality in LCNEC prognosis.

The Gradient Boosting model provides a reliable and clinically relevant tool for predicting 5-year survival for LCNEC patients. By incorporating routinely collected variables such as age, tumor size, stage, and treatment modalities, the model can be easily integrated into clinical workflows to support decision-making in MDT discussions. Its ability to outperform traditional methods highlights the potential of machine learning in enhancing patient-specific prognostication and treatment planning.

Limitations

Despite its strengths, there are certain limitations in this study. First, external validation using independent datasets was not conducted, which may affect the generalizability of our findings. However, robust internal-external validation mitigates this concern, as recent guidelines emphasize its importance during model development (17,46). Second, the SEER database lacks certain clinical variables, such as smoking status, surgical techniques, chemotherapy regimens, and geographic information, which could influence survival predictions. Future studies incorporating more comprehensive datasets with additional clinical and biological predictors are needed to further refine and validate the models.


Conclusions

This study demonstrates the feasibility, robustness, and clinical relevance of machine learning models, particularly Gradient Boosting, for survival prediction for LCNEC patients. Through rigorous internal-external cross-validation, the model exhibited superior discrimination, calibration, and clinical utility while maintaining fairness across subgroups. These findings represent a significant advancement in LCNEC prognostication. Future research should focus on external validation and the incorporation of additional clinical and molecular variables to further improve model accuracy and generalizability.


Acknowledgments

None.


Footnote

Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://tlcr.amegroups.com/article/view/10.21037/tlcr-2025-130/rc

Peer Review File: Available at https://tlcr.amegroups.com/article/view/10.21037/tlcr-2025-130/prf

Funding: This study was partially sponsored by Fujian Provincial Health Technology Project (2024QNA022, 2022GGA021), National Natural Science Foundation of China (82203307), the Fujian Provincial Natural Science Foundation of China (2022J01241), the Talent Fund Project of Fujian Medical University Union Hospital (2021XH029), and Joint Fund for the innovation of science and Technology, Fujian province (2023Y9204).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://tlcr.amegroups.com/article/view/10.21037/tlcr-2025-130/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Fasano M, Della Corte CM, Papaccio F, et al. Pulmonary Large-Cell Neuroendocrine Carcinoma: From Epidemiology to Therapy. J Thorac Oncol 2015;10:1133-41. [Crossref] [PubMed]
  2. Travis WD, Linnoila RI, Tsokos MG, et al. Neuroendocrine tumors of the lung with proposed criteria for large-cell neuroendocrine carcinoma. An ultrastructural, immunohistochemical, and flow cytometric study of 35 cases. Am J Surg Pathol 1991;15:529-53. [Crossref] [PubMed]
  3. Sánchez de Cos Escuín J. Diagnosis and treatment of neuroendocrine lung tumors. Arch Bronconeumol 2014;50:392-6. [Crossref] [PubMed]
  4. Lindsay CR, Shaw EC, Moore DA, et al. Large cell neuroendocrine lung carcinoma: consensus statement from The British Thoracic Oncology Group and the Association of Pulmonary Pathologists. Br J Cancer 2021;125:1210-6. [Crossref] [PubMed]
  5. Deng C, Wu SG, Tian Y. Lung Large Cell Neuroendocrine Carcinoma: An Analysis of Patients from the Surveillance, Epidemiology, and End-Results (SEER) Database. Med Sci Monit 2019;25:3636-46. [Crossref] [PubMed]
  6. Derks JL, van Suylen RJ, Thunnissen E, et al. Chemotherapy for pulmonary large cell neuroendocrine carcinomas: does the regimen matter? Eur Respir J 2017;49:1601838. [Crossref] [PubMed]
  7. Galvani L, Zappi A, Pusceddu S, et al. Capecitabine and temozolomide or temozolomide alone in patients with atypical carcinoids. Endocrine 2025;88:660-7. [Crossref] [PubMed]
  8. Moons KG, Royston P, Vergouwe Y, et al. Prognosis and prognostic research: what, why, and how? BMJ 2009;338:b375. [Crossref] [PubMed]
  9. de Hond AAH, Leeuwenberg AM, Hooft L, et al. Guidelines and quality criteria for artificial intelligence-based prediction models in healthcare: a scoping review. NPJ Digit Med 2022;5:2. [Crossref] [PubMed]
  10. She Y, Jin Z, Wu J, et al. Development and Validation of a Deep Learning Model for Non-Small Cell Lung Cancer Survival. JAMA Netw Open 2020;3:e205842. [Crossref] [PubMed]
  11. Tong H, Sun J, Fang J, et al. A Machine Learning Model Based on PET/CT Radiomics and Clinical Characteristics Predicts Tumor Immune Profiles in Non-Small Cell Lung Cancer: A Retrospective Multicohort Study. Front Immunol 2022;13:859323. [Crossref] [PubMed]
  12. Guo Y, Li L, Zheng K, et al. Development and validation of a survival prediction model for patients with advanced non-small cell lung cancer based on LASSO regression. Front Immunol 2024;15:1431150. [Crossref] [PubMed]
  13. Zhao J, Wang L, Zhou A, et al. Decision model for durable clinical benefit from front- or late-line immunotherapy alone or with chemotherapy in non-small cell lung cancer. Med 2024;5:981-997.e4.
  14. Xia K, Chen D, Jin S, et al. Prediction of lung papillary adenocarcinoma-specific survival using ensemble machine learning models. Sci Rep 2023;13:14827. [Crossref] [PubMed]
  15. Lu YF, Chang YH, Chen YJ, et al. Proteomic profiling of tumor microenvironment and prognosis risk prediction in stage I lung adenocarcinoma. Lung Cancer 2024;191:107791. [Crossref] [PubMed]
  16. Fang X, Liu X, Weng C, et al. Construction and Validation of a Protein Prognostic Model for Lung Squamous Cell Carcinoma. Int J Med Sci 2020;17:2718-27. [Crossref] [PubMed]
  17. Collins GS, Dhiman P, Ma J, et al. Evaluation of clinical prediction models (part 1): from development to external validation. BMJ 2024;384:e074819. [Crossref] [PubMed]
  18. Nguyen L, Van Hoeck A, Cuppen E. Machine learning-based tissue of origin classification for cancer of unknown primary diagnostics using genome-wide mutation features. Nat Commun 2022;13:4013. [Crossref] [PubMed]
  19. Guo Z, Zhang Z, Liu L, et al. Machine learning for predicting liver and/or lung metastasis in colorectal cancer: A retrospective study based on the SEER database. Eur J Surg Oncol 2024;50:108362. [Crossref] [PubMed]
  20. Alabi RO, Elmusrati M, Leivo I, et al. Machine learning explainability in nasopharyngeal cancer survival using LIME and SHAP. Sci Rep 2023;13:8984. [Crossref] [PubMed]
  21. Lynch CM, Abdollahi B, Fuqua JD, et al. Prediction of lung cancer patient survival via supervised machine learning classification techniques. Int J Med Inform 2017;108:1-8. [Crossref] [PubMed]
  22. Xu X, Liu B, Su Y, et al. The prognostic analysis and a machine-learning based disease-specific survival state model in pulmonary large-cell neuroendocrine carcinomas. J Thorac Dis 2024;16:5152-66. [Crossref] [PubMed]
  23. Liang M, Singh S, Huang J. Implementing machine learning to predict survival outcomes in patients with resected pulmonary large cell neuroendocrine carcinoma. Expert Rev Anticancer Ther 2024;24:1041-53. [Crossref] [PubMed]
  24. Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024;385:e078378. [Crossref] [PubMed]
  25. Riley RD, Ensor J, Snell KIE, et al. Calculating the sample size required for developing a clinical prediction model. BMJ 2020;368:m441. [Crossref] [PubMed]
  26. Infante G, Miceli R, Ambrogi F. Sample size and predictive performance of machine learning methods with survival data: A simulation study. Stat Med 2023;42:5657-75. [Crossref] [PubMed]
  27. van der Ploeg T, Austin PC, Steyerberg EW. Modern modelling techniques are data hungry: a simulation study for predicting dichotomous endpoints. BMC Med Res Methodol 2014;14:137. [Crossref] [PubMed]
  28. Sterne JA, White IR, Carlin JB, et al. Multiple imputation for missing data in epidemiological and clinical research: potential and pitfalls. BMJ 2009;338:b2393. [Crossref] [PubMed]
  29. Clift AK, Dodwell D, Lord S, et al. Development and internal-external validation of statistical and machine learning models for breast cancer prognostication: cohort study. BMJ 2023;381:e073800. [Crossref] [PubMed]
  30. Clift AK, Coupland CAC, Keogh RH, et al. Living risk prediction algorithm (QCOVID) for risk of hospital admission and mortality from coronavirus 19 in adults: national derivation and validation cohort study. BMJ 2020;371:m3731. [Crossref] [PubMed]
  31. Berg B, Gorosito MA, Fjeld O, et al. Machine Learning Models for Predicting Disability and Pain Following Lumbar Disc Herniation Surgery. JAMA Netw Open 2024;7:e2355024. [Crossref] [PubMed]
  32. Sen T, Dotsu Y, Corbett V, et al. Pulmonary neuroendocrine neoplasms: the molecular landscape, therapeutic challenges, and diagnosis and management strategies. Lancet Oncol 2025;26:e13-33. [Crossref] [PubMed]
  33. Crowson CS, Larson DR, Devick KL, et al. Living With Survival Analysis in Orthopedics. J Arthroplasty 2021;36:3358-61. [Crossref] [PubMed]
  34. Dey T, Lipsitz SR, Cooper Z, et al. Survival analysis-time-to-event data and censoring. Nat Methods 2022;19:906-8. [Crossref] [PubMed]
  35. Rajkomar A, Hardt M, Howell MD, et al. Ensuring Fairness in Machine Learning to Advance Health Equity. Ann Intern Med 2018;169:866-72. [Crossref] [PubMed]
  36. Chen RJ, Wang JJ, Williamson DFK, et al. Algorithmic fairness in artificial intelligence for medicine and healthcare. Nat Biomed Eng 2023;7:719-42. [Crossref] [PubMed]
  37. Chohlas-Wood A, Coots M, Goel S, et al. Designing equitable algorithms. Nat Comput Sci 2023;3:601-10. [Crossref] [PubMed]
  38. Poirion OB, Jing Z, Chaudhary K, et al. DeepProg: an ensemble of deep-learning and machine-learning models for prognosis prediction using multi-omics data. Genome Med 2021;13:112. [Crossref] [PubMed]
  39. Fisch D, Bozorgmehr F, Kazdal D, et al. Comprehensive Dissection of Treatment Patterns and Outcome for Patients With Metastatic Large-Cell Neuroendocrine Lung Carcinoma. Front Oncol 2021;11:673901. [Crossref] [PubMed]
  40. Xia L, Wang L, Zhou Z, et al. Treatment outcome and prognostic analysis of advanced large cell neuroendocrine carcinoma of the lung. Sci Rep 2022;12:16562. [Crossref] [PubMed]
  41. Altuhaifa FA, Win KT, Su G. Predicting lung cancer survival based on clinical data using machine learning: A review. Comput Biol Med 2023;165:107338. [Crossref] [PubMed]
  42. Christodoulou E, Ma J, Collins GS, et al. A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. J Clin Epidemiol 2019;110:12-22. [Crossref] [PubMed]
  43. Le Treut J, Sault MC, Lena H, et al. Multicentre phase II study of cisplatin-etoposide chemotherapy for advanced large-cell neuroendocrine lung carcinoma: the GFPC 0302 study. Ann Oncol 2013;24:1548-52. [Crossref] [PubMed]
  44. Sun JM, Ahn MJ, Ahn JS, et al. Chemotherapy for pulmonary large cell neuroendocrine carcinoma: similar to that for small cell lung cancer or non-small cell lung cancer? Lung Cancer 2012;77:365-70. [Crossref] [PubMed]
  45. La Salvia A, Persano I, Siciliani A, et al. Prognostic significance of laterality in lung neuroendocrine tumors. Endocrine 2022;76:733-46. [Crossref] [PubMed]
  46. Efthimiou O, Seo M, Chalkou K, et al. Developing clinical prediction models: a step-by-step guide. BMJ 2024;386:e078276. [Crossref] [PubMed]
Cite this article as: Gong X, Pan M, Lin Y, Ye X, Qian J, Liao G, Du J, Zheng B, Chen C, Yang Z. Prognostic models for large cell neuroendocrine lung carcinoma: a machine learning and regression approach. Transl Lung Cancer Res 2025;14(7):2470-2482. doi: 10.21037/tlcr-2025-130

Download Citation