Development and validation of machine learning diagnostic models integrating clinical, CT, and laboratory features to differentiate lung cancer from pulmonary tuberculosis in patients with solitary pulmonary nodules: a single-center retrospective study
Original Article

Development and validation of machine learning diagnostic models integrating clinical, CT, and laboratory features to differentiate lung cancer from pulmonary tuberculosis in patients with solitary pulmonary nodules: a single-center retrospective study

Yuhang Li1,2#, Weijun Wu3#, Huiling Xu3, Jun Ma4, Minwei Bao5, Junjie Zhu6, Shanhao Chen4

1Department of Medical Oncology, Shanghai Pulmonary Hospital, School of Medicine, Tongji University, Shanghai, China; 2School of Medicine, Tongji University, Shanghai, China; 3Department of Radiology, Tongren Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China; 4Clinic and Research Center of Tuberculosis, Shanghai Pulmonary Hospital, School of Medicine, Tongji University, Shanghai, China; 5Department of Thoracic Surgery, Shanghai Pulmonary Hospital, School of Medicine, Tongji University, Shanghai, China; 6Innovation and Incubation Center (IIC), Shanghai Pulmonary Hospital, School of Medicine, Tongji University, Shanghai, China

#These authors contributed equally to this work as co-first authors.

Correspondence to: Shanhao Chen, MD. Clinic and Research Center of Tuberculosis, Shanghai Pulmonary Hospital, Tongji University School of Medicine, No. 507 Zhengmin Road, Shanghai 200433, China. Email: cshddd@163.com; Junjie Zhu, PhD. Innovation and Incubation Center (IIC), Shanghai Pulmonary Hospital, School of Medicine, Tongji University, No. 507 Zhengmin Road. Shanghai 200433, China. Email: zhujunjie@tongji.edu.cn.

Background: Differentiating lung cancer from pulmonary tuberculosis in patients with solitary pulmonary nodules (SPNs) remains clinically challenging, particularly in tuberculosis-endemic settings, because these two conditions may show overlapping computed tomography (CT) morphological features. This study aimed to develop and internally validate machine learning (ML) diagnostic models integrating these routinely available variables and to identify stable discriminative features using feature-importance analyses.

Methods: This single-center retrospective diagnostic prediction model development and internal validation study included adult patients with CT-detected SPNs measuring ≤3 cm and a definitive diagnosis of primary lung cancer or pulmonary tuberculosis between May 2020 and June 2024. Candidate predictors were extracted from baseline clinical information, tuberculosis-related tests, manually assessed CT morphological features, circulating tumor cell indicators, routine laboratory tests, and blood gas analysis variables obtained within 7 days before surgery or biopsy. Variables with more than 5% missingness were excluded, and 72 variables were retained before least absolute shrinkage and selection operator (LASSO) feature selection. The dataset was randomly divided into training and validation sets in a stratified 7:3 ratio. Seven ML models were subsequently constructed, including logistic regression (LR), random forest (RF), extra trees (ET), radial basis function support vector machine (RBF-SVM), k-nearest neighbors (KNN), multilayer perceptron (MLP), and gradient boosting decision tree (GBDT). Model performance was evaluated using the area under the curve (AUC) for the receiver operating characteristic (ROC) curve, sensitivity, specificity, and balanced accuracy. Permutation importance and SHapley Additive exPlanations (SHAP) analyses were further performed to assess model interpretability.

Results: A total of 431 patients were included, comprising 168 patients with pathologically confirmed lung cancer and 263 patients with pulmonary tuberculosis. The median age was 60.00 years (interquartile range, 53.00–67.00 years), 277 patients (64.3%) were male, 349 patients (81.0%) had solid nodules, and 277 patients (64.3%) had positive QuantiFERON-TB (QFT) results. After missingness filtering, 72 variables were retained, and 29 features were selected by LASSO for model development. Among the seven models, the RBF-SVM model achieved the highest validation AUC of 0.803, with a sensitivity of 0.686, specificity of 0.734, and balanced accuracy of 0.710. The ET model showed comparable validation performance, with an AUC of 0.791, whereas LR achieved the highest balanced accuracy of 0.711. Cross-model interpretability analyses identified nodule type and age as the most stable core features.

Conclusions: ML models integrating routine clinical, manually assessed CT, and laboratory features showed moderate discriminative performance for differentiating lung cancer from pulmonary tuberculosis in patients with SPNs. The RBF-SVM model achieved the highest validation AUC, and cross-model interpretability analyses identified nodule type and age as stable discriminative features.

Keywords: Solitary pulmonary nodule (SPN); lung cancer; pulmonary tuberculosis; differential diagnosis; machine learning (ML)


Submitted May 08, 2026. Accepted for publication May 29, 2026. Published online Jul 20, 2026.

doi: 10.21037/tlcr-2026-0559


Highlight box

Key findings

• Machine learning (ML) models that integrate baseline clinical information, manually assessed computed tomography morphological features, and routine laboratory parameters showed moderate discriminative performance in distinguishing lung cancer from pulmonary tuberculosis in patients with solitary pulmonary nodules (SPNs).

• Radial basis function support vector machine achieved the highest validation area under the curve , while logistic regression and extra trees showed comparable balanced classification performance.

• Nodule type and age were the most stable core discriminative features across the models.

What is known, and what is new?

• Previous studies have shown that integrating multimodal information can improve the differentiation between lung cancer and pulmonary tuberculosis in patients with SPNs.

• This study focused on the highly challenging clinical scenario of distinguishing lung cancer from pulmonary tuberculosis. Without relying on costly molecular testing, we established a multi-model ML framework based on routinely available clinical data and further identified stable cross-model features using permutation importance and SHapley Additive exPlanations analyses.

What is the implication, and what should change now?

• Joint modeling of routine clinical, imaging, and laboratory information has potential clinical utility and may provide useful support for the differential diagnosis of SPNs.

• Future studies should further validate the robustness and generalizability of these models in multicenter external cohorts, and facilitate their translation into clinical decision-support tools.


Introduction

With the widespread use of low-dose computed tomography (CT) in lung cancer screening and pulmonary nodule management, the detection rate of pulmonary nodules, particularly solitary pulmonary nodules (SPNs), has increased substantially (1-3). An SPN is generally defined as a single, round or oval intrapulmonary lesion measuring no more than 3 cm in diameter, completely surrounded by lung parenchyma, and not associated with obvious atelectasis, hilar or mediastinal lymphadenopathy, or pleural effusion (4-6). Although most SPNs are ultimately benign, a proportion represent early-stage lung cancer. Therefore, accurate risk stratification and differentiation between lung cancer and pulmonary tuberculosis in patients with SPNs are of great importance for optimizing clinical decision-making and improving patient outcomes (1,7).

Among the differential diagnoses of SPNs, distinguishing pulmonary tuberculosis from early-stage lung cancer is particularly challenging, especially in regions with a high tuberculosis burden, where this issue has even greater clinical relevance (8-10). Pulmonary tuberculosis presenting as an SPN may closely resemble early lung cancer in terms of nodule type, lobulation, spiculation, pleural indentation, and patterns of calcification; in particular, non-calcified tuberculosis-related solitary nodules may closely mimic malignant lesions. Previous studies specifically addressing the differentiation between pulmonary tuberculosis and lung cancer have suggested that reliance on imaging interpretation alone is associated with a substantial risk of misclassification, which may explain why such lesions are often regarded as “radiologically suspicious for tumor, yet difficult to definitively characterize before pathology” in clinical practice (11).

At present, the evaluation of SPNs mainly relies on imaging follow-up, percutaneous biopsy, and bronchoscopic examination (12). In the differential diagnosis of SPNs, distinguishing pulmonary tuberculosis from early-stage lung cancer remains a persistent clinical challenge, particularly in high-burden tuberculosis settings. Although manually assessed CT morphological features provide important baseline information for clinical risk stratification, the substantial overlap in radiological phenotypes between lung cancer and pulmonary tuberculosis may lead to misclassification when diagnostic decisions rely solely on imaging interpretation. Therefore, integrating clinical data with routine laboratory indicators may offer a more comprehensive assessment (13-15). Although invasive procedures can provide pathological evidence, their utility is constrained by lesion location, nodule size, sampling conditions, and the risk of procedure-related complications. Meanwhile, traditional serum biomarkers, liquid biopsy, and other emerging biomarkers have shown certain promise; however, their widespread implementation in real-world practice remains limited by cost, accessibility, and insufficient evidence of clinical utility (16,17). Therefore, developing a comprehensive evaluation tool that is based on routine clinical workflows, relatively non-invasive, accurate, and broadly applicable has clear clinical value.

Previous studies have explored several approaches for differentiating lung cancer from tuberculosis-related or granulomatous nodules, including conventional CT morphological assessment, clinical-imaging models, radiomics models, deep learning-based imaging algorithms, and multimodal biomarker models. Comparative studies and systematic reviews have shown that clinical and CT features, such as age, lesion size, lesion margin, calcification pattern, lobulation, spiculation, pleural indentation, and enhancement-related characteristics, may help distinguish lung cancer from pulmonary tuberculosis or granulomatous lesions, although substantial radiological overlap remains (18,19). Radiomics-based studies have further incorporated high-dimensional quantitative imaging features, sometimes in combination with clinical variables, to improve the preoperative differentiation of tuberculous granulomas or granulomatous lesions from lung adenocarcinoma (10,20,21). Deep learning-based imaging models have also been developed for this specific diagnostic scenario, reflecting the recent progress of artificial intelligence methods in pulmonary nodule classification (22). In addition, circulating tumor cell (CTC)-based diagnostic models and combined clinical, imaging, and cell-free DNA methylation biomarker models have been proposed for pulmonary nodule classification (14,15). For example, in a SPN cohort, folate receptor-positive circulating tumor cell (FR+ CTC) was identified as an independent predictor, but its standalone diagnostic performance was moderate, with AUCs of 0.650 and 0.700 in the training and validation sets, respectively (14).

Although these studies have reported favorable discriminatory performance in selected cohorts, their clinical translation remains limited by several factors. Radiomics and deep learning models usually require lesion segmentation, quantitative feature extraction, high-dimensional feature selection, and dedicated computational workflows, which may restrict their routine use and increase the risk of over-parameterization or unstable feature selection in limited datasets (10,20-22). Molecular biomarker-based models may also be constrained by cost, accessibility, assay availability, and differences in clinical implementation across institutions (14,15). Furthermore, complex imaging algorithms are often difficult to interpret clinically, and models developed from limited or selected cohorts still require further validation before broad clinical application (10,22). Many previous studies have focused primarily on imaging features or molecular biomarkers, whereas the contribution of routinely available laboratory indicators and tuberculosis-related immunological information has not been fully clarified in interpretable models specifically designed for differentiating lung cancer from pulmonary tuberculosis. Therefore, an interpretable model integrating routine clinical variables, manually assessed CT features, tuberculosis-related immunological information, and laboratory data may provide a practical complementary framework for preliminary differential diagnosis.

In recent years, machine learning (ML) has attracted increasing attention in pulmonary nodule risk assessment and the assisted diagnosis of lung cancer (23-25). Compared with traditional univariable analyses or linear statistical approaches, ML is better suited to integrating multidimensional heterogeneous data, including clinical information, imaging features, and laboratory indicators, thereby enabling the identification of potential nonlinear relationships and complex interaction patterns (23). Previous studies have shown that combining clinical data, imaging features, and laboratory indicators can improve the diagnostic accuracy of pulmonary nodules. However, in the specific clinical setting of differentiating lung cancer from pulmonary tuberculosis, where imaging manifestations overlap substantially, studies that systematically integrate manually assessed CT morphological features with basic clinical information and routine laboratory data remain relatively limited, and the stability and relative contribution of key features have not yet been fully evaluated.

Thus, this study focused on the clinically important and specific problem of differentiating lung cancer from pulmonary tuberculosis in patients with SPNs. Based on basic clinical information, we further integrated manually assessed CT morphological features and multidimensional routine laboratory indicators into discriminative models to systematically evaluate their value in differentiating lung cancer from pulmonary tuberculosis in patients with SPNs. Rather than merely emphasizing the superiority of a particular algorithm, this study sought to construct integrated discriminative models based on routinely available data with potential clinical utility. In addition, feature importance and interpretability analyses were used to identify stable key feature sets, thereby enhancing the clinical interpretability and potential translational value of the findings. This research strategy is also consistent with current expectations for standardized development, transparent reporting, and the appropriate validation of clinical prediction models (26-28). We present this article in accordance with the TRIPOD reporting checklist (available at https://tlcr.amegroups.com/article/view/10.21037/tlcr-2026-0559/rc).


Methods

Study design and ethical approval

This was a single-center retrospective diagnostic prediction model development and internal validation study. The study included patients with SPNs treated at Shanghai Pulmonary Hospital between May 2020 and June 2024. The study protocol was approved by the Ethics Committee of Shanghai Pulmonary Hospital (No. K23-233Y), which waived the requirement for informed consent due to the retrospective design of the study. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.

Study population

Patient screening and enrollment

Consecutive patients with SPNs detected on chest CT were retrospectively identified through the hospital electronic medical record system. An SPN was defined as a solitary round or oval lesion measuring ≤3 cm in diameter, completely surrounded by lung parenchyma, with clear margins and without atelectasis, hilar enlargement, or pleural effusion. The final diagnosis of lung cancer was established by pathological examination. The diagnosis of pulmonary tuberculosis was established based on pathological findings in conjunction with clinical, radiological, and, when available, microbiological evidence.

Inclusion and exclusion criteria

The inclusion criteria were as follows: (I) age ≥18 years; (II) an SPN detected on chest CT, with a maximum diameter ≤3 cm; (III) a definitive diagnosis of primary lung cancer or pulmonary tuberculosis; and (IV) availability of clinical, imaging, and laboratory data within 7 days before surgery or biopsy. The exclusion criteria were as follows: (I) a history of other malignancies within the previous 5 years; (II) concomitant pulmonary infectious diseases other than active pulmonary tuberculosis; (III) previous anti-tuberculosis or anti-tumor treatment; (IV) severe hepatic or renal dysfunction, autoimmune disease, or ongoing immunosuppressive therapy; and/or (V) substantial missing key information precluding model development. Ultimately, 431 patients were included in the study, of whom 168 had lung cancer and 263 had pulmonary tuberculosis.

Variable extraction

Candidate predictors were selected based on clinical relevance, previous literature, and data availability. The selected variables were intended to capture three clinically relevant dimensions involved in the differential diagnosis of lung cancer and pulmonary tuberculosis in patients with SPNs. First, local nodule morphology was represented by manually assessed CT features, including nodule type, calcification, spiculation, lobulation, pleural indentation, and cavitation. Second, host background and tuberculosis-related information were represented by demographic variables and QuantiFERON-TB (QFT) results. Third, systemic biological status was represented by CTC indicators, routine hematological, coagulation, biochemical, and blood gas analysis variables obtained within 7 days before surgery or biopsy.

CT morphological features were assessed using a two-reader plus senior-adjudication workflow. Two radiologists with experience in thoracic imaging independently assessed CT morphological features while blinded to the final pathological or clinical diagnosis. Before formal assessment, both radiologists reviewed predefined imaging criteria for nodule type, calcification, spiculation, lobulation, pleural indentation, and cavitation to standardize feature interpretation. Each CT feature was first recorded independently. Interobserver agreement between the two initial readers for categorical CT features was assessed using Cohen’s kappa coefficients. Discrepancies were subsequently resolved by a senior thoracic radiologist, and the final adjudicated results were used for model development.

The predicted outcome was the final diagnosis, defined as lung cancer versus pulmonary tuberculosis. Lung cancer was coded as the positive class (label =1), and pulmonary tuberculosis was coded as the negative class (label =0).

Data preprocessing

Before model development, candidate variables were assessed for missingness in the overall cohort, and those with a missing rate >5% were excluded. After missingness filtering and numeric usability assessment, 72 variables were retained for model development. The dataset was then randomly divided into a training set (n=301) and a validation set (n=130) in a stratified 7:3 ratio. Subsequent preprocessing was performed in a training set-driven manner to avoid data leakage. Specifically, categorical variables were imputed using the mode, while continuous variables were imputed using iterative imputation; median imputation was used when iterative imputation could not be implemented. Continuous variables were then standardized using Z-score normalization. The imputation and scaling parameters were estimated in the training set and subsequently applied to the validation set.

Feature selection

After missing-value imputation and standardization, feature selection was performed using the training set only. Specifically, an L1-regularization-based least absolute shrinkage and selection operator (LASSO) screening procedure with five-fold cross-validation was used for feature selection in the training set. The penalty parameter was tuned over a prespecified logarithmic sequence, and variables with non-zero coefficients at the optimal penalty level were retained for subsequent model development.

Model development and hyperparameter optimization

Based on the feature set selected by LASSO, seven binary classification models were developed, including logistic regression (LR), random forest (RF), extra trees (ET), radial basis function support vector machine (RBF-SVM), k-nearest neighbors (KNN), multilayer perceptron (MLP), and gradient boosting decision tree (GBDT).

Hyperparameter optimization for all models was performed exclusively in the training set using GridSearchCV with three-fold stratified cross-validation, with the area under the curve (AUC) for the receiver operating characteristic (ROC) curve serving as the criterion for optimal parameter selection. After the optimal hyperparameter combination was identified, each corresponding model was refitted using the full training set to generate the final model for subsequent threshold optimization and performance evaluation.

Threshold selection and model evaluation

To avoid the potential bias associated with directly using a fixed threshold of 0.5 for classification, the classification threshold was further optimized in the training set. Specifically, for each optimally tuned model, five-fold stratified cross-validation was performed in the training set to obtain out-of-fold (OOF) predicted probabilities for each training sample. Candidate thresholds were evaluated over 501 grid points between 0 and 1, and the final threshold was selected using a predefined training-set-based criterion designed to balance sensitivity and specificity. Threshold determination was based solely on the OOF predictions from the training set, while the internal hold-out validation set was not involved in threshold optimization and was used only for final validation.

The performance of each model was evaluated in the training set and further validated in the internal hold-out validation set. Evaluation metrics included the AUC, sensitivity, specificity, and balanced accuracy. OOF predictions obtained in the training set were used only for threshold determination, while the internal hold-out validation set was used for final performance assessment.

Given the exploratory nature of this internally validated model, no definitive clinical decision threshold was prespecified for direct clinical implementation. Instead, thresholds were selected to balance sensitivity and specificity within the training set. The resulting models were intended to serve as preliminary decision-support tools to assist risk stratification, rather than as stand-alone tools for ruling in or ruling out lung cancer or pulmonary tuberculosis.

Model interpretability analysis

In all binary classification analyses, lung cancer was coded as the positive class (label =1) and pulmonary tuberculosis as the negative class (label =0). To evaluate the contributions of input features to the predictions of the final models, feature importance and interpretability analyses were further performed. First, permutation importance was applied in the internal hold-out validation set to assess feature importance. Specifically, based on the final models obtained after hyperparameter optimization and refitting on the full training set, each feature in the validation set was randomly permuted in turn. The AUC was used as the scoring metric, and feature importance was quantified by the mean and standard deviation over 30 repetitions.

Subsequently, SHapley Additive exPlanations (SHAP) analysis was used to interpret model predictions. SHAP analysis was conducted on the feature matrix after missing-value imputation, standardization, and feature selection. To balance computational efficiency and result stability, 200 samples were randomly selected from the training set as the background dataset, and all samples from the validation set were included for explanation. TreeExplainer was used for the tree-based models, while KernelExplainer was used for the non-tree-based models. For the binary classification task, SHAP values corresponding to the positive class were extracted, and the global importance of each feature was quantified by the mean absolute SHAP value. SHAP summary bar plots and summary dot plots were also generated.

For cross-model comparison, permutation importance values and mean absolute SHAP values were normalized in each model before calculating cross-model averages. All analyses were performed in Python, mainly using the pandas, numpy, scikit-learn, shap, and matplotlib packages.

Sample size consideration and parsimonious-model sensitivity analysis

Because this was a retrospective diagnostic prediction model development and internal validation study, no formal prospective sample size calculation was performed before patient enrollment. Instead, all eligible patients during the study period were included to maximize the available sample size. The final cohort included 431 patients, including 168 lung cancer events. In the primary analysis, 29 predictors were retained after LASSO feature selection, corresponding to an event-per-variable ratio of approximately 5.8. Given the relatively limited number of positive events in relation to the selected feature set, additional sensitivity analyses were performed to assess the robustness of the modeling results.

All feature selection, hyperparameter tuning, and threshold optimization procedures were performed using the training set only, and the internal hold-out validation set was reserved exclusively for final performance assessment. First, a parsimonious-model sensitivity analysis was conducted using only the top 10 predictors ranked by the absolute values of the LASSO coefficients derived from the training set. All data splitting, imputation, scaling, hyperparameter tuning, and threshold-selection procedures were otherwise unchanged. This analysis was intended to assess model parsimony and robustness rather than to replace external validation.

Complete-case sensitivity analysis

As a sensitivity analysis for missing-data handling, we repeated model development and evaluation using only complete cases without missing-value imputation. Complete cases were defined as patients with non-missing values for the 29 predictors selected in the primary analysis. The original stratified 7:3 training-validation split was retained, and patients with missing values in any of these 29 predictors were excluded separately from the training and internal hold-out validation sets. The same 29 predictors were then used for model training and evaluation without re-running LASSO feature selection. This analysis was performed to evaluate whether the main findings were materially influenced by missing-value imputation.

Statistical analysis

Continuous variables were summarized as median and interquartile range, and categorical variables were summarized as counts and percentages. Between-group comparisons were performed using appropriate nonparametric tests for continuous variables and chi-square or Fisher’s exact tests for categorical variables, as appropriate. Model discrimination was evaluated using the area under the receiver operating characteristic curve (AUC). Sensitivity, specificity, and balanced accuracy were calculated according to the training-set-derived classification threshold. All preprocessing, feature selection, model tuning, and threshold optimization procedures were performed using the training set only, and the validation set was reserved for final performance assessment. Statistical analyses were performed using Python packages including pandas, numpy, scikit-learn, shap, and matplotlib.


Results

Results of data preprocessing and feature selection

A total of 431 patients with SPNs were included in this study, including 168 patients with pathologically confirmed lung cancer and 263 patients with SPNs ultimately diagnosed as pulmonary tuberculosis. Candidate variable missingness was first assessed in the overall cohort, and variables with a missing rate >5% were removed. After missingness filtering and numeric usability assessment, 72 variables were retained for subsequent model development. The dataset was then randomly divided into a training set and a validation set at a ratio of 7:3 using stratified sampling. The final training set included 301 patients and the validation set included 130 patients. The overall distributions of the main input variables were broadly similar between the training and validation sets, with no obvious imbalance (Figure 1 and Table 1).

Figure 1 Study design and machine learning workflow. A total of 438 patients with SPNs from Shanghai Pulmonary Hospital were initially screened, and 431 patients were included in the final analytical cohort. The final cohort was randomly divided into a training set (n=301) and a validation set (n=130) at a ratio of 7:3. The candidate variables included imaging features, tumor biomarkers, routine serological indices, and tuberculosis-related markers. After feature selection using the LASSO, the selected variables were used to develop multiple ML models, including RBF-SVM, RF, KNN, LR, ET, MLP, and GBDT. Model performance was evaluated using ROC curves, and feature importance analysis was performed to assess the contribution of individual variables to model predictions. AUC, area under the curve; ET, extra trees; GBDT, gradient boosting decision tree; KNN, k-nearest neighbors; LASSO, least absolute shrinkage and selection operator; LR, logistic regression; ML, machine learning; MLP, multilayer perceptron; RBF-SVM, radial basis function support vector machine; RF, random forest; ROC, receiver operating characteristic; SPNs, solitary pulmonary nodules; TB, tuberculosis.

Table 1

Baseline characteristics of the study population and their distribution in the training and validation sets

Variable Overall (n=431) Training set (n=301) Validation set (n=130) P value Missing overall Missing train Missing validation
QFT 0.72 0 0 0
   Negative 144 (33.4) 102 (33.9) 42 (32.3)
   Positive 277 (64.3) 191 (63.5) 86 (66.2)
   Uncertain 10 (2.3) 8 (2.7) 2 (1.5)
Gender 0.38 0 0 0
   Female 154 (35.7) 103 (34.2) 51 (39.2)
   Male 277 (64.3) 198 (65.8) 79 (60.8)
Nodule type 0.84 0 0 0
   Ground-glass nodule 82 (19.0) 56 (18.6) 26 (20.0)
   Solid nodule 349 (81.0) 245 (81.4) 104 (80.0)
Calcification 0.69 0 0 0
   Absent 412 (95.6) 289 (96.0) 123 (94.6)
   Present 19 (4.4) 12 (4.0) 7 (5.4)
Spiculation 0.89 0 0 0
   Absent 335 (77.7) 235 (78.1) 100 (76.9)
   Present 96 (22.3) 66 (21.9) 30 (23.1)
Lobulation 0.77 0 0 0
   Absent 301 (69.8) 212 (70.4) 89 (68.5)
   Present 130 (30.2) 89 (29.6) 41 (31.5)
Pleural indentation 0.43 0 0 0
   Absent 340 (78.9) 241 (80.1) 99 (76.2)
   Present 91 (21.1) 60 (19.9) 31 (23.8)
Cavitation 0.40 0 0 0
   Absent 369 (85.6) 261 (86.7) 108 (83.1)
   Present 62 (14.4) 40 (13.3) 22 (16.9)
Age (years) 60.00 (53.00, 67.00) 60.00 (54.00, 67.00) 60.00 (51.00, 66.00) 0.46 0 0 0
CTC (counts/3 mL) 10.72 (7.76, 15.37) 10.72 (7.76, 15.38) 10.74 (7.78, 15.09) 0.79 0 0 0
Hb (g/L) 142.00 (130.00, 152.00) 142.00 (130.00, 153.00) 141.00 (129.25, 151.00) 0.56 0 0 0
RBC_abs (×1012/L) 4.62 (4.25, 4.95) 4.63 (4.26, 4.96) 4.61 (4.25, 4.88) 0.40 0 0 0
WBC_abs (×109/L) 5.92 (4.97, 7.04) 5.88 (4.99, 7.07) 5.99 (4.89, 6.94) 0.84 0 0 0
N (%) 63.60 (58.45, 69.65) 63.70 (59.20, 70.10) 63.55 (57.35, 69.15) 0.40 0 0 0
L (%) 26.40 (21.20, 31.10) 26.30 (20.90, 30.90) 26.60 (22.50, 31.93) 0.24 0 0 0
M (%) 7.20 (6.00, 8.45) 7.10 (6.00, 8.40) 7.30 (6.10, 8.50) 0.65 0 0 0
E (%) 1.40 (0.80, 2.40) 1.40 (0.80, 2.50) 1.30 (0.80, 2.27) 0.55 0 0 0
B (%) 0.50 (0.30, 0.70) 0.50 (0.30, 0.70) 0.50 (0.40, 0.60) 0.35 0 0 0
N_abs (×109/L) 3.70 (3.04, 4.70) 3.68 (3.07, 4.73) 3.74 (2.94, 4.68) 0.88 0 0 0
L_abs (×109/L) 1.52 (1.21, 1.90) 1.50 (1.18, 1.91) 1.58 (1.29, 1.90) 0.15 0 0 0
M_abs (×109/L) 0.41 (0.34, 0.52) 0.41 (0.34, 0.51) 0.44 (0.34, 0.53) 0.41 0 0 0
E_abs (×109/L) 0.08 (0.04, 0.14) 0.08 (0.04, 0.14) 0.08 (0.04, 0.14) 0.62 0 0 0
B_abs (×109/L) 0.03 (0.02, 0.04) 0.03 (0.02, 0.04) 0.03 (0.02, 0.04) 0.42 0 0 0
PLT (×109/L) 218.00 (182.50, 264.00) 218.00 (182.00, 261.00) 219.00 (186.25, 268.75) 0.40 0 0 0
PCV (L/L) 0.42 (0.39, 0.45) 0.42 (0.39, 0.45) 0.42 (0.39, 0.45) 0.36 1 1 0
MCV (fL) 91.30 (88.85, 93.85) 91.40 (88.70, 93.90) 91.05 (89.10, 93.70) 0.70 0 0 0
MCH (pg) 30.70 (29.75, 31.70) 30.70 (29.60, 31.70) 30.80 (30.00, 31.77) 0.61 0 0 0
MCHC (g/L) 336.00 (328.00, 342.00) 336.00 (329.00, 342.00) 336.00 (327.25, 344.00) 0.49 0 0 0
RDW-CV (%) 12.70 (12.20, 13.25) 12.70 (12.30, 13.30) 12.60 (12.10, 13.20) 0.13 0 0 0
RDW-SD (fL) 42.20 (40.45, 44.60) 42.20 (40.60, 44.90) 42.20 (39.92, 44.38) 0.23 0 0 0
PCT (%) 0.23 (0.20, 0.27) 0.23 (0.20, 0.27) 0.23 (0.20, 0.28) 0.36 5 4 1
PDW (fL) 12.20 (10.90, 13.88) 12.30 (10.90, 13.70) 12.10 (10.80, 14.20) 0.90 5 4 1
P-LCR (%) 29.35 (23.82, 35.18) 29.50 (23.90, 34.80) 29.00 (23.00, 36.20) 0.96 5 4 1
MPV (fL) 10.50 (9.90, 11.20) 10.60 (9.90, 11.20) 10.50 (9.80, 11.30) >0.99 5 4 1
PT (s) 11.00 (10.40, 11.50) 11.00 (10.40, 11.50) 11.00 (10.38, 11.50) 0.78 2 0 2
INR 0.99 (0.93, 1.03) 0.99 (0.93, 1.03) 0.98 (0.94, 1.03) 0.91 2 0 2
APTT (s) 30.80 (28.90, 32.70) 31.00 (29.00, 32.70) 30.40 (28.80, 32.35) 0.29 2 0 2
TT (s) 14.20 (13.40, 14.90) 14.20 (13.40, 14.90) 14.25 (13.50, 15.20) 0.69 2 0 2
Fbg (g/L) 2.98 (2.70, 3.37) 3.00 (2.74, 3.35) 2.88 (2.58, 3.38) 0.23 2 0 2
FDP (μg/mL) 0.71 (0.37, 1.23) 0.71 (0.34, 1.26) 0.74 (0.45, 1.14) 0.64 3 1 2
ATIII (%) 104.00 (94.00, 113.00) 105.00 (95.00, 114.00) 103.00 (93.00, 112.00) 0.31 2 0 2
GGT (U/L) 19.50 (13.85, 30.00) 19.80 (14.00, 29.10) 19.15 (13.35, 31.45) 0.97 0 0 0
ALT (U/L) 18.00 (13.70, 25.35) 18.10 (13.50, 25.50) 18.00 (14.10, 24.35) 0.93 0 0 0
AST (U/L) 18.70 (14.60, 22.70) 19.00 (15.10, 23.00) 18.00 (13.85, 21.93) 0.08 0 0 0
ALP (U/L) 76.00 (62.00, 90.35) 76.00 (62.00, 90.40) 76.80 (63.00, 90.00) 0.76 1 0 1
TBil (μmol/L) 11.70 (9.00, 15.00) 11.90 (9.00, 15.40) 11.10 (8.85, 14.67) 0.27 0 0 0
DBil (μmol/L) 3.70 (2.90, 5.00) 3.80 (3.00, 5.10) 3.45 (2.80, 4.97) 0.20 0 0 0
TP (g/L) 71.00 (68.00, 74.05) 71.10 (68.00, 74.50) 70.90 (67.80, 74.00) 0.27 0 0 0
Albumin (g/L) 43.30 (41.15, 45.40) 43.20 (41.20, 45.50) 43.45 (41.05, 45.08) 0.90 0 0 0
Globin (g/L) 27.70 (25.05, 30.00) 28.00 (25.60, 30.10) 27.00 (24.25, 29.88) 0.07 0 0 0
A/G 1.60 (1.40, 1.75) 1.60 (1.40, 1.70) 1.60 (1.40, 1.80) 0.15 2 2 0
Uric acid (μmol/L) 320.80 (260.90, 387.05) 321.00 (264.70, 394.50) 319.40 (253.48, 373.62) 0.52 0 0 0
BUN (mmol/L) 5.96 (4.90, 7.11) 5.97 (4.90, 7.10) 5.91 (4.90, 7.11) 0.94 0 0 0
Cre (μmol/L) 64.00 (55.00, 74.15) 65.20 (56.00, 74.30) 61.00 (53.17, 73.00) 0.03 0 0 0
Glucose (mmol/L) 6.22 (5.10, 8.63) 6.18 (5.10, 8.86) 6.30 (5.10, 8.09) 0.95 13 11 2
K (mmol/L) 3.95 (3.73, 4.20) 3.94 (3.73, 4.18) 3.98 (3.73, 4.22) 0.83 0 0 0
Na (mmol/L) 141.60 (140.10, 143.70) 141.70 (140.10, 143.60) 141.55 (139.80, 143.88) 0.71 0 0 0
Cl (mmol/L) 105.60 (103.65, 107.30) 105.50 (103.90, 107.30) 105.90 (103.22, 107.20) 0.71 0 0 0
Ca (mmol/L) 2.39 (2.30, 2.50) 2.39 (2.30, 2.50) 2.38 (2.30, 2.49) 0.51 0 0 0
pH 7.41 (7.40, 7.43) 7.41 (7.40, 7.43) 7.41 (7.40, 7.43) 0.96 21 17 4
PaCO2 (mmHg) 39.65 (37.40, 42.10) 39.90 (37.50, 42.23) 38.95 (36.45, 41.50) 0.02 21 17 4
THB (g/dL) 14.60 (13.30, 15.70) 14.50 (13.30, 15.70) 14.70 (13.20, 15.70) 0.77 21 17 4
SpO2 (%) 97.15 (96.50, 97.70) 97.10 (96.50, 97.70) 97.20 (96.60, 97.70) 0.80 21 17 4
HbO2 (%) 95.60 (94.90, 96.20) 95.50 (94.90, 96.10) 95.70 (95.10, 96.40) 0.12 21 17 4
HbCO (%) 0.90 (0.70, 1.20) 0.90 (0.70, 1.20) 0.90 (0.70, 1.10) 0.34 21 17 4
Reduced Hb (%) 2.80 (2.23, 3.50) 2.80 (2.20, 3.50) 2.80 (2.30, 3.38) 0.88 21 17 4
Methemoglobin (%) 0.50 (0.40, 0.80) 0.55 (0.40, 0.80) 0.50 (0.30, 0.78) 0.053 21 17 4
Lactic acid (mmol/L) 1.10 (0.80, 1.44) 1.10 (0.80, 1.50) 1.10 (0.79, 1.40) 0.56 21 17 4
SBE (mmol/L) 0.60 (−0.90, 1.98) 0.70 (−0.70, 2.10) 0.10 (−1.77, 1.77) 0.02 21 17 4
ABE (mmol/L) 0.60 (−0.60, 1.80) 0.65 (−0.40, 1.90) 0.10 (−1.45, 1.58) 0.03 21 17 4
AB (mmol/L) 25.10 (23.60, 26.40) 25.20 (23.98, 26.50) 24.55 (22.90, 26.08) 0.01 21 17 4
SB (mmol/L) 25.00 (23.90, 26.00) 25.00 (24.08, 26.10) 24.50 (23.23, 25.77) 0.04 21 17 4

Data are presented as n (%) for categorical variables and median (interquartile range) for continuous variables. Missing data are presented as n. A/G, albumin/globulin ratio; AB, actual bicarbonate; ABE, actual base excess; ALP, alkaline phosphatase; ALT, alanine aminotransferase; APTT, activated partial thromboplastin time; AST, aspartate aminotransferase; ATIII, antithrombin III; B_abs, absolute basophil count; B%, basophil percentage; BUN, blood urea nitrogen; Cl, chloride; Cre, creatinine; CTC, circulating tumor cell; DBil, direct bilirubin; E_abs, absolute eosinophil count; E%, eosinophil percentage; Fbg, fibrinogen; FDP, fibrin degradation products; GGT, gamma-glutamyl transferase; Hb, hemoglobin; HbCO, carboxyhemoglobin; HbO2, oxyhemoglobin; INR, international normalized ratio; L_abs, absolute lymphocyte count; L%, lymphocyte percentage; M_abs, absolute monocyte count; M%, monocyte percentage; MCH, mean corpuscular hemoglobin; MCHC, mean corpuscular hemoglobin concentration; MCV, mean corpuscular volume; MPV, mean platelet volume; N_abs, absolute neutrophil count; N%, neutrophil percentage; P-LCR, platelet large cell ratio; PaCO2, partial pressure of carbon dioxide; PCT, plateletcrit; PCV, packed cell volume; PDW, platelet distribution width; pH, potential of hydrogen; PLT, platelet count; PT, prothrombin time; QFT, QuantiFERON-TB; RBC, red blood cell; RDW-CV, red cell distribution width-coefficient of variation; RDW-SD, red cell distribution width-standard deviation; Reduced Hb, reduced hemoglobin; SB, standard bicarbonate; SBE, standard base excess; SpO2, peripheral oxygen saturation; TBil, total bilirubin; THB, total hemoglobin measured by blood gas analysis; TP, total protein; TT, thrombin time; WBC, white blood cell.

Interobserver agreement between the two initial radiologists was assessed for manually evaluated CT morphological features. The Cohen’s kappa coefficients were 1.000 for nodule type, 1.000 for calcification, 0.973 for spiculation, 0.984 for lobulation, 0.993 for pleural indentation, and 1.000 for cavitation, indicating almost perfect interobserver agreement. Discrepant cases were resolved by a senior thoracic radiologist before model development.

After data splitting, variables were numerically encoded and imputed using variable type-specific strategies: categorical variables were imputed with the mode, while continuous variables were preferentially imputed using iterative imputation; if iterative imputation failed, median imputation was used. After imputation, all retained variables were standardized, and all 72 variables entered the feature selection stage. Based on the standardized training set, LASSO regression was used for feature shrinkage and selection. The optimal penalty parameter was determined by five-fold cross-validation (α=0.017270), and 29 features with non-zero coefficients were ultimately retained for model development. The selected features included QFT, nodule type, calcification, lobulation, pleural indentation, age, CTC, basophil percentage (B%), absolute lymphocyte count (L_abs), absolute eosinophil count (E_abs), packed cell volume (PCV), red cell distribution width-standard deviation (RDW-SD), platelet large cell ratio (P-LCR), activated partial thromboplastin time (APTT), thrombin time (TT), fibrin degradation products (FDP), aspartate aminotransferase (AST), alkaline phosphatase (ALP), uric acid, blood urea nitrogen (BUN), creatinine (Cre), glucose, potassium (K), chloride (Cl), partial pressure of carbon dioxide (PaCO2), total hemoglobin measured by blood gas analysis (THB), oxyhemoglobin (HbO2), reduced hemoglobin, and lactic acid. Overall, the final selected features covered imaging morphological characteristics, baseline clinical information, and hematologic, metabolic, and coagulation-related indicators, suggesting that the differentiation of SPNs does not rely on a single dimension of information, but rather on the joint contribution of multimodal features.

Development and performance comparison of multiple models

Based on the 29 features selected by LASSO, seven ML models were further developed and compared, including LR, RF, ET, RBF-SVM, KNN, MLP, and GBDT. Hyperparameters for all models were optimized in the training set using GridSearchCV with three-fold stratified cross-validation, with AUC as the criterion for optimal parameter selection. Subsequently, based on the five-fold OOF predicted probabilities in the training set (TRAIN-OOF), the optimal classification threshold was selected from 501 candidate thresholds using the prespecified training-set-based sensitivity-specificity criterion (Figure 2 and Table 2).

Figure 2 Integrated evaluation of classification performance across models. (A) Comparison of the ROC curves for the seven ML models in the training set. (B) Comparison of the ROC curves for the seven ML models in the validation set. (C) Comparison of the sensitivity, specificity, and balanced accuracy of each model in the training set. (D) Comparison of the sensitivity, specificity, and balanced accuracy of each model in the validation set. Overall, all models showed good discriminative ability in the training set, while performance declined to some extent in the validation set. RBF-SVM achieved the highest validation AUC, while ET and LR showed comparable balanced classification performance under the selected thresholds. AUC, area under the curve; ET, extra trees; GBDT, gradient boosting decision tree; KNN, k-nearest neighbors; LR, logistic regression; ML, machine learning; MLP, multilayer perceptron; RBF-SVM, radial basis function support vector machine; RF, random forest; ROC, receiver operating characteristic.

Table 2

Comparative diagnostic performance of the ML models in the training and validation sets

Model Training set Validation set
AUC Sensitivity Specificity Accuracy Balanced Acc AUC Sensitivity Specificity Accuracy Balanced Acc
LR 0.859 0.769 0.804 0.791 0.787 0.783 0.725 0.696 0.708 0.711
RF 0.922 0.838 0.804 0.817 0.821 0.756 0.647 0.709 0.685 0.678
ET 0.863 0.769 0.783 0.777 0.776 0.791 0.686 0.734 0.715 0.710
RBF-SVM 0.899 0.821 0.832 0.827 0.826 0.803 0.686 0.734 0.715 0.710
KNN 0.815 0.744 0.75 0.748 0.747 0.779 0.706 0.696 0.700 0.701
MLP 0.852 0.821 0.707 0.751 0.764 0.689 0.667 0.646 0.654 0.656
GBDT 0.914 0.812 0.832 0.824 0.822 0.753 0.569 0.734 0.669 0.651

AUC, area under the curve; Balanced Acc, balanced accuracy; ET, extra trees; GBDT, gradient boosting decision tree; KNN, k-nearest neighbors; LR, logistic regression; ML, machine learning; MLP, multilayer perceptron; RBF-SVM, radial basis function support vector machine; RF, random forest.

Overall, the training AUCs of the seven models ranged from 0.815 to 0.922, but performance declined to varying degrees in the TRAIN-OOF and validation analyses, indicating differences in generalizability across models. Using validation AUC as the primary performance metric, RBF-SVM achieved the highest validation AUC of 0.803, with a sensitivity of 0.686, specificity of 0.734, and balanced accuracy of 0.710. ET showed comparable validation performance, with an AUC of 0.791, sensitivity of 0.686, specificity of 0.734, and balanced accuracy of 0.710. LR achieved a validation AUC of 0.783 and showed the highest balanced accuracy of 0.711. KNN achieved a validation AUC of 0.779, with a sensitivity of 0.706, specificity of 0.696, and balanced accuracy of 0.701. RF achieved the highest training AUC but showed a lower validation AUC of 0.756, suggesting reduced generalization. In contrast, RBF-SVM showed the best validation AUC in the current analysis. MLP showed the lowest validation AUC (0.689), while GBDT achieved a validation AUC of 0.753.

Threshold optimization analysis showed that objective thresholds determined based on TRAIN-OOF predictions generally improved the balance of classification performance for some models in the validation set. For example, in the KNN model, the validation balanced accuracy increased from 0.606 under the fixed threshold of 0.5 to 0.701 after OOF-based threshold optimization, mainly because of an improvement in sensitivity. These findings suggest that threshold optimization mainly adjusted the operating points of the models by changing the trade-off between sensitivity and specificity, rather than uniformly improving all models.

Feature importance analysis (permutation importance)

To identify the key variables driving model discrimination, permutation importance was assessed in the validation set to evaluate the importance of the input features, using the AUC as the scoring metric and 30 repetitions to reduce random fluctuation. Cross-model permutation importance results showed that nodule type contributed most prominently across all models, with the highest cross-model mean normalized importance of 47.7%, followed by age at 19.8%, indicating that basic nodule morphology and patient age were the most critical discriminative factors for distinguishing lung cancer from pulmonary tuberculosis. In addition, glucose ranked third with a cross-model mean normalized importance of 6.7%, indicating stable contributions of metabolic and hematologic features to model performance (Figure 3).

Figure 3 Overview of permutation importance across models. (A) Global heatmap of permutation importance for 29 features across the seven ML models. Color intensity represents the magnitude of permutation importance for each feature in different models; darker colors indicate a greater decline in model performance after feature permutation, suggesting a larger contribution to prediction. (B) Bubble plot of the top 10 features ranked by permutation importance across the models. Both bubble size and color indicate the relative importance of each feature in different models, highlighting the consistency and variability of important features across models. (C) Summary plot of permutation importance for the top 10 features across models. Dots indicate mean values, thick lines indicate mean ± standard deviation, and thin lines indicate the minimum-to-maximum range. Overall, nodule type and age were the dominant variables, while glucose, lobulation, pleural indentation, RDW-SD, QFT, uric acid, calcification, and E_abs also showed additional contributions, suggesting consistent discriminative value for differentiating lung cancer from pulmonary tuberculosis. ALP, alkaline phosphatase; APTT, activated partial thromboplastin time; AST, aspartate aminotransferase; B%, basophil percentage; BUN, blood urea nitrogen; Cl, chloride; Cre, creatinine; CTC, circulating tumor cell; ET, extra trees; GBDT, gradient boosting decision tree; HbO2, oxyhemoglobin; K, potassium; KNN, k-nearest neighbors; L_abs, absolute lymphocyte count; LR, logistic regression; ML, machine learning; MLP, multilayer perceptron; PCV, packed cell volume; P-LCR, platelet large cell ratio; PaCO2, partial pressure of carbon dioxide; QFT, QuantiFERON-TB; RBF-SVM, radial basis function support vector machine; RDW-SD, red cell distribution width-standard deviation; Reduced Hb, reduced hemoglobin; RF, random forest; THB, total hemoglobin measured by blood gas analysis; TT, thrombin time.

Beyond these core variables, lobulation (3.4%), pleural indentation (2.4%), RDW-SD (2.4%), QFT (2.2%), uric acid (2.0%), calcification (1.9%), and E_abs (1.9%) were also among the top 10 features across models. The heatmap further showed that nodule type and age displayed relatively consistent importance patterns across different models, while the importance of other variables exhibited some degree of model dependence. These findings suggest that the differentiation of SPNs does not rely solely on a single imaging feature, but also involves the joint contribution of multidimensional biological information, including local morphology, host background, metabolic status, hematologic status, and tuberculosis-related immunological information.

SHAP interpretability analysis and cross-model consistency

To further explain the prediction mechanisms of different models, SHAP analysis was performed to assess the contributions of individual features to model outputs. The results showed that SHAP and permutation importance were consistent in identifying nodule type and age as the dominant variables, although the rankings of subsequent variables differed between the two methods. Nodule type remained the most influential variable, with a cross-model mean normalized mean |SHAP| value of 24.1%, followed by age at 12.1%. The subsequent features were THB (7.0%), RDW-SD (5.0%), lobulation (4.5%), PCV (4.4%), TT (3.7%), glucose (3.6%), pleural indentation (3.3%), and platelet large cell ratio (P-LCR; 3.3%) (Figure 4). These results indicate that, beyond the dominant effects of nodule morphology and age, hematologic and coagulation-related variables also contributed to model decision-making.

Figure 4 Overview of SHAP importance across models. (A) Global heatmap of SHAP importance for 29 features across the seven ML models. Color intensity represents the mean absolute SHAP value of each feature in the different models, reflecting its overall contribution to model output. (B) Bubble plot of the top 10 features ranked by mean absolute SHAP values across the models. Both bubble size and color indicate the relative importance of each feature in the different models. (C) Summary plot of SHAP importance for the top 10 features across the models. Dots indicate mean values, thick lines indicate mean ± standard deviation, and thin lines indicate the minimum-to-maximum range. Overall, nodule type and age showed the most stable explanatory contributions, while THB, RDW-SD, lobulation, PCV, TT, glucose, pleural indentation, and P-LCR also contributed to model outputs, suggesting relatively stable discriminative value for differentiating lung cancer from pulmonary tuberculosis. ALP, alkaline phosphatase; APTT, activated partial thromboplastin time; AST, aspartate aminotransferase; B%, basophil percentage; BUN, blood urea nitrogen; Cl, chloride; Cre, creatinine; CTC, circulating tumor cell; ET, extra trees; GBDT, gradient boosting decision tree; HbO2, oxyhemoglobin; K, potassium; KNN, k-nearest neighbors; L_abs, absolute lymphocyte count; LR, logistic regression; ML, machine learning; MLP, multilayer perceptron; PCV, packed cell volume; P-LCR, platelet large cell ratio; PaCO2, partial pressure of carbon dioxide; QFT, QuantiFERON-TB; RBF-SVM, radial basis function support vector machine; RDW-SD, red cell distribution width-standard deviation; Reduced Hb, reduced hemoglobin; RF, random forest; SHAP, SHapley Additive exPlanations; THB, total hemoglobin measured by blood gas analysis; TT, thrombin time.

Taken together, the SHAP and permutation importance analyses consistently identified nodule type and age as the two most stable core features. However, the rankings of secondary variables differed between the two interpretability methods, suggesting that some laboratory contributions were method-dependent and model-dependent. Permutation importance highlighted glucose, lobulation, pleural indentation, QFT, uric acid, and calcification, whereas SHAP emphasized THB, RDW-SD, PCV, TT, and P-LCR. This discrepancy suggests that the dominant discriminative signals were stable, while secondary variables should be interpreted cautiously. Overall, these interpretability analyses support the central finding of this study: the differentiation of lung cancer and pulmonary tuberculosis in patients with SPNs is a multidimensional process jointly driven by nodule morphology, age-related background, and systemic biological status.

Parsimonious-model sensitivity analysis using the top 10 LASSO-ranked predictors

In the parsimonious-model sensitivity analysis, the top 10 LASSO-ranked predictors were nodule type, age, FDP, THB, E_abs, APTT, lobulation, TT, RDW-SD, and pleural indentation. The reduced-feature models achieved validation AUCs ranging from 0.746 to 0.789 (Table 3). The best-performing reduced-feature model was ET, with a validation AUC of 0.789, sensitivity of 0.686, specificity of 0.722, and balanced accuracy of 0.704. KNN achieved the highest balanced accuracy of 0.723. This reduced-feature analysis increased the event-per-feature ratio to 16.8 in the overall cohort and 11.7 in the training set. These findings suggest that the main discriminative signal was not entirely dependent on the full 29-predictor feature set.

Table 3

Top 10 LASSO reduced-feature sensitivity analysis

Model Training set Validation set
AUC Sensitivity Specificity Accuracy Balanced Acc AUC Sensitivity Specificity Accuracy Balanced Acc
LR 0.825 0.709 0.783 0.754 0.746 0.759 0.647 0.759 0.715 0.703
RF 0.843 0.786 0.734 0.754 0.760 0.776 0.706 0.696 0.700 0.701
ET 0.823 0.735 0.750 0.744 0.743 0.789 0.686 0.722 0.708 0.704
RBF-SVM 0.821 0.735 0.761 0.751 0.748 0.765 0.686 0.709 0.700 0.698
KNN 0.813 0.701 0.750 0.731 0.725 0.746 0.686 0.759 0.731 0.723
MLP 0.834 0.564 0.886 0.761 0.725 0.754 0.490 0.861 0.715 0.675
GBDT 0.828 0.778 0.717 0.741 0.748 0.780 0.667 0.696 0.685 0.681

Top 10 LASSO-ranked predictors were nodule type, age, FDP, THB, E_abs, APTT, lobulation, TT, RDW-SD, and pleural indentation. APTT, activated partial thromboplastin time; AUC, area under the curve; Balanced Acc, balanced accuracy; E_abs, absolute eosinophil count; ET, extra trees; FDP, fibrin degradation products; GBDT, gradient boosting decision tree; KNN, k-nearest neighbors; LASSO, least absolute shrinkage and selection operator; LR, logistic regression; MLP, multilayer perceptron; RBF-SVM, radial basis function support vector machine; RDW-SD, red cell distribution width-standard deviation; RF, random forest; THB, total hemoglobin measured by blood gas analysis; TT, thrombin time.

Sensitivity analysis using complete-case data

In the complete-case no-imputation sensitivity analysis, complete cases were defined as patients with non-missing values for the 29 predictors selected in the primary analysis. After excluding patients with missing values in any of these predictors, 268 of 301 patients in the training set and 122 of 130 patients in the internal hold-out validation set were retained. The complete-case training set included 158 patients with pulmonary tuberculosis and 110 patients with lung cancer, whereas the validation set included 73 patients with pulmonary tuberculosis and 49 patients with lung cancer.

The same 29 predictors were used for model training and evaluation. Validation AUCs ranged from 0.722 to 0.814 across the seven models (Table 4). ET achieved the highest validation AUC of 0.814, with a sensitivity of 0.755, specificity of 0.767, and balanced accuracy of 0.761. The RBF-SVM model maintained a validation AUC of 0.804. Overall, these results suggest that the main discrimination findings were not primarily driven by missing-value imputation.

Table 4

Complete-case sensitivity analysis of diagnostic performance without missing-value imputation

Model Training set Validation set
AUC Sensitivity Specificity Accuracy Balanced Acc AUC Sensitivity Specificity Accuracy Balanced Acc
LR 0.884 0.800 0.816 0.810 0.808 0.787 0.714 0.726 0.721 0.720
RF 0.968 0.891 0.911 0.903 0.901 0.772 0.673 0.767 0.730 0.720
ET 0.928 0.891 0.804 0.840 0.847 0.814 0.755 0.767 0.762 0.761
RBF-SVM 0.905 0.800 0.835 0.821 0.818 0.804 0.653 0.795 0.738 0.724
KNN 0.832 0.727 0.709 0.716 0.718 0.757 0.673 0.795 0.746 0.734
MLP 0.859 0.782 0.804 0.795 0.793 0.722 0.633 0.699 0.672 0.666
GBDT 0.930 0.864 0.804 0.828 0.834 0.771 0.694 0.740 0.721 0.717

AUC, area under the curve; Balanced Acc, balanced accuracy; ET, extra trees; GBDT, gradient boosting decision tree; KNN, k-nearest neighbors; LR, logistic regression; MLP, multilayer perceptron; RBF-SVM, radial basis function support vector machine; RF, random forest.


Discussion

The main finding of this study is that an integrated analytical framework based on pretreatment baseline clinical information, manually assessed imaging features, and multidimensional routine laboratory parameters in patients with SPNs can, to a certain extent, effectively differentiate lung cancer from pulmonary tuberculosis. More importantly, our findings suggest that the distinction between lung cancer and pulmonary tuberculosis does not rely on any single isolated indicator, but rather on the combined contribution of local nodule morphology, host age-related background, and systemic state-related signals. In recent years, studies on pulmonary nodule biomarkers and multimodal predictive models have repeatedly emphasized that a single test is often insufficient to meet the demands of clinical decision-making for complex nodules, whereas integration of multi-source information is more likely to improve classification of indeterminate pulmonary nodules (7,29-31).

The significance of this study lies first in its focus on the specific clinical scenario of differentiating lung cancer from pulmonary tuberculosis, a highly challenging condition with a distinct epidemiological background, rather than on a generic binary classification of benign versus malignant pulmonary nodules. In regions with a high tuberculosis burden, pulmonary tuberculosis presenting as non-calcified solitary solid nodules often mimics lung cancer on imaging, posing substantial diagnostic challenges. Previous studies directly comparing pulmonary tuberculosis and lung cancer have shown that although lesion margin characteristics, calcification, and overall morphology may provide useful insights, CT-based empirical assessment alone is far from sufficient for resolving all cases independently (32-34). Therefore, the present study did not merely add another ML result; rather, it attempted to provide an integrative diagnostic basis for this longstanding clinical dilemma using data that are readily available in routine clinical practice.

In terms of feature composition, the top-ranked variables can be broadly categorized into three groups. The first group comprises local imaging morphology, represented mainly by nodule type, lobulation, pleural indentation, and calcification. The second group comprises host background information, represented primarily by age. The third group includes systemic biological indicators, including glucose, THB, RDW-SD, PCV, TT, P-LCR, QFT, uric acid, and E_abs. This feature architecture itself helps explain why the present study achieved relatively stable discriminative performance: the models were not merely learning what the nodules look like, nor simply whether a specific blood biomarker is abnormal; rather, the models were simultaneously leveraging complementary information derived from local lesion phenotype and systemic biological status (35-38).

Among these variables, the high ranking of nodule type, lobulation, and calcification is clinically plausible. Nodule type essentially reflects lesion density, internal homogeneity, and growth pattern, and is one of the most fundamental and important imaging dimensions in pulmonary nodule risk assessment (4,39). Lobulation generally suggests heterogeneous growth rates across different regions of the lesion, and is often associated with tumor heterogeneity and local invasive behavior. However, spiculation was not among the most stable top-ranked variables in the current cross-model analysis, suggesting that its contribution may be partly captured by other morphology-related features (40,41). Conversely, calcification often serves as an important indicator supporting a benign or granulomatous lesion. High-quality management guidelines and disease-specific studies have consistently shown that morphological signs remain central to pulmonary nodule evaluation, and the difficulty in differentiating pulmonary tuberculosis from lung cancer lies precisely in the overlap of some morphological features between these two entities, despite differences in their overall combinational patterns and relative weights (2,19,42-45).

Age was consistently ranked among the top features in this study. Systematic reviews of pulmonary nodule risk prediction have shown that age is one of the most common independent variables included in malignancy probability models. Although age does not directly describe lesion morphology, it captures baseline cancer risk and cumulative exposure history at the host level. In the specific context of lung cancer versus pulmonary tuberculosis, the importance of age suggests that the models identified not only imaging differences, but also differences in baseline disease probability (46-48).

QFT provided an important tuberculosis-related immunological indicator for differentiating lung cancer from pulmonary tuberculosis. As an interferon-gamma release assay (IGRA), a positive QFT result is more suggestive of a tuberculosis-related background. Kobashi et al. reported that QFT-2G showed a positivity rate of 79% in patients with tuberculosis and a false-positive rate of only 5% in patients with non-tuberculous pulmonary diseases, indicating relatively good direct discriminative value. However, meta-analyses have also shown that IGRAs have limitations in the diagnosis of active tuberculosis and cannot be used independently without consideration of the clinical and imaging context. In the present study, QFT was among the top-ranked variables in permutation importance, supporting its complementary value as a tuberculosis-related immunological marker. However, it was less prominent in the SHAP ranking, suggesting that its contribution may be method-dependent and should not be interpreted as a standalone diagnostic determinant (49,50).

The prominence of laboratory variables further indicates that the model also captured information related to systemic metabolism, inflammatory response, and host condition. Among these variables, the importance of glucose is better interpreted as reflecting metabolic differences between lung cancer and pulmonary tuberculosis, rather than serving as an independent diagnostic marker. Previous studies have shown reproducible separation patterns between lung cancer and pulmonary tuberculosis at the level of serum metabolomics. Chen et al. developed a serum metabolomics-based model using 694 participants to distinguish lung cancer from pulmonary tuberculosis, and reported an AUC of 0.89 for the core metabolite markers in the validation cohort, suggesting that metabolism-related information itself has substantial discriminative potential (51). In addition, a study by Zheng et al. on non-small cell lung cancer and tuberculosis further showed that combining glucose metabolism–related and tumor burden indicators with cell-free DNA improved discrimination compared with either signal source alone, supporting the practical value of glucose metabolism–related information in differentiating the two conditions (52).

The contribution of hematologic and coagulation-related variables should be interpreted at the level of systemic biological status rather than as evidence that a single blood marker independently distinguishes lung cancer from pulmonary tuberculosis. Earlier discriminant analysis studies incorporated Hb into the laboratory framework for differentiating lung cancer from pulmonary tuberculosis, and showed that laboratory parameters including Hb had practical classification value, with an overall correct classification probability of approximately 72.6–78.0% (53). Meanwhile, more recent studies have suggested that anemia in patients with pulmonary tuberculosis is associated with greater disease severity, more pronounced wasting status, and slower recovery during treatment (54). In the current analysis, variables such as THB, RDW-SD, PCV, TT, and P-LCR showed explanatory contributions, suggesting that differences in oxygen-carrying capacity, red cell distribution, coagulation status, and platelet-related indices may provide complementary information.

Another aspect of this study that deserves emphasis is that feature importance and interpretability analyses were considered as important as discriminative performance itself. For medical ML models, reporting only the AUC, sensitivity, and specificity is insufficient to explain what the model has actually learned. Conversely, identifying relatively stable key features under different interpretability frameworks not only enhances model transparency, but may also improve clinical acceptability. Methodological studies on explainable artificial intelligence in healthcare generally agree that interpretability is essential for building clinical trust, facilitating deployment, and helping researchers move from a mere predictive tool toward an understandable decision-support tool (55,56).

Misclassification should be carefully considered when interpreting the clinical implications of the model. A false-negative prediction for lung cancer may delay further diagnostic work-up, surgical evaluation, or oncological treatment, whereas a false-positive prediction for lung cancer may lead to unnecessary invasive procedures, additional imaging or biopsy, and psychological burden. Conversely, misclassifying pulmonary tuberculosis as lung cancer may delay timely anti-tuberculosis therapy, whereas misclassifying lung cancer as tuberculosis may postpone cancer diagnosis and treatment. Therefore, the model is not intended to serve as a stand-alone screening, rule-in, or rule-out tool. Its appropriate role is to provide supportive risk stratification in combination with CT assessment, laboratory testing, microbiological evidence, pathological confirmation when feasible, and multidisciplinary clinical judgment.

Several limitations should be acknowledged. First, this was a single-center retrospective study, and only internal hold-out validation was performed. The validation set was generated by random stratified partitioning of the same institutional cohort; therefore, it should not be interpreted as evidence of external generalizability. Multicenter external validation is required to assess transportability across different institutions, scanners, patient populations, and tuberculosis-prevalence settings. Second, the number of positive lung cancer events was limited relative to the number of selected predictors. Although LASSO feature selection, cross-validation-based hyperparameter tuning, training-set-only threshold optimization, OOF prediction, internal hold-out validation, and complete-case no-imputation sensitivity analysis were used to reduce optimism and evaluate robustness, residual overfitting cannot be excluded. Third, no prespecified clinical rule-in or rule-out threshold was defined, and the observed sensitivity and specificity should be interpreted as exploratory performance measures rather than clinically validated decision thresholds. Future studies based on multicenter external validation should further incorporate calibration analysis, clinical utility assessment, and model simplification to improve clinical feasibility and translational potential (27).

Overall, the study findings support the potential value of routinely available multimodal clinical information in the differential diagnosis of lung cancer and pulmonary tuberculosis in patients with SPNs.


Conclusions

ML models that integrate baseline clinical information, manually assessed imaging features, and routine laboratory parameters demonstrated potential value for differentiating lung cancer from pulmonary tuberculosis in patients with SPNs. Nodule type, age, and several key variables showed relatively stable importance across models.


Acknowledgments

None.


Footnote

Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://tlcr.amegroups.com/article/view/10.21037/tlcr-2026-0559/rc

Data Sharing Statement: Available at https://tlcr.amegroups.com/article/view/10.21037/tlcr-2026-0559/dss

Peer Review File: Available at https://tlcr.amegroups.com/article/view/10.21037/tlcr-2026-0559/prf

Funding: This study was supported by the National Natural Science Foundation of China (grant No. 52271248), and the Clinical Research Project of the Health Industry of the Shanghai Municipal Health Commission (grant No. 202340244).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://tlcr.amegroups.com/article/view/10.21037/tlcr-2026-0559/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study protocol was approved by the Ethics Committee of Shanghai Pulmonary Hospital (No. K23-233Y), which waived the requirement for informed consent due to the retrospective design of the study. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Jonas DE, Reuland DS, Reddy SM, et al. Screening for Lung Cancer With Low-Dose Computed Tomography: Updated Evidence Report and Systematic Review for the US Preventive Services Task Force. JAMA 2021;325:971-87. [Crossref] [PubMed]
  2. Zhong D, Sidorenkov G, Jacobs C, et al. Lung Nodule Management in Low-Dose CT Screening for Lung Cancer: Lessons from the NELSON Trial. Radiology 2024;313:e240535. [Crossref] [PubMed]
  3. Asija A, Manickam R, Aronow WS, et al. Pulmonary nodule: a comprehensive review and update. Hosp Pract (1995) 2014;42:7-16. [Crossref] [PubMed]
  4. Gould MK, Donington J, Lynch WR, et al. Evaluation of individuals with pulmonary nodules: when is it lung cancer? Diagnosis and management of lung cancer, 3rd ed: American College of Chest Physicians evidence-based clinical practice guidelines. Chest 2013;143:e93S-e120S.
  5. MacMahon H, Naidich DP, Goo JM, et al. Guidelines for Management of Incidental Pulmonary Nodules Detected on CT Images: From the Fleischner Society 2017. Radiology 2017;284:228-43. [Crossref] [PubMed]
  6. Mazzone PJ, Lam L. Evaluating the Patient With a Pulmonary Nodule: A Review. JAMA 2022;327:264-73. [Crossref] [PubMed]
  7. Paez R, Kammer MN, Tanner NT, et al. Update on Biomarkers for the Stratification of Indeterminate Pulmonary Nodules. Chest 2023;164:1028-41. [Crossref] [PubMed]
  8. Zheng Z, Pan Y, Guo F, et al. Multimodality FDG PET/CT appearance of pulmonary tuberculoma mimicking lung cancer and pathologic correlation in a tuberculosis-endemic country. South Med J 2011;104:440-5. [Crossref] [PubMed]
  9. Lang S, Sun J, Wang X, et al. Asymptomatic pulmonary tuberculosis mimicking lung cancer on imaging: A retrospective study. Exp Ther Med 2017;14:2180-8. [Crossref] [PubMed]
  10. Yang L, Jiang Z, Tong J, et al. Development and validation of a preoperative CT‑based radiomics nomogram to differentiate tuberculosis granulomas from lung adenocarcinomas: an external validation study. BMC Cancer 2024;24:670. [Crossref] [PubMed]
  11. Sathekge MM, Maes A, Pottel H, et al. Dual time-point FDG PET-CT for differentiating benign from malignant solitary pulmonary nodules in a TB endemic area. S Afr Med J 2010;100:598-601. [Crossref] [PubMed]
  12. Expert Panel on Thoracic Imaging. ACR Appropriateness Criteria® Incidentally Detected Indeterminate Pulmonary Nodule. J Am Coll Radiol 2023;20:S455-70.
  13. Nair A, Bartlett EC, Walsh SLF, et al. Variable radiological lung nodule evaluation leads to divergent management recommendations. Eur Respir J 2018;52:1801359. [Crossref] [PubMed]
  14. Wang D, Li P, Fei X, et al. A combined diagnostic model based on circulating tumor cell in patients with solitary pulmonary nodules. J Gene Med 2023;25:e3529. [Crossref] [PubMed]
  15. He J, Wang B, Tao J, et al. Accurate classification of pulmonary nodules by a combined model of clinical, imaging, and cell-free DNA methylation biomarkers: a model development and external validation study. Lancet Digit Health 2023;5:e647-56. [Crossref] [PubMed]
  16. Sethi S, Cicenia J. Role of biomarkers in lung nodule evaluation. Curr Opin Pulm Med 2022;28:275-81. [Crossref] [PubMed]
  17. Kammer MN, Lakhani DA, Balar AB, et al. Integrated Biomarkers for the Management of Indeterminate Pulmonary Nodules. Am J Respir Crit Care Med 2021;204:1306-16. [Crossref] [PubMed]
  18. Sun W, Zhang L, Liang J, et al. Comparison of clinical and imaging features between pulmonary tuberculosis complicated with lung cancer and simple pulmonary tuberculosis: a systematic review and meta-analysis. Epidemiol Infect 2022;150:e43. [Crossref] [PubMed]
  19. Gao X, Tan H, Zhu M, et al. Construction and validation of a clinical differentiation model between peripheral lung cancer and solitary pulmonary tuberculosis. Lung Cancer 2024;193:107851. [Crossref] [PubMed]
  20. Feng B, Chen X, Chen Y, et al. Radiomics nomogram for preoperative differentiation of lung tuberculoma from adenocarcinoma in solitary pulmonary solid nodule. Eur J Radiol 2020;128:109022. [Crossref] [PubMed]
  21. Dong Q, Wen Q, Li N, et al. Radiomics combined with clinical features in distinguishing non-calcifying tuberculosis granuloma and lung adenocarcinoma in small pulmonary nodules. PeerJ 2022;10:e14127. [Crossref] [PubMed]
  22. Feng B, Chen X, Chen Y, et al. Solitary solid pulmonary nodules: a CT-based deep learning nomogram helps differentiate tuberculosis granulomas from lung adenocarcinomas. Eur Radiol 2020;30:6497-507. [Crossref] [PubMed]
  23. de Margerie-Mellon C, Chassagnon G. Artificial intelligence: A critical review of applications for lung nodule and lung cancer. Diagn Interv Imaging 2023;104:11-7. [Crossref] [PubMed]
  24. Wu Z, Wang F, Cao W, et al. Lung cancer risk prediction models based on pulmonary nodules: A systematic review. Thorac Cancer 2022;13:664-77. [Crossref] [PubMed]
  25. Ost DE. Artificial intelligence applications for the diagnosis of pulmonary nodules. Curr Opin Pulm Med 2025;31:344-51. [Crossref] [PubMed]
  26. Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024;385:e078378. [Crossref] [PubMed]
  27. Collins GS, Dhiman P, Ma J, et al. Evaluation of clinical prediction models (part 1): from development to external validation. BMJ 2024;384:e074819. [Crossref] [PubMed]
  28. Collins GS, Reitsma JB, Altman DG, et al. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD statement. BMJ 2015;350:g7594. [Crossref] [PubMed]
  29. Liang W, Tao J, Cheng C, et al. A clinically effective model based on cell-free DNA methylation and low-dose CT for risk stratification of pulmonary nodules. Cell Rep Med 2024;5:101750. [Crossref] [PubMed]
  30. Yang M, Yu H, Feng H, et al. Enhancing the differential diagnosis of small pulmonary nodules: a comprehensive model integrating plasma methylation, protein biomarkers, and LDCT imaging features. J Transl Med 2024;22:984. [Crossref] [PubMed]
  31. Li X, Zhang Q, Jin X, et al. Combining serum miRNAs, CEA, and CYFRA21-1 with imaging and clinical features to distinguish benign and malignant pulmonary nodules: a pilot study : Xianfeng Li et al.: Combining biomarker, imaging, and clinical features to distinguish pulmonary nodules. World J Surg Oncol 2017;15:107. [Crossref] [PubMed]
  32. Liu Z, Ran H, Yu X, et al. Immunocyte count combined with CT features for distinguishing pulmonary tuberculoma from malignancy among non-calcified solitary pulmonary solid nodules. J Thorac Dis 2023;15:386-98. [Crossref] [PubMed]
  33. Zhang J, Han T, Ren J, et al. Discriminating Small-Sized (2 cm or Less), Noncalcified, Solitary Pulmonary Tuberculoma and Solid Lung Adenocarcinoma in Tuberculosis-Endemic Areas. Diagnostics (Basel) 2021;11:930. [Crossref] [PubMed]
  34. Patnam N, Trivedi K, Janu A, et al. Cross-sectional imaging review of common to uncommon lung cancer mimickers in a tertiary care oncology center. Acta Radiol 2023;64:2731-47. [Crossref] [PubMed]
  35. Gould MK, Ananth L, Barnett PG, et al. A clinical model to estimate the pretest probability of lung cancer in patients with solitary pulmonary nodules. Chest 2007;131:383-8. [Crossref] [PubMed]
  36. Liu Y, Balagurunathan Y, Atwater T, et al. Radiological Image Traits Predictive of Cancer Status in Pulmonary Nodules. Clin Cancer Res 2017;23:1442-9. [Crossref] [PubMed]
  37. Sim YT, Poon FW. Imaging of solitary pulmonary nodule-a clinical review. Quant Imaging Med Surg 2013;3:316-26. [Crossref] [PubMed]
  38. Langan RC, Goodbred AJ. Pulmonary Nodules: Common Questions and Answers. Am Fam Physician 2023;107:282-91.
  39. Callister ME, Baldwin DR, Akram AR, et al. British Thoracic Society guidelines for the investigation and management of pulmonary nodules. Thorax 2015;70:ii1-ii54. [Crossref] [PubMed]
  40. Erasmus JJ, Connolly JE, McAdams HP, et al. Solitary pulmonary nodules: Part I. Morphologic evaluation for differentiation of benign and malignant lesions. Radiographics 2000;20:43-58.
  41. Snoeckx A, Reyntiens P, Desbuquoit D, et al. Evaluation of the solitary pulmonary nodule: size matters, but do not ignore the power of morphology. Insights Imaging 2018;9:73-86. [Crossref] [PubMed]
  42. Brims F, McWilliams A, Williamson J, et al. The TSANZ Practical Guide for Clinicians in the Management of Screen- and Incidentally-Detected Nodules. Respirology 2025;30:558-73. [Crossref] [PubMed]
  43. Xu HB, Lv FJ, Ding C, et al. Exploring and verifying key thin-section computed tomography features for accurately differentiating granulomas and peripheral lung cancers. J Thorac Dis 2025;17:2827-40. [Crossref] [PubMed]
  44. Xu HB, Ding C, Zhao M, et al. Exploring the key clinical and CT characteristics of granulomas mimicking peripheral lung cancers: a case-control study. Insights Imaging 2025;16:157. [Crossref] [PubMed]
  45. Zhuo Y, Zhan Y, Zhang Z, et al. Clinical and CT Radiomics Nomogram for Preoperative Differentiation of Pulmonary Adenocarcinoma From Tuberculoma in Solitary Solid Nodule. Front Oncol 2021;11:701598. [Crossref] [PubMed]
  46. Senent-Valero M, Librero J, Pastor-Valero M. Solitary pulmonary nodule malignancy predictive models applicable to routine clinical practice: a systematic review. Syst Rev 2021;10:308. [Crossref] [PubMed]
  47. Chung K, Mets OM, Gerke PK, et al. Brock malignancy risk calculator for pulmonary nodules: validation outside a lung cancer screening population. Thorax 2018;73:857-63. [Crossref] [PubMed]
  48. Al-Ameri A, Malhotra P, Thygesen H, et al. Risk of malignancy in pulmonary nodules: A validation study of four prediction models. Lung Cancer 2015;89:27-30. [Crossref] [PubMed]
  49. Kobashi Y, Mouri K, Yagi S, et al. Usefulness of the QuantiFERON TB-2G test for the differential diagnosis of pulmonary tuberculosis. Intern Med 2008;47:237-43. [Crossref] [PubMed]
  50. Lu P, Chen X, Zhu LM, Yang HT. Interferon-Gamma Release Assays for the Diagnosis of Tuberculosis: A Systematic Review and Meta-analysis. Lung 2016;194:447-58. [Crossref] [PubMed]
  51. Chen S, Li C, Qin Z, et al. Serum Metabolomic Profiles for Distinguishing Lung Cancer From Pulmonary Tuberculosis: Identification of Rapid and Noninvasive Biomarker. J Infect Dis 2023;228:1154-65. [Crossref] [PubMed]
  52. Zheng W, Quan B, Gao G, et al. Combination of Circulating Cell-Free DNA and Positron Emission Tomography to Distinguish Non-Small Cell Lung Cancer from Tuberculosis. Lab Med 2023;54:130-41. [Crossref] [PubMed]
  53. Agapova RK, Bogadel’nikova IV, Sergeev AS, et al. Discriminant analysis of clinical laboratory data for diagnosis and treatment of tuberculosis and lung cancer. Vestn Ross Akad Med Nauk 1999;47-51.
  54. Ashenafi S, Bekele A, Aseffa G, et al. Anemia Is a Strong Predictor of Wasting, Disease Severity, and Progression, in Clinical Tuberculosis (TB). Nutrients 2022;14:3318. [Crossref] [PubMed]
  55. Loh HW, Ooi CP, Seoni S, et al. Application of explainable artificial intelligence for healthcare: A systematic review of the last decade (2011-2022). Comput Methods Programs Biomed 2022;226:107161. [Crossref] [PubMed]
  56. Ali S, Akhlaq F, Imran AS, et al. The enlightening role of explainable artificial intelligence in medical & healthcare domains: A systematic literature review. Comput Biol Med 2023;166:107555. [Crossref] [PubMed]
Cite this article as: Li Y, Wu W, Xu H, Ma J, Bao M, Zhu J, Chen S. Development and validation of machine learning diagnostic models integrating clinical, CT, and laboratory features to differentiate lung cancer from pulmonary tuberculosis in patients with solitary pulmonary nodules: a single-center retrospective study. Transl Lung Cancer Res 2026;15(7):209. doi: 10.21037/tlcr-2026-0559

Download Citation