Lung cancer survivorship: natural language processing for automated abstraction of follow-up computed tomography indication
Editorial Commentary

Lung cancer survivorship: natural language processing for automated abstraction of follow-up computed tomography indication

Brane Grambozov, Elvis Ruznic, Franz Zehentmayr ORCID logo

Department of Radiation Oncology, Paracelsus Medical University, Salzburg, Austria

Correspondence to: Franz Zehentmayr, MD. Department of Radiation Oncology, Paracelsus Medical University, Müllner Hauptstrasse 48, Salzburg A-5020, Austria. Email: f.zehentmayr@salk.at.

Comment on: Khan A, Choi E, Su C, et al. Automatic Abstraction of Computed Tomography Imaging Indication Using Natural Language Processing for Evaluation of Surveillance Patterns in Long-Term Lung Cancer Survivors. JCO Clin Cancer Inform 2025;9:e2400279.


Keywords: Cancer survivorship; lung cancer; natural language processing (NLP); overall survival (OS)


Submitted Nov 24, 2025. Accepted for publication Dec 04, 2025. Published online Jan 16, 2026.

doi: 10.21037/tlcr-2025-1-1344


Lung cancer is still the leading cause of cancer deaths worldwide (1). Structured follow-up regimens are paramount to secure long-term tumor control. In this respect, guidelines usually refer to standardized oncological follow-up procedures for the first 5 years after diagnosis but no evidence-based follow-up exists thereafter. This is somewhat astonishing since the risk of recurrence is 2–14% per patient year and the risk of second primary lung cancer is 1–4% (2) or—as stated in the current study—even 8.36% within 10 years (3).

In this context, one must keep in mind that the problem of long-term lung cancer survivorship is relatively new. In dependence on stage, the 5-year overall survival for lung cancer Union for International Cancer Control (UICC) stages I–III varies between 12% and 90% (4) and for UICC IV, it is below 10%. With the widespread clinical use of immunotherapy after the publication of the PACIFIC trial (5-9), the rate for locally advanced stages UICC IIIa-c has risen from 20% (10) to approximately 40–50% during the past 5 years (11). The 585 patients analyzed in the current paper, however, were treated between 2000 and 2017, i.e., in the pre-immunotherapy era, with more than half of the patients having UICC I lung cancer. Although the majority of these individuals (more than 90%) will never—even in the long run—experience a second lung primary after treatment (8.36%), the design of stringent follow-up schedules for these patients based on robust evidence is laudable. As the authors point out in the introduction, at present, it is often impossible to judge from radiology reports whether follow-up computer tomography (CT) scans were performed for surveillance or other purposes. To find this out by manually reviewing the patient charts in retrospect is time-consuming work for which most institutions lack resources. Hence, the aim of the current paper was to design a natural language processing (NLP)-based model for CT indication abstraction and to test the clinical effectiveness of the model by predicting overall survival (OS). Of the initial cohort of more than 7,000 patients, 585 were eligible for analysis. The primary outcome was a statistical measure, i.e., the area under the curve (AUC), to predict whether the model correctly classified a CT scan indicated for surveillance or “other reasons”. The second—and from the clinical point of view—more important outcome measure was OS in correlation with the receipt of surveillance CT after 5 years. The hybrid model composed of the NLP features derived from the radiology reports combined with the variables from the structured electronic health record (EHR) outperformed the models simply based on either approach with respect to automated abstraction of CT indications [AUC 0.86, 95% confidence interval (CI): 0.82–0.90]. Three-quarters of the patients (n=438) had at least one CT for surveillance after 5 years, whereas 147 patients had a CT for other reasons. Among the latter group were primarily patients with severe lung conditions that require immediate intervention. As for the secondary endpoint, OS in correlation with CT indications, the mere comparison between patients with and without CT did not yield a difference. However, a significant difference [P=0.02; hazard ratio (HR) 0.60; 95% CI: 0.41–0.89] could be detected when only those patients who had a surveillance CT scan were compared to those without. From this, the authors conclude that surveillance CT beyond 5 years for long-term lung cancer survivors is clinically useful. This type of follow-up CTs can be reliably indicated from real-world data sets using the proposed hybrid NLP model, which could result in standardized traceable follow-up and finally better long-term clinical outcome (12).

Thus far, the timing of surveillance imaging after 5 years of oncological follow-up depends on findings in retrospective series based on cumbersome manual chart review (2). Generally, automated approaches can analyze large datasets with great efficiency, which enhances robustness but lacks clinical granularity (2). To balance these requirements, the best method seems to be a combination of manual abstraction and NLP (2). The interest of the scientific community in NLP research increases constantly, with a doubling of the number of related publications between 2016 [reviewed by Pons 2016 (13)] and 2021 [reviewed by Casey (14)]. NLP can extract structured information from free-text radiology reports. Subsequently, it breaks down these narratives to the most essential semantic information, which may then be used to train a new model (12,14). In the current study, the level of granularity is ensured by a combination of first-step manual annotation under the supervision of a medical oncologist that serves as the basis for further NLP, which includes correlating the extracted terms with quantitative measures [mean + standard deviation (SD), min, max, median, interquartile difference]. This—like in all automated approaches—reduces workload with the idea of informing clinical decision processes. NLP covers five areas of application: diagnostic surveillance, cohort building for epidemiological studies, query-based case retrieval, quality assessment of radiological practice, and clinical support services (13). Diagnostic surveillance, which is the focus of the current study, was investigated using thoracic CT scans. This is important since both the disease entity of interest and each type of imaging modality come along with a domain specific language, which—if NLP approaches should be successful in the medical context—must be specifically accounted for in order to render well-performing solutions. NLP models, especially deep neural network-based large language models (LLMs), have the advantage of being more specifically fine-tuned to the task that they should fulfill. This means that by abstracting problem-specific key vocabulary (in this case: indication for surveillance CT imaging), the model will yield precise results in terms of this specific problem. However, models that are too specialized, i.e., models that are trained with the “vocabulary” or “features” relevant for this very specific problem only, may deliver misleading suggestions if used in a different domain.

The above-mentioned issue of task-specificity sheds light on the following aspects: key phrases, heterogeneity of the studied patient population and OS as an endpoint. which should be critically appraised. First, the key phrases identified by the proposed NLP model to indicate a follow-up CT scan after 5 years and beyond are “surveillance” and “other”, the latter of which includes a variety of acute severe health conditions. From a clinical perspective, the stratification criterion shows that patients receiving a CT scan tagged as “surveillance” are in better general condition with lower tumor burden than those receiving a CT scan for “other” reasons. It is obvious that the first group of patients is more likely to achieve better clinical outcomes in terms of survival than the second. Hence, it seems that the CT scan indications “surveillance” or “other” are surrogates for OS probability since the signal words are rather connected to oncological conditions (Tab. S1: metastasis, recurrence, stable, nodule, pleural effusion, systemic treatment) rather than mere radiological indications. Secondly, a major problem is the heterogeneity of the patient population. While the authors refer to “early lung cancer” (first line of the “introduction”), in fact, only 50.4% of the patients have early-stage lung cancer UICC I (Tab. S7). The other half of the patient population consists of 24.8% locally advanced stages UICC II to IIIc and 18.3% advanced stages UICC IV. This very last group of patients, which makes up almost one-fifth of the whole cohort, has a much worse prognosis than the 50.4% early stages. Therefore, indications for follow-up CT will differ substantially from the first group. Given the fact that the proportion of patients with a Charlson Comorbidity Index (CCI) >7 is higher by a factor of 3 (42.9% vs. 12.6%) in the “other reason” CT group (Tab. S9), the decision for CT scan is more driven by clinical necessity for acute health conditions such as pleural effusion, metastases, atelectasis, than by elective surveillance purposes. Again, this underlines the impression that CT indications dichotomized by the terms “surveillance” and “other” are surrogates for different lung cancers, which—by the natural course of the disease—entail different survival outcomes. Thirdly, a major potential for improvement of the model is—as duly noted by the authors under limitations—that it currently does not incorporate factors like acute health conditions. In this context, the question arises whether OS is the appropriate endpoint. It is noteworthy that the median age of the 128 control patients without CT is substantially higher (Tab. S7: 68.8 years) than that of the 438 selected patients with surveillance CT scan (Tab. S9: 65.9 years), so that the OS rate is expected to be higher in the second group. Unfortunately, the study does not provide descriptive statistics comparing these two cohorts. It also remains unclear how many of the patients really died of lung cancer. This may call into question the clinical utility of the model tested by a time-to-event analysis with OS as an endpoint. Fig. A depicts a comparison of patients, including 18.3% stage IV (n=585) with CT, compared to a relatively well-selected group of lung cancer patients without CT (n=128) containing only 9.4% stage IV patients (Tab. S7). This difference by a factor of 2, alongside an approximately 10% difference in stage I patients (50.4% vs. 60.9%; Tab. S7), may be clinically relevant. There is no statistical difference between the two groups in terms of OS. From a clinical viewpoint the most probable reason for this finding is that patients with higher tumor burden (UICC IV) and frail general condition (CCI >7) die of their disease regardless of CTs for surveillance or other reasons, whereas the group of selected lung cancer patients with less stage IV (n=128) have a lower probability for death since they have per definition a better prognosis even if they do not get a surveillance CT. Fig. B shows the comparison of selected mainly early stage patients from the experimental cohort (n=438) who—because of the course of their disease, are less prone to acute health conditions than late stages—have regular (at least 1) surveillance CTs (n=438). The difference in clinical features and the potential selection bias become obvious from Tab. S9 with more than twice as many stage I and only one fifth of stage IV patients in this “well-selected” cohort (n=438) with surveillance CT compared to the 147 patients with CT scans for “other reasons”. While the 147 patients are eliminated from the experimental cohort, the low tumor burden patients with CT (n=438) live significantly longer than a comparable cohort without CT (n=128). The fact that the first comparison—in contrast to the second—revealed no difference, underpins the notion that CT indication represents a surrogate for disease stage.

In summary, the conclusion drawn by the authors [“This analysis showed a potential association between better OS and the receipt of at least one surveillance CT beyond 5-year survival.” (3)] might rather be modified as follows: patients with low tumor burden and good CCI benefit from regular CT surveillance compared to those without.


Acknowledgments

None.


Footnote

Provenance and Peer Review: This article was commissioned by the Editorial Office, Translational Lung Cancer Research. The article did not undergo external peer review.

Funding: None.

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://tlcr.amegroups.com/article/view/10.21037/tlcr-2025-1-1344/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Sung H, Ferlay J, Siegel RL, et al. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin 2021;71:209-49. [Crossref] [PubMed]
  2. Byrd C, Ajawara U, Laundry R, et al. Performance of a rule-based semi-automated method to optimize chart abstraction for surveillance imaging among patients treated for non-small cell lung cancer. BMC Med Inform Decis Mak 2022;22:148. [Crossref] [PubMed]
  3. Khan A, Choi E, Su C, et al. Automatic Abstraction of Computed Tomography Imaging Indication Using Natural Language Processing for Evaluation of Surveillance Patterns in Long-Term Lung Cancer Survivors. JCO Clin Cancer Inform 2025;9:e2400279. [Crossref] [PubMed]
  4. Goldstraw P, Chansky K, Crowley J, et al. The IASLC Lung Cancer Staging Project: Proposals for Revision of the TNM Stage Groupings in the Forthcoming (Eighth) Edition of the TNM Classification for Lung Cancer. J Thorac Oncol 2016;11:39-51. [Crossref] [PubMed]
  5. Antonia SJ, Villegas A, Daniel D, et al. Overall Survival with Durvalumab after Chemoradiotherapy in Stage III NSCLC. N Engl J Med 2018;379:2342-50. [Crossref] [PubMed]
  6. Girard N, Bar J, Garrido P, et al. Treatment Characteristics and Real-World Progression-Free Survival in Patients With Unresectable Stage III NSCLC Who Received Durvalumab After Chemoradiotherapy: Findings From the PACIFIC-R Study. J Thorac Oncol 2023;18:181-93. [Crossref] [PubMed]
  7. Zehentmayr F, Feurstein P, Ruznic E, et al. Durvalumab Prolongs Overall Survival, Whereas Radiation Dose Escalation > 66 Gy Might Improve Long-Term Local Control in Unresectable NSCLC Stage III: Updated Analysis of the Austrian Radio-Oncological Lung Cancer Study Association Registry (ALLSTAR). Cancers (Basel) 2025; [Crossref]
  8. Park CK, Oh HJ, Kim YC, et al. Korean Real-World Data on Patients With Unresectable Stage III NSCLC Treated With Durvalumab After Chemoradiotherapy: PACIFIC-KR. J Thorac Oncol 2023;18:1042-54. [Crossref] [PubMed]
  9. Faehling M, Schumann C, Christopoulos P, et al. Durvalumab after definitive chemoradiotherapy in locally advanced unresectable non-small cell lung cancer (NSCLC): Real-world data on survival and safety from the German expanded-access program (EAP). Lung Cancer 2020;150:114-22. [Crossref] [PubMed]
  10. Aupérin A, Le Péchoux C, Rolland E, et al. Meta-analysis of concomitant versus sequential radiochemotherapy in locally advanced non-small-cell lung cancer. J Clin Oncol 2010;28:2181-90. [Crossref] [PubMed]
  11. Spigel DR, Faivre-Finn C, Gray JE, et al. Five-Year Survival Outcomes From the PACIFIC Trial: Durvalumab After Chemoradiotherapy in Stage III Non-Small-Cell Lung Cancer. J Clin Oncol 2022;40:1301-11. [Crossref] [PubMed]
  12. Lou R, Lalevic D, Chambers C, et al. Automated Detection of Radiology Reports that Require Follow-up Imaging Using Natural Language Processing Feature Engineering and Machine Learning Classification. J Digit Imaging 2020;33:131-6. [Crossref] [PubMed]
  13. Pons E, Braun LM, Hunink MG, et al. Natural Language Processing in Radiology: A Systematic Review. Radiology 2016;279:329-43. [Crossref] [PubMed]
  14. Casey A, Davidson E, Poon M, et al. A systematic review of natural language processing applied to radiology reports. BMC Med Inform Decis Mak 2021;21:179. [Crossref] [PubMed]
Cite this article as: Grambozov B, Ruznic E, Zehentmayr F. Lung cancer survivorship: natural language processing for automated abstraction of follow-up computed tomography indication. Transl Lung Cancer Res 2026;15(1):2. doi: 10.21037/tlcr-2025-1-1344

Download Citation