Predicting the risk of lung cancer using machine learning: A large study based on UK Biobank

Author:

Zhang Siqi1,Yang Liangwei2,Xu Weiwen2,Wang Yue3,Han Liyuan4,Zhao Guofang2,Cai Ting4ORCID

Affiliation:

1. The Second School of Clinical Medicine, Zhejiang Chinese Medical University, Hangzhou, China

2. Department of Cardiothoracic Surgery, Ningbo No. 2 Hospital, Ningbo, China

3. School of Public Health, Medical College of Soochow University, Suzhou, China

4. Center for Cardiovascular and Cerebrovascular Epidemiology and Translational Medicine, Ningbo Institute of Life and Health Industry, University of Chinese Academy of Sciences, Ningbo, China.

Abstract

In response to the high incidence and poor prognosis of lung cancer, this study tends to develop a generalizable lung-cancer prediction model by using machine learning to define high-risk groups and realize the early identification and prevention of lung cancer. We included 467,888 participants from UK Biobank, using lung cancer incidence as an outcome variable, including 49 previously known high-risk factors and less studied or unstudied predictors. We developed multivariate prediction models using multiple machine learning models, namely logistic regression, naïve Bayes, random forest, and extreme gradient boosting models. The performance of the models was evaluated by calculating the areas under their receiver operating characteristic curves, Brier loss, log loss, precision, recall, and F1 scores. The Shapley additive explanations interpreter was used to visualize the models. Three were ultimately 4299 cases of lung cancer that were diagnosed in our sample. The model containing all the predictors had good predictive power, and the extreme gradient boosting model had the best performance with an area under curve of 0.998. New important predictive factors for lung cancer were also identified, namely hip circumference, waist circumference, number of cigarettes previously smoked daily, neuroticism score, age, and forced expiratory volume in 1 second. The predictive model established by incorporating novel predictive factors can be of value in the early identification of lung cancer. It may be helpful in stratifying individuals and selecting those at higher risk for inclusion in screening programs.

Publisher

Ovid Technologies (Wolters Kluwer Health)

Reference38 articles.

1. Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries.;Sung;CA Cancer J Clin,2021

2. The eighth edition lung cancer stage classification.;Detterbeck;Chest,2017

3. Lung cancer LDCT screening and mortality reduction – evidence, pitfalls and future perspectives.;Oudkerk;Nat Rev Clin Oncol,2021

4. Impact of low-dose computed tomography (LDCT) screening on lung cancer-related mortality.;Bonney;Cochrane Database Syst Rev,2022

5. Variations in lung cancer risk among smokers.;Bach;J Natl Cancer Inst,2003

Cited by 1 articles. 订阅此论文施引文献 订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3