Facilitating clinical trials in Polycythemia vera (PV) by identifying patient cohorts at high near-term risk of thrombosis using rich data and machine learning

Author:

Abu-Zeinah Ghaith,Krichevsky SpencerORCID,Erdos Katie,Silver Richard T.,Scandura Joseph M.ORCID

Abstract

AbstractThrombosis remains the leading cause of morbidity and mortality for patients (pts) with polycythemia vera (PV), yet PV clinical trials are not powered to identify interventions that improve thrombosis-free survival (TFS). Such trials are infeasible in a contemporary PV cohort, even when selecting “high-risk” pts based on Age >60 and thrombosis history, because thousands of patients would be required for a short-term study to meet TFS endpoint. To address this problem, we used artificial intelligence and machine learning (ML) to dynamically predict near-term (1-year) thrombosis risk in PV pts with high sensitivity and positive predictive value (PPV) to enhance pts selection. Our automation-driven data extraction methods yielded more than 16 million data elements across 1,448 unique variables (parameters) from 11,123 clinical visits for 470 pts. Using the AutoGluon framework, the Random Forest ML classification algorithm was selected as the top performer. The full (309-parameter) model performed very well (F1=0.91, AUC=0.84) when compared with the current ELN gold-standard for thrombosis risk stratification in PV (F1=0.1, AUC=0.39). Parameter engineering, guided by Gini feature importance identified the 21 parameters (top-21) most important for accurate prediction. The top-21 parameters included known, suspected and previously unappreciated thrombosis risk factors. To identify the minimum number of parameters required for the accurate ML prediction, we tested the performance of every possible combination of 3-9 parameters from top-21 (>1.6M combinations). High-performing models (F1> 0.8) most frequently included age (continuous), time since dx, time since thrombosis, complete blood count parameters, blood type, body mass index, and JAK2 mutant allele frequency. Having trained at tested over 1.6M practical ML models with a feasible number of parameters (3-9 parameters in top-21 most predictive), it is clear that study cohorts of patients with PV at high near-term thrombosis risk can be identified with high enough sensitivity and PPV to power a clinical trial for TFS. Further validation with external, multicenter cohorts is ongoing to establish a universal ML model for PV thrombosis that would facilitate clinical trials aimed at improving TFS.

Publisher

Cold Spring Harbor Laboratory

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3