Protein feature engineering framework for AMPylation site prediction-Reference-Cited by-同舟云学术

Protein feature engineering framework for AMPylation site prediction

Published:2024-04-15 Issue:1 Volume:14 Page:
ISSN:2045-2322
Container-title:Scientific Reports
language:en
Short-container-title:Sci Rep

Author:

Prabhu Hardik,Bhosale Hrushikesh,Sane Aamod,Dhadwal Renu,Ramakrishnan Vigneshwar,Valadi Jayaraman

Abstract

AbstractAMPylation is a biologically significant yet understudied post-translational modification where an adenosine monophosphate (AMP) group is added to Tyrosine and Threonine residues primarily. While recent work has illuminated the prevalence and functional impacts of AMPylation, experimental identification of AMPylation sites remains challenging. Computational prediction techniques provide a faster alternative approach. The predictive performance of machine learning models is highly dependent on the features used to represent the raw amino acid sequences. In this work, we introduce a novel feature extraction pipeline to encode the key properties relevant to AMPylation site prediction. We utilize a recently published dataset of curated AMPylation sites to develop our feature generation framework. We demonstrate the utility of our extracted features by training various machine learning classifiers, on various numerical representations of the raw sequences extracted with the help of our framework. Tenfold cross-validation is used to evaluate the model’s capability to distinguish between AMPylated and non-AMPylated sites. The top-performing set of features extracted achieved MCC score of 0.58, Accuracy of 0.8, AUC-ROC of 0.85 and F1 score of 0.73. Further, we elucidate the behaviour of the model on the set of features consisting of monogram and bigram counts for various representations using SHapley Additive exPlanations.

Publisher

Springer Science and Business Media LLC

Link

https://www.nature.com/articles/s41598-024-58450-8.pdf

Reference67 articles.

1. Ramazi, S. & Zahiri, J. Post-translational modifications in proteins: Resources, tools and prediction methods. Database 2021, baab012 (2021).

2. Alquezar, C., Arya, S. & Kao, A. W. Tau post-translational modifications: Dynamic transformers of tau function, degradation, and aggregation. Front. Neurol. 11, 595532 (2021).

3. Gong, C.-X., Liu, F., Grundke-Iqbal, I. & Iqbal, K. Post-translational modifications of tau protein in Alzheimer’s disease. J. Neural Transm. 112, 813–838 (2005).

4. Liu, J., Wang, Q., Kang, Y., Xu, S. & Pang, D. Unconventional protein post-translational modifications: The Helmsmen in breast cancer. Cell Biosci. 12, 1–28 (2022).

5. Kapoor, K., Chen, T. & Tajkhorshid, E. Posttranslational modifications optimize the ability of SARS-COV-2 spike for effective interaction with host cell receptors. Proc. Natl. Acad. Sci. 119, e2119761119 (2022).

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Explainable Machine Learning Model to Accurately Predict Protein-Binding Peptides;Algorithms;2024-09-12