GASPIDs Versus Non-GASPIDs - Differentiation Based on Machine Learning Approach

Author:

Ahmad Fawad1ORCID,Ikram Saima1,Ahmad Jamshaid1,Ullah Waseem2,Hassan Fahad1,Khattak Saeed Ullah1,Irshad Ur Rehman 1

Affiliation:

1. Centre of Biotechnology & Microbiology, University of Peshawar, Peshawar, Pakistan

2. College of Software Convergence, Sejong University, Seoul, South Korea

Abstract

Background: Peptidases are a group of enzymes which catalyze the cleavage of peptide bonds. Around 2-3% of the whole genome codes for proteases and about one-third of all known proteases are serine proteases which are divided into 13 clans and 40 families. They are involved in diverse physiological roles such as digestion, coagulation of blood, fibrinolysis, processing of proteins and prohormones, signaling pathways, complement fixation, and have a vital role in the immune defense system. Based on their functions, they can broadly be divided into two classes; GASPIDs (Granule Associated Serine Peptidases involved in Immune Defense System) and Non- GASPIDs. GASPIDs, in particular are involved in immune-associated functions i.e. initiating apoptosis to kill virally infected and cancerous cells, cytokine modulation for the generation of inflammatory responses, and direct killing of pathogens through phagosomes. Methods: In this study, sequence-based characterization of these two types of serine proteases is performed. We first identified sequences by analyzing multiple online databases as well as by analyzing whole genomes of different species from different orthologous and non-orthologous species. Sequences were identified by devising a distinct criterion to differentiate GASPIDs from Non-GASPIDs. The translated version of these sequences was then subjected to feature extraction. Using these distinctive features, we differentiated GASPIDs from Non-GASPIDs by applying multiple supervised machine learning models. Results and Conclusion: Our results show that, among the three classifiers used in this study, SVM classifier coupled with tripeptide as feature method has shown the best accuracy in classification of sequences as GASPIDs and Non-GASPIDs.

Publisher

Bentham Science Publishers Ltd.

Subject

Computational Mathematics,Genetics,Molecular Biology,Biochemistry

Reference28 articles.

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3