Assessing the low complexity of protein sequences via the low complexity triangle-Reference-Cited by-同舟云学术

Assessing the low complexity of protein sequences via the low complexity triangle

Published:2020-12-30 Issue:12 Volume:15 Page:e0239154
ISSN:1932-6203
Container-title:PLOS ONE
language:en
Short-container-title:PLoS ONE

Author:

Mier Pablo^ORCID,Andrade-Navarro Miguel A.

Abstract

Background Proteins with low complexity regions (LCRs) have atypical sequence and structural features. Their amino acid composition varies from the expected, determined proteome-wise, and they do not follow the rules of structural folding that prevail in globular regions. One way to characterize these regions is by assessing the repeatability of a sequence, that is, calculating the local propensity of a region to be part of a repeat. Results We combine two local measures of low complexity, repeatability (using the RES algorithm) and fraction of the most frequent amino acid, to evaluate different proteomes, datasets of protein regions with specific features, and individual cases of proteins with extreme compositions. We apply a representation called ‘low complexity triangle’ as a proof-of-concept to represent the low complexity measured values. Results show that proteomes have distinct signatures in the low complexity triangle, and that these signatures are associated to complexity features of the sequences. We developed a web tool called LCT (http://cbdm-01.zdv.uni-mainz.de/~munoz/lct/) to allow users to calculate the low complexity triangle of a given protein or region of interest. Conclusions The low complexity triangle proves to be a suitable procedure to represent the general low complexity of a sequence or protein dataset. Homorepeats, direpeats, compositionally biased regions and globular regions occupy characteristic positions in the triangle. The described pipeline can be used to characterize LCRs and may help in quantifying the content of degenerated tandem repeats in proteins and proteomes.

Funder

Deutsche Forschungsgemeinschaft

Publisher

Public Library of Science (PLoS)

Subject

Multidisciplinary

Reference39 articles.

1. Exceptionally Abundant Exceptions: Comprehensive Characterization of Intrinsic Disorder in All Domains of Life;Z Peng;Cell Mol Life Sci,2015

2. Protein Homorepeats Sequences, Structures, Evolution, and Functions.;J Jorda;Adv Protein Chem Struct Biol,2010

3. Tandem Repeats in Proteins: From Sequence to Structure;AV Kajava;J Struct Biol,2012

4. Tandem and Cryptic Amino Acid Repeats Accumulate in Disordered Regions of Proteins;M Simon;Genome Biol,2009

5. Disentangling the Complexity of Low Complexity Proteins;P Mier;Brief Bioinform,2020

Cited by 7 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Terminal regions of a protein are a hotspot for low complexity regions and selection;Open Biology;2024-06

2. Optimizing strategy for the discovery of compositionally-biased or low-complexity regions in proteins;Scientific Reports;2024-01-05

3. Patterns of low-complexity regions in human genes;2023-12-04

4. Terminal regions of a protein are a hotspot for low complexity regions (LCRs) and selection;2023-07-07

5. fLPS 2.0: rapid annotation of compositionally-biased regions in biological sequences;PeerJ;2021-10-28