Format

Send to

Choose Destination
Biol Direct. 2018 May 9;13(1):8. doi: 10.1186/s13062-018-0209-6.

Use of designed sequences in protein structure recognition.

Author information

1
Lab 103, Molecular Biophysics Unit, Indian Institute of Science, Bangalore, Karnataka, 560012, India.
2
Present address: Institute for Research in Biomedicine (IRB), Parc Cientific de Barcelona, C/ Baldiri Reixac 10, 08028, Barcelona, Spain.
3
Lab 103, Molecular Biophysics Unit, Indian Institute of Science, Bangalore, Karnataka, 560012, India. ns@iisc.ac.in.
4
Lab 103, Molecular Biophysics Unit, Indian Institute of Science, Bangalore, Karnataka, 560012, India. sandhyas@iisc.ac.in.

Abstract

BACKGROUND:

Knowledge of the protein structure is a pre-requisite for improved understanding of molecular function. The gap in the sequence-structure space has increased in the post-genomic era. Grouping related protein sequences into families can aid in narrowing the gap. In the Pfam database, structure description is provided for part or full-length proteins of 7726 families. For the remaining 52% of the families, information on 3-D structure is not yet available. We use the computationally designed sequences that are intermediately related to two protein domain families, which are already known to share the same fold. These strategically designed sequences enable detection of distant relationships and here, we have employed them for the purpose of structure recognition of protein families of yet unknown structure.

RESULTS:

We first measured the success rate of our approach using a dataset of protein families of known fold and achieved a success rate of 88%. Next, for 1392 families of yet unknown structure, we made structural assignments for part/full length of the proteins. Fold association for 423 domains of unknown function (DUFs) are provided as a step towards functional annotation.

CONCLUSION:

The results indicate that knowledge-based filling of gaps in protein sequence space is a lucrative approach for structure recognition. Such sequences assist in traversal through protein sequence space and effectively function as 'linkers', where natural linkers between distant proteins are unavailable.

REVIEWERS:

This article was reviewed by Oliviero Carugo, Christine Orengo and Srikrishna Subramanian.

KEYWORDS:

Fold-assignment; Function annotation; Homology detection; Sequence-structure gap; Structural domain assignment; Structure recognition

PMID:
29776380
PMCID:
PMC5960202
DOI:
10.1186/s13062-018-0209-6
[Indexed for MEDLINE]
Free PMC Article

Supplemental Content

Full text links

Icon for BioMed Central Icon for PubMed Central
Loading ...
Support Center