Format

Send to

Choose Destination
Nucleic Acids Res. 2014 Feb;42(4):e25. doi: 10.1093/nar/gkt1141. Epub 2013 Nov 19.

UnSplicer: mapping spliced RNA-Seq reads in compact genomes and filtering noisy splicing.

Author information

1
Joint Georgia Tech and Emory Wallace H. Coulter Department of Biomedical Engineering, Atlanta, GA 30332, USA, Department of Bioengineering, University of Illinois at Urbana-Champaign, IL 61801, USA, Institute for Genomic Biology, University of Illinois at Urbana-Champaign, IL 61801, USA, School of Computational Science & Engineering, Georgia Tech, Atlanta, GA 30332, USA and Department of Bioinformatics, Moscow Institute of Physics and Technology, Moscow, 141700, Russia.

Abstract

Accurate mapping of spliced RNA-Seq reads to genomic DNA has been known as a challenging problem. Despite significant efforts invested in developing efficient algorithms, with the human genome as a primary focus, the best solution is still not known. A recently introduced tool, TrueSight, has demonstrated better performance compared with earlier developed algorithms such as TopHat and MapSplice. To improve detection of splice junctions, TrueSight uses information on statistical patterns of nucleotide ordering in intronic and exonic DNA. This line of research led to yet another new algorithm, UnSplicer, designed for eukaryotic species with compact genomes where functional alternative splicing is likely to be dominated by splicing noise. Genome-specific parameters of the new algorithm are generated by GeneMark-ES, an ab initio gene prediction algorithm based on unsupervised training. UnSplicer shares several components with TrueSight; the difference lies in the training strategy and the classification algorithm. We tested UnSplicer on RNA-Seq data sets of Arabidopsis thaliana, Caenorhabditis elegans, Cryptococcus neoformans and Drosophila melanogaster. We have shown that splice junctions inferred by UnSplicer are in better agreement with knowledge accumulated on these well-studied genomes than predictions made by earlier developed tools.

PMID:
24259430
PMCID:
PMC3936741
DOI:
10.1093/nar/gkt1141
[Indexed for MEDLINE]
Free PMC Article

Supplemental Content

Full text links

Icon for Silverchair Information Systems Icon for PubMed Central
Loading ...
Support Center