Display Settings:

Format

Send to:

Choose Destination
See comment in PubMed Commons below
Genome Res. 2004 Mar;14(3):426-41. Epub 2004 Feb 12.

The multiassembly problem: reconstructing multiple transcript isoforms from EST fragment mixtures.

Author information

  • 1UCLA-DOE Center for Genomics and Proteomics, Molecular Biology Institute and Department of Chemistry & Biochemistry, University of California, Los Angeles, Los Angeles, California 90095-1570, USA.

Abstract

Recent evidence of abundant transcript variation (e.g., alternative splicing, alternative initiation, alternative polyadenylation) in complex genomes indicates that cataloging the complete set of transcripts from an organism is an important project. One challenge is the fact that most high-throughput experimental methods for characterizing transcripts (such as EST sequencing) give highly detailed information about short fragments of transcripts or protein products, instead of a complete characterization of a full-length form. We analyze this "multiassembly problem"-reconstructing the most likely set of full-length isoform sequences from a mixture of EST fragment data-and present a graph-based algorithm for solving it. In a variety of tests, we demonstrate that this algorithm deals appropriately with coupling of distinct alternative splicing events, increasing fragmentation of the input data and different types of transcript variation (such as alternative splicing, initiation, polyadenylation, and intron retention). To test the method's performance on pure fragment (EST) data, we removed all mRNA sequences, and found it produced no errors in 40 cases tested. Using this algorithm, we have constructed an Alternatively Spliced Proteins database (ASP) from analysis of human expressed and genomic sequences, consisting of 13,384 protein isoforms of 4422 genes, yielding an average of 3.0 protein isoforms per gene.

PMID:
14962984
[PubMed - indexed for MEDLINE]
PMCID:
PMC353230
Free PMC Article

Images from this publication.See all images (11)Free text

Figure 1
Figure 3
Figure 2
Figure 5
Figure 4
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
PubMed Commons home

PubMed Commons

0 comments
How to join PubMed Commons

    Supplemental Content

    Full text links

    Icon for HighWire Icon for PubMed Central
    Loading ...
    Write to the Help Desk