Format

Send to

Choose Destination
See comment in PubMed Commons below
Mol Cell Proteomics. 2006 Aug;5(8):1520-32. Epub 2006 Jun 11.

DivergentSet, a tool for picking non-redundant sequences from large sequence collections.

Author information

1
Department of Chemistry and Biochemistry, University of Colorado, Boulder, Colorado 80309, USA.

Abstract

DivergentSet addresses the important but so far neglected bioinformatics task of choosing a representative set of sequences from a larger collection. We found that using a phylogenetic tree to guide the construction of divergent sets of sequences can be up to 2 orders of magnitude faster than the naive method of using a full distance matrix. By providing a user-friendly interface (available online) that integrates the tasks of finding additional sequences, building and refining the divergent set, producing random divergent sets from the same sequences, and exporting identifiers, this software facilitates a wide range of bioinformatics analyses including finding significant motifs and covariations. As an example application of DivergentSet, we demonstrate that the motifs identified by the motif-finding package MEME (Motif Elicitation by Maximum Entropy) are highly unstable with respect to the specific choice of sequences. This instability suggests that the types of sensitivity analysis enabled by DivergentSet may be widely useful for identifying the motifs of biological significance.

PMID:
16769708
DOI:
10.1074/mcp.T600022-MCP200
[Indexed for MEDLINE]
Free full text
PubMed Commons home

PubMed Commons

0 comments
How to join PubMed Commons

    Supplemental Content

    Full text links

    Icon for HighWire
    Loading ...
    Support Center