Format

Send to

Choose Destination
PLoS One. 2013;8(3):e56859. doi: 10.1371/journal.pone.0056859. Epub 2013 Mar 11.

Edge principal components and squash clustering: using the special structure of phylogenetic placement data for sample comparison.

Author information

1
Fred Hutchinson Cancer Research Center, Seattle, Washington, United States of America. matsen@fhcrc.org

Erratum in

  • PLoS One. 2013;8(6). doi:10.1371/annotation/40cb3123-845a-43e7-b4c0-9fb00b6e2212.

Abstract

Principal components analysis (PCA) and hierarchical clustering are two of the most heavily used techniques for analyzing the differences between nucleic acid sequence samples taken from a given environment. They have led to many insights regarding the structure of microbial communities. We have developed two new complementary methods that leverage how this microbial community data sits on a phylogenetic tree. Edge principal components analysis enables the detection of important differences between samples that contain closely related taxa. Each principal component axis is a collection of signed weights on the edges of the phylogenetic tree, and these weights are easily visualized by a suitable thickening and coloring of the edges. Squash clustering outputs a (rooted) clustering tree in which each internal node corresponds to an appropriate "average" of the original samples at the leaves below the node. Moreover, the length of an edge is a suitably defined distance between the averaged samples associated with the two incident nodes, rather than the less interpretable average of distances produced by UPGMA, the most widely used hierarchical clustering method in this context. We present these methods and illustrate their use with data from the human microbiome.

PMID:
23505415
PMCID:
PMC3594297
DOI:
10.1371/journal.pone.0056859
[Indexed for MEDLINE]
Free PMC Article

Supplemental Content

Full text links

Icon for Public Library of Science Icon for PubMed Central
Loading ...
Support Center