SAME-clustering: Single-cell Aggregated Clustering via Mixture Model Ensemble

descriptionPublicationkeyboard_double_arrow_right Article , Other literature type , Preprint 24 May 2019Publisher:openRxivJournal:Nucleic Acids Research, volume 48, pages 86-95 (issn: 0305-1048, eissn: 1362-4962,

Copyright policy )Funded by:NIH | Design and Analysis of Se..., NIH | Genetic Studies of Blood ...

Authors: Ruth Huh; Yuchen Yang; Yuchao Jiang; Yin Shen; Yun Li;

doi: 10.1101/645820 , 10.1093/nar/gkz959

pmid: 31777938

pmc: PMC6943136

SAME-clustering: Single-cell Aggregated Clustering via Mixture Model Ensemble

- Summary
- Subjects
- Metrics

Abstract

ABSTRACT Clustering is an essential step in the analysis of single cell RNA-seq (scRNA-seq) data to shed light on tissue complexity including the number of cell types and transcriptomic signatures of each cell type. Due to its importance, novel methods have been developed recently for this purpose. However, different approaches generate varying estimates regarding the number of clusters and the single-cell level cluster assignments. This type of unsupervised clustering is challenging and it is often times hard to gauge which method to use because none of the existing methods outperform others across all scenarios. We present SAME-clustering, a mixture model-based approach that takes clustering solutions from multiple methods and selects a maximally diverse subset to produce an improved ensemble solution. We tested SAME-clustering across 15 scRNA-seq datasets generated by different platforms, with number of clusters varying from 3 to 15, and number of single cells from 49 to 32,695. Results show that our SAME-clustering ensemble method yields enhanced clustering, in terms of both cluster assignments and number of clusters. The mixture model ensemble clustering is not limited to clustering scRNA-seq data and may be useful to a wide range of clustering applications.

Related Organizations

University of California
United States
University of North Carolina at Chapel Hill
United States
Department of Neurology University of California
United States
University of California, San Francisco
United States
University of California, Berkeley
United States

View all View all

Keywords

Sequence Analysis, RNA, Gene Expression Profiling, Computational Biology, Cluster Analysis, Datasets as Topic, Humans, Single-Cell Analysis, Transcriptome, Algorithms

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	75
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 1%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 10%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 1%