
This repo contain supplementary tables from the manuscript entitled: "Scalable nonparametric clustering with unified marker gene selection for single-cell RNA-seq data". Clustering is a common way to identify cell types in single-cell RNA-sequencing (scRNA-seq) data. Unfortunately, current methods (i) require users to make human-in-the-loop decisions, which adds significant runtime to bioinformatic analyses, and (ii) reuse the same data twice when testing for differentially expressed genes, which can lead to an increased number of false discoveries. In this work, we overcome these limitations with NCLUSION: a Bayesian nonparametric method that simultaneously clusters cells and selects marker genes. NCLUSION operates without user-defined heuristics to set model parameters and leverages variational expectation-maximization (EM) for posterior inference which allows it to scale well up to 1 million cells. By analyzing publicly available datasets, we illustrate that NCLUSION matches the state-of-the-art clustering performance of competing approaches, achieves improved computational efficiency, and directly enables identification of biologically relevant gene sets driving cluster definitions.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
