
pmid: 29925568
pmc: PMC6010767
Abstract Multi-omic studies promise the improved characterization of biological processes across molecular layers. However, methods for the unsupervised integration of the resulting heterogeneous datasets are lacking. We present Multi-Omics Factor Analysis (MOFA), a computational method for discovering the principal sources of variation in multi-omic datasets. MOFA infers a set of (hidden) factors that capture biological and technical sources of variability. It disentangles axes of heterogeneity that are shared across multiple modalities and those specific to individual data modalities. The learnt factors enable a variety of downstream analyses, including identification of sample subgroups, data imputation, and the detection of outlier samples. We applied MOFA to a cohort of 200 patient samples of chronic lymphocytic leukaemia, profiled for somatic mutations, RNA expression, DNA methylation and ex-vivo drug responses. MOFA identified major dimensions of disease heterogeneity, including immunoglobulin heavy chain variable region status, trisomy of chromosome 12 and previously underappreciated drivers, such as response to oxidative stress. In a second application, we used MOFA to analyse single-cell multiomics data, identifying coordinated transcriptional and epigenetic changes along cell differentiation.
Medicine (General), QH301-705.5, Datasets as Topic, 610 Medicine & health, Antineoplastic Agents, 1100 General Agricultural and Biological Sciences, General Biochemistry, Genetics and Molecular Biology, R5-920, 2604 Applied Mathematics, 1300 General Biochemistry, Genetics and Molecular Biology, 2400 General Immunology and Microbiology, multi‐omics, Methods, Humans, Computer Simulation, single‐cell omics, Biology (General), 610 Medicine & health, data integration, dimensionality reduction, Models, Statistical, General Immunology and Microbiology, Applied Mathematics, Computational Biology, personalized medicine, Leukemia, Lymphocytic, Chronic, B-Cell, Oxidative Stress, Computational Theory and Mathematics, 10032 Clinic for Oncology and Hematology, Data Integration ; Dimensionality Reduction ; Multi-omics ; Personalized Medicine ; Single-cell Omics, General Agricultural and Biological Sciences, Transcriptome, Software, Information Systems
Medicine (General), QH301-705.5, Datasets as Topic, 610 Medicine & health, Antineoplastic Agents, 1100 General Agricultural and Biological Sciences, General Biochemistry, Genetics and Molecular Biology, R5-920, 2604 Applied Mathematics, 1300 General Biochemistry, Genetics and Molecular Biology, 2400 General Immunology and Microbiology, multi‐omics, Methods, Humans, Computer Simulation, single‐cell omics, Biology (General), 610 Medicine & health, data integration, dimensionality reduction, Models, Statistical, General Immunology and Microbiology, Applied Mathematics, Computational Biology, personalized medicine, Leukemia, Lymphocytic, Chronic, B-Cell, Oxidative Stress, Computational Theory and Mathematics, 10032 Clinic for Oncology and Hematology, Data Integration ; Dimensionality Reduction ; Multi-omics ; Personalized Medicine ; Single-cell Omics, General Agricultural and Biological Sciences, Transcriptome, Software, Information Systems
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 1K | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 0.01% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Top 1% | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 0.1% |
