
pmid: 26147637
arXiv: 1501.01231
Canonical correlation analysis (CCA) describes the associations between two sets of variables by maximizing the correlation between linear combinations of the variables in each dataset. However, in high‐dimensional settings where the number of variables exceeds the sample size or when the variables are highly correlated, traditional CCA is no longer appropriate. This paper proposes a method for sparse CCA. Sparse estimation produces linear combinations of only a subset of variables from each dataset, thereby increasing the interpretability of the canonical variates. We consider the CCA problem from a predictive point of view and recast it into a regression framework. By combining an alternating regression approach together with a lasso penalty, we induce sparsity in the canonical vectors. We compare the performance with other sparse CCA techniques in different simulation settings and illustrate its usefulness on a genomic dataset.
FOS: Computer and information sciences, Biometry, Measures of association (correlation, canonical correlation, etc.), sparsity, Genomics, Applications of statistics to biology and medical sciences; meta analysis, Methodology (stat.ME), Regression Analysis, genomic data, penalized regression, Genetics and epigenetics, lasso, Statistics - Methodology, Algorithms, canonical correlation analysis
FOS: Computer and information sciences, Biometry, Measures of association (correlation, canonical correlation, etc.), sparsity, Genomics, Applications of statistics to biology and medical sciences; meta analysis, Methodology (stat.ME), Regression Analysis, genomic data, penalized regression, Genetics and epigenetics, lasso, Statistics - Methodology, Algorithms, canonical correlation analysis
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 46 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Top 10% | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
