Relative-Error $CUR$ Matrix Decompositions

descriptionPublicationkeyboard_double_arrow_right Article , Preprint 01 Jan 2008Embargo end date: 01 Jan 2007 English Publisher:Society for Industrial & Applied Mathematics (SIAM)Journal:SIAM Journal on Matrix Analysis and Applications, volume 30, pages 844-881 (issn: 0895-4798, eissn: 1095-7162,

Copyright policy )

Authors: Petros Drineas; Michael W. Mahoney; S. Muthukrishnan 0001;

doi: 10.1137/07070471x , 10.48550/arxiv.0708.3696

arXiv: 0708.3696

Relative-Error $CUR$ Matrix Decompositions

- Summary
- Subjects
- Metrics

Abstract

Many data analysis applications deal with large matrices and involve approximating the matrix using a small number of ``components.'' Typically, these components are linear combinations of the rows and columns of the matrix, and are thus difficult to interpret in terms of the original features of the input data. In this paper, we propose and study matrix approximations that are explicitly expressed in terms of a small number of columns and/or rows of the data matrix, and thereby more amenable to interpretation in terms of the original data. Our main algorithmic results are two randomized algorithms which take as input an $m \times n$ matrix $A$ and a rank parameter $k$. In our first algorithm, $C$ is chosen, and we let $A'=CC^+A$, where $C^+$ is the Moore-Penrose generalized inverse of $C$. In our second algorithm $C$, $U$, $R$ are chosen, and we let $A'=CUR$. ($C$ and $R$ are matrices that consist of actual columns and rows, respectively, of $A$, and $U$ is a generalized inverse of their intersection.) For each algorithm, we show that with probability at least $1-δ$: $$ ||A-A'||_F \leq (1+ε) ||A-A_k||_F, $$ where $A_k$ is the ``best'' rank-$k$ approximation provided by truncating the singular value decomposition (SVD) of $A$. The number of columns of $C$ and rows of $R$ is a low-degree polynomial in $k$, $1/ε$, and $\log(1/δ)$. Our two algorithms are the first polynomial time algorithms for such low-rank matrix approximations that come with relative-error guarantees; previously, in some cases, it was not even known whether such matrix decompositions exist. Both of our algorithms are simple, they take time of the order needed to approximately compute the top $k$ singular vectors of $A$, and they use a novel, intuitive sampling method called ``subspace sampling.''

40 pages, 10 figures

Related Organizations

Rensselaer Polytechnic Institute
United States
Google (United States)
United States
Yahoo!
United States

Keywords

FOS: Computer and information sciences, Computer Science - Data Structures and Algorithms, Data Structures and Algorithms (cs.DS)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	232
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 1%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 1%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 10%

Found an issue? Give us feedback

232

Top 1%

Top 10%

Green

bronze

Fields of Science (4) View all

Fields of Science