
arXiv: 1707.09752
Real data often contain anomalous cases, also known as outliers. These may spoil the resulting analysis but they may also contain valuable information. In either case, the ability to detect such anomalies is essential. A useful tool for this purpose is robust statistics, which aims to detect the outliers by first fitting the majority of the data and then flagging data points that deviate from it. We present an overview of several robust methods and the resulting graphical outlier detection tools. We discuss robust procedures for univariate, low‐dimensional, and high‐dimensional data, such as estimating location and scatter, linear regression, principal component analysis, classification, clustering, and functional data analysis. Also the challenging new topic of cellwise outliers is introduced.WIREs Data Mining Knowl Discov2018, 8:e1236. doi: 10.1002/widm.1236This article is categorized under:Algorithmic Development > Spatial and Temporal Data MiningTechnologies > ClassificationTechnologies > Structure Discovery and ClusteringTechnologies > Visualization
SELECTION, FOS: Computer and information sciences, Technology, Science & Technology, 0804 Data Format, Machine Learning (stat.ML), MULTIVARIATE LOCATION, Computer Science, Artificial Intelligence, MATLAB LIBRARY, 4605 Data management and data science, DISPERSION, Computer Science, Theory & Methods, Statistics - Machine Learning, 0806 Information Systems, Computer Science, REGRESSION, 0801 Artificial Intelligence and Image Processing, PRINCIPAL COMPONENT ANALYSIS, LINEAR-MODEL, ALGORITHM, COVARIANCE, ESTIMATORS
SELECTION, FOS: Computer and information sciences, Technology, Science & Technology, 0804 Data Format, Machine Learning (stat.ML), MULTIVARIATE LOCATION, Computer Science, Artificial Intelligence, MATLAB LIBRARY, 4605 Data management and data science, DISPERSION, Computer Science, Theory & Methods, Statistics - Machine Learning, 0806 Information Systems, Computer Science, REGRESSION, 0801 Artificial Intelligence and Image Processing, PRINCIPAL COMPONENT ANALYSIS, LINEAR-MODEL, ALGORITHM, COVARIANCE, ESTIMATORS
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 208 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 1% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Top 1% | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 1% |
