Extended Isolation Forest

descriptionPublicationkeyboard_double_arrow_right Article , Preprint , Other literature type 01 Apr 2021Embargo end date: 01 Jan 2018Publisher:Institute of Electrical and Electronics Engineers (IEEE)Journal:IEEE Transactions on Knowledge and Data Engineering, volume 33, pages 1,479-1,489 (issn: 1041-4347, eissn: 2326-3865,

Copyright policy )

Authors: Sahand Hariri; Matias Carrasco Kind; Robert J. Brunner;

doi: 10.1109/tkde.2019.2947676 , 10.48550/arxiv.1811.02141

arXiv: 1811.02141

Extended Isolation Forest

- Summary
- Subjects
- Metrics

Abstract

We present an extension to the model-free anomaly detection algorithm, Isolation Forest. This extension, named Extended Isolation Forest (EIF), resolves issues with assignment of anomaly score to given data points. We motivate the problem using heat maps for anomaly scores. These maps suffer from artifacts generated by the criteria for branching operation of the binary tree. We explain this problem in detail and demonstrate the mechanism by which it occurs visually. We then propose two different approaches for improving the situation. First we propose transforming the data randomly before creation of each tree, which results in averaging out the bias. Second, which is the preferred way, is to allow the slicing of the data to use hyperplanes with random slopes. This approach results in remedying the artifact seen in the anomaly score heat maps. We show that the robustness of the algorithm is much improved using this method by looking at the variance of scores of data points distributed along constant level sets. We report AUROC and AUPRC for our synthetic datasets, along with real-world benchmark datasets. We find no appreciable difference in the rate of convergence nor in computation time between the standard Isolation Forest and EIF.

12 pages; 21 figures, Published. Open source code in https://github.com/sahandha/eif

Related Organizations

National Center for Supercomputing Applications
United States
University of Illinois at Urbana Champaign
United States
University of Illinois Urbana-Champaign
United States
University of Illinois at Urbana–Champaign
United States

Keywords

FOS: Computer and information sciences, Computer Science - Machine Learning, Statistics - Machine Learning, FOS: Physical sciences, Machine Learning (stat.ML), Astrophysics - Instrumentation and Methods for Astrophysics, Instrumentation and Methods for Astrophysics (astro-ph.IM), Machine Learning (cs.LG)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	200
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 0.1%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 1%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 0.1%