Random kernel k-nearest neighbors regression

descriptionPublicationkeyboard_double_arrow_right Article 01 Jul 2024Publisher:Frontiers Media SAJournal:Frontiers in Big Data, volume 7 (eissn: 2624-909X,

Copyright policy )

Authors: Patchanok Srisuradetchai; Korn Suksrikran;

doi: 10.3389/fdata.2024.1402384

pmid: 39011467

pmc: PMC11246867

Random kernel k-nearest neighbors regression

- Summary
- Subjects
- Metrics

Abstract

The k-nearest neighbors (KNN) regression method, known for its nonparametric nature, is highly valued for its simplicity and its effectiveness in handling complex structured data, particularly in big data contexts. However, this method is susceptible to overfitting and fit discontinuity, which present significant challenges. This paper introduces the random kernel k-nearest neighbors (RK-KNN) regression as a novel approach that is well-suited for big data applications. It integrates kernel smoothing with bootstrap sampling to enhance prediction accuracy and the robustness of the model. This method aggregates multiple predictions using random sampling from the training dataset and selects subsets of input variables for kernel KNN (K-KNN). A comprehensive evaluation of RK-KNN on 15 diverse datasets, employing various kernel functions including Gaussian and Epanechnikov, demonstrates its superior performance. When compared to standard KNN and the random KNN (R-KNN) models, it significantly reduces the root mean square error (RMSE) and mean absolute error, as well as improving R-squared values. The RK-KNN variant that employs a specific kernel function yielding the lowest RMSE will be benchmarked against state-of-the-art methods, including support vector regression, artificial neural networks, and random forests.

Related Organizations

Northwest Normal University
China (People's Republic of)
Northeast Normal University
China (People's Republic of)
Shenzhen University
China (People's Republic of)
Thammasat University
Thailand
University of Adelaide
Australia

Keywords

k-nearest neighbors regression, Big Data, feature selection, bootstrapping, Information technology, kernel k-nearest neighbors, T58.5-58.64, state-of-the-art (SOTA)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	58
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 1%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 10%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 1%

Found an issue? Give us feedback

58

Top 1%

Top 10%

Top 1%

Green

gold

Fields of Science (4) View all

Fields of Science