Optimal ratio for data splitting

descriptionPublicationkeyboard_double_arrow_right Article , Preprint 04 Apr 2022Embargo end date: 01 Jan 2022 English Publisher:WileyJournal:Statistical Analysis and Data Mining: The ASA Data Science Journal, volume 15, pages 531-538 (issn: 1932-1864, eissn: 1932-1872,

Copyright policy )Funded by:NSF | Integrating Data- and Mod..., NSF | DMREF: Design of Organic-...

Authors: V. Roshan Joseph;

doi: 10.1002/sam.11583 , 10.48550/arxiv.2202.03326

arXiv: 2202.03326

Optimal ratio for data splitting

- Summary
- Subjects
- Metrics

Abstract

AbstractIt is common to split a dataset into training and testing sets before fitting a statistical or machine learning model. However, there is no clear guidance on how much data should be used for training and testing. In this article, we show that the optimal training/testing splitting ratio is , where is the number of parameters in a linear regression model that explains the data well.

Related Organizations

Georgia Institute of Technology
United States
GEORGIA TECH RESEARCH CORPORATION
United States

Keywords

FOS: Computer and information sciences, Computer Science - Machine Learning, Statistics - Machine Learning, Machine Learning (stat.ML), Machine Learning (cs.LG)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	596
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 0.1%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 1%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 0.01%