Data Readiness for AI: A 360-Degree Survey

Name: Data Readiness for AI: A 360-Degree Survey
Keywords: FOS: Computer and information sciences, Computer Science - Machine Learning, Networking and Information Technology R&D (NITRD) (rcdc), AI-ready data, Computer Science - Artificial Intelligence, Generic health relevance (hrcs-hc), 08 Information and Computing Sciences (for), I.2.0, Machine Learning (cs.LG), 46 Information and Computing Sciences (for-2020)

Kaveen Hiniduma; Suren Byna; Jean Luca Bez

Found an issue? Give us feedback

ACM Computing Survey...arrow_drop_down

ACM Computing Surveys

Article . 2025 . Peer-reviewed

License: CC BY

Data sources: Crossref

arXiv.org e-Print Archive

Preprint . 2024

Data sources: arXiv.org e-Print Archive

eScholarship - University of California

Article . 2025

Data sources: eScholarship - University of California

https://dx.doi.org/10.48550/ar...

Article . 2024

License: CC BY NC ND

Data sources: Datacite

DBLP

Article

Data sources: DBLP

DBLP

Article

Data sources: DBLP

Data Readiness for AI: A 360-Degree Survey

descriptionPublicationkeyboard_double_arrow_right Article , Preprint 04 Apr 2025Embargo end date: 01 Jan 2024 United States English Publisher:Association for Computing Machinery (ACM)Journal:ACM Computing Surveys, volume 57, pages 1-39 (issn: 0360-0300, eissn: 1557-7341,

Copyright policy )

Authors: Kaveen Hiniduma; Suren Byna; Jean Luca Bez;

doi: 10.1145/3722214 , 10.48550/arxiv.2404.05779

arXiv: 2404.05779

Data Readiness for AI: A 360-Degree Survey

- Summary
- Subjects
- Related research
  (1)
- Metrics

Abstract

Artificial Intelligence (AI) applications critically depend on data. Poor-quality data produces inaccurate and ineffective AI models that may lead to incorrect or unsafe use. Evaluation of data readiness is a crucial step in improving the quality and appropriateness of data usage for AI. R&D efforts have been spent on improving data quality. However, standardized metrics for evaluating data readiness for use in AI training are still evolving. In this study, we perform a comprehensive survey of metrics used to verify data readiness for AI training. This survey examines more than 140 papers published by ACM Digital Library, IEEE Xplore, journals such as Nature, Springer, and Science Direct, and online articles published by prominent AI experts. This survey aims to propose a taxonomy of data readiness for AI (DRAI) metrics for structured and unstructured datasets. We anticipate that this taxonomy will lead to new standards for DRAI metrics that would be used for enhancing the quality, accuracy, and fairness of AI training and inference.

Country

United States

Related Organizations

University of California, San Francisco
United States
Lawrence Berkeley National Laboratory
United States
The Ohio State University
United States

Keywords

FOS: Computer and information sciences, Computer Science - Machine Learning, Networking and Information Technology R&D (NITRD) (rcdc), AI-ready data, Computer Science - Artificial Intelligence, Generic health relevance (hrcs-hc), 08 Information and Computing Sciences (for), I.2.0, Machine Learning (cs.LG), 46 Information and Computing Sciences (for-2020), I.2.0; E.m, Artificial Intelligence (cs.AI), Information Systems (science-metrix), 46 Information and computing sciences (for-2020), data quality metrics, 4608 Human-Centred Computing (for-2020), Machine Learning and Artificial Intelligence (rcdc), Data readiness, E.m

1 Research products, page 1 of 1

piq software on GitHub
IsRelatedTo

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	14
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 10%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 10%

Found an issue? Give us feedback

14

Top 10%

Green

hybrid

Data Readiness for AI: A 360-Degree Survey

Data Readiness for AI: A 360-Degree Survey

1 Research products, page 1 of 1

piq software on GitHub