
The increased dependence of industries on data-driven decision-making practices makes it obvious that quality data and governance systems are necessary. Contemporary organisations have to deal with extensive and heterogeneous data sets and are highly concerned about the quality of data and compliance. This survey overcomes such obstacles by suggesting a taxonomy of traditional data-quality dimensions, e.g., accuracy, completeness, and timeliness, and machine-learning-specific dimensions, e.g., training-serving skew, label noise, feature freshness, and feedback-loop contamination. It further offers a structured comparison, mapping validation, drift detection, and monitoring methods to the circumstances in which each method is best suited. The reviewed articles cover the period of 2022-2024 of algorithmic drift detection, serverless monitoring architecture, deep-learning-based validation, and industry-based reliability. With this synthesis, a number of relevant trends can be identified: model-based drift detection, cost-efficient data cleaning and adaptive monitoring. At the same time, there are still gaps, e.g., the lack of standardized benchmarks, the limited scalability of available solutions, and insufficient enterprise preparedness to implement them. This survey thus makes data-quality engineering an important foundation for trustworthy artificial intelligence, identifying its benefits and drawbacks. Sometimes, more specific taxonomies and practical frameworks can inform practitioners to build pipelines that sustain the results of the model and guarantee future confidence in real-world deployments.
Artificial intelligence, Data monitoring and drift detection, Data quality engineering, Data validation, Concept drift
Artificial intelligence, Data monitoring and drift detection, Data quality engineering, Data validation, Concept drift
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
