- Model performance is bounded by training data quality.
- MIT found 3.4% label errors in major ML benchmark datasets.
- ISO/IEC 5259 (2024) standardizes AI data quality measurement.
Data quality in AI is the measure of how accurate, complete, consistent, and representative a dataset is for training, testing, and validating machine learning models.
Why it matters
The performance of any machine learning model is bounded by the quality of the data it learns from. Andrew Ng, co-founder of Google Brain and Stanford's AI Lab, formalized this principle in 2021 when he launched the data-centric AI movement at NeurIPS. His argument was direct: for most practical applications, improving the data produces better results than improving the model.
Key figure
82%
ML projects that stall due to data quality issues (MIT)
The consequences of poor data quality are measurable. MIT researchers found an average of 3.4% label errors across ten commonly used machine learning benchmark datasets, including over 2,900 mislabeled images in ImageNet's validation set alone. Ng identified inconsistent labeling in 20 to 30 percent of computer vision datasets. These errors do not simply reduce accuracy. They introduce systematic biases that compound through training.
In healthcare AI, the effects can be stark. Diagnostic models trained on imbalanced datasets have missed 44% of liver disease cases in women, compared to 23% in men. Facial recognition systems have shown error rates up to 35% for dark-skinned women while dropping below 1% for white men, as documented in MIT Media Lab's Gender Shades project (Buolamwini and Gebru, 2018). In each case, the algorithm performed exactly as its training data instructed.
How it works
Data quality in AI is measured across several dimensions, now codified in the ISO/IEC 5259 standard series (published 2024). The standard defines five core characteristics: accuracy (whether values correctly represent reality), completeness (whether all necessary elements are present), consistency (whether data remains uniform across sources), timeliness (whether data reflects current conditions), and representativeness (whether the dataset reflects the real-world population the model will serve).
Key figure
5 parts
ISO/IEC 5259 standard series for AI data quality (2024)
Traditional software engineering adopted the phrase "garbage in, garbage out" decades ago. In machine learning, the principle operates with greater force. A 15% inaccuracy rate in training data can undermine model performance entirely. Unlike traditional software bugs, data quality problems often remain invisible until a model is deployed and begins producing biased or unreliable outputs.
Addressing data quality involves several stages: profiling (understanding what the data contains), cleaning (correcting errors and removing duplicates), augmentation (expanding datasets to improve representativeness), and monitoring (tracking quality over time as distributions shift). An estimated 91% of deployed ML models experience temporal degradation, where accuracy drops within months as real-world conditions diverge from training data.
Key context
The concept of data quality predates AI by decades. The phrase "garbage in, garbage out" appeared in computing literature as early as the late 1950s. But the field gained specific urgency for machine learning in the 2010s, as deep learning models began consuming datasets of unprecedented scale. When OpenAI trained GPT-3 on 570 GB of internet text in 2020, questions about what that text contained, and what it excluded, became unavoidable.
In 2021, Andrew Ng's data-centric AI workshops at NeurIPS marked a formal turning point. The workshops argued that the ML community's emphasis on model architecture had overshadowed the more fundamental question of data quality. As of 2024, ISO/IEC 5259 provides the first international standard framework for measuring and managing data quality specifically for AI and machine learning applications.
FAQ
What is the difference between data quality and data quantity in AI?
Data quantity is how much data you have. Data quality is how reliable, accurate, and representative that data is. A smaller, well-curated dataset often outperforms a larger, noisy one. Andrew Ng's data-centric AI research demonstrated this repeatedly across computer vision and natural language tasks.
Can AI itself fix data quality problems?
Partially. Automated tools can detect duplicates, flag outliers, and identify labeling inconsistencies. But determining whether data is representative of the real-world population, or whether labels reflect genuine ground truth, still requires human domain expertise. Fully automated data cleaning risks introducing new biases.
Why do data quality problems disproportionately affect certain groups?
Training datasets often overrepresent dominant demographic groups. When a facial recognition system trains primarily on light-skinned faces, it learns those features more precisely. The MIT Media Lab's Gender Shades project (Buolamwini and Gebru, 2018) documented error rate gaps of over 34 percentage points between demographic groups, tracing the disparity directly to training data composition.
How does the ISO/IEC 5259 standard help?
Published in 2024, ISO/IEC 5259 provides standardized definitions, measurement methods, and management guidelines for data quality in AI. It gives organizations a shared vocabulary and assessment framework, making it possible to audit and compare data quality practices across projects and industries.
Sources
- Primary: Data-Centric AI Competition, NeurIPS 2021 (Ng, A. et al., 2021)
- Primary: Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks (Northcutt, C.G., Athalye, A., Mueller, J., 2021)
- Standard: ISO/IEC 5259-1:2024, Artificial Intelligence: Data Quality for Analytics and Machine Learning
- Primary: Gender Shades: Intersectional Accuracy Disparities (Buolamwini, J. and Gebru, T., 2018)
- Survey: Data-Centric Artificial Intelligence: A Survey (Zha, D. et al., 2023)
Related Reading




Fact Check: Claim-by-Claim Verification Verified
All 11 claims verified. Key statistics on label error rates, demographic bias gaps, and ISO/IEC 5259 publication confirmed against primary sources.
Sources used for verification
- NeurIPS 2021 Data-Centric AI Workshop - neurips.cc
- Pervasive Label Errors in Test Sets - dl.acm.org
- ISO/IEC 5259-1:2024 - iso.org
- Gender Shades - proceedings.mlr.press
- Andrew Ng: Unbiggen AI - spectrum.ieee.org
