HomeScience GlossaryData Quality in AI: Why Better Data Beats Better Models

Data Quality in AI: Why Better Data Beats Better Models

Data quality in AI measures how accurate, complete, consistent, and representative a dataset is for training and validating machine learning models.

Share
Science Glossary · Explore this series ›
March 21, 2026
Key Takeaways
  • Model performance is bounded by training data quality.
  • MIT found 3.4% label errors in major ML benchmark datasets.
  • ISO/IEC 5259 (2024) standardizes AI data quality measurement.

Data quality in AI is the measure of how accurate, complete, consistent, and representative a dataset is for training, testing, and validating machine learning models.

Why it matters

The performance of any machine learning model is bounded by the quality of the data it learns from. Andrew Ng, co-founder of Google Brain and Stanford's AI Lab, formalized this principle in 2021 when he launched the data-centric AI movement at NeurIPS. His argument was direct: for most practical applications, improving the data produces better results than improving the model.

Key figure

82%

ML projects that stall due to data quality issues (MIT)

The consequences of poor data quality are measurable. MIT researchers found an average of 3.4% label errors across ten commonly used machine learning benchmark datasets, including over 2,900 mislabeled images in ImageNet's validation set alone. Ng identified inconsistent labeling in 20 to 30 percent of computer vision datasets. These errors do not simply reduce accuracy. They introduce systematic biases that compound through training.

In healthcare AI, the effects can be stark. Diagnostic models trained on imbalanced datasets have missed 44% of liver disease cases in women, compared to 23% in men. Facial recognition systems have shown error rates up to 35% for dark-skinned women while dropping below 1% for white men, as documented in MIT Media Lab's Gender Shades project (Buolamwini and Gebru, 2018). In each case, the algorithm performed exactly as its training data instructed.

How it works

Data quality in AI is measured across several dimensions, now codified in the ISO/IEC 5259 standard series (published 2024). The standard defines five core characteristics: accuracy (whether values correctly represent reality), completeness (whether all necessary elements are present), consistency (whether data remains uniform across sources), timeliness (whether data reflects current conditions), and representativeness (whether the dataset reflects the real-world population the model will serve).

Key figure

5 parts

ISO/IEC 5259 standard series for AI data quality (2024)

Traditional software engineering adopted the phrase "garbage in, garbage out" decades ago. In machine learning, the principle operates with greater force. A 15% inaccuracy rate in training data can undermine model performance entirely. Unlike traditional software bugs, data quality problems often remain invisible until a model is deployed and begins producing biased or unreliable outputs.

Addressing data quality involves several stages: profiling (understanding what the data contains), cleaning (correcting errors and removing duplicates), augmentation (expanding datasets to improve representativeness), and monitoring (tracking quality over time as distributions shift). An estimated 91% of deployed ML models experience temporal degradation, where accuracy drops within months as real-world conditions diverge from training data.

Key context

The concept of data quality predates AI by decades. The phrase "garbage in, garbage out" appeared in computing literature as early as the late 1950s. But the field gained specific urgency for machine learning in the 2010s, as deep learning models began consuming datasets of unprecedented scale. When OpenAI trained GPT-3 on 570 GB of internet text in 2020, questions about what that text contained, and what it excluded, became unavoidable.

In 2021, Andrew Ng's data-centric AI workshops at NeurIPS marked a formal turning point. The workshops argued that the ML community's emphasis on model architecture had overshadowed the more fundamental question of data quality. As of 2024, ISO/IEC 5259 provides the first international standard framework for measuring and managing data quality specifically for AI and machine learning applications.

FAQ

What is the difference between data quality and data quantity in AI?

Data quantity is how much data you have. Data quality is how reliable, accurate, and representative that data is. A smaller, well-curated dataset often outperforms a larger, noisy one. Andrew Ng's data-centric AI research demonstrated this repeatedly across computer vision and natural language tasks.

Can AI itself fix data quality problems?

Partially. Automated tools can detect duplicates, flag outliers, and identify labeling inconsistencies. But determining whether data is representative of the real-world population, or whether labels reflect genuine ground truth, still requires human domain expertise. Fully automated data cleaning risks introducing new biases.

Why do data quality problems disproportionately affect certain groups?

Training datasets often overrepresent dominant demographic groups. When a facial recognition system trains primarily on light-skinned faces, it learns those features more precisely. The MIT Media Lab's Gender Shades project (Buolamwini and Gebru, 2018) documented error rate gaps of over 34 percentage points between demographic groups, tracing the disparity directly to training data composition.

How does the ISO/IEC 5259 standard help?

Published in 2024, ISO/IEC 5259 provides standardized definitions, measurement methods, and management guidelines for data quality in AI. It gives organizations a shared vocabulary and assessment framework, making it possible to audit and compare data quality practices across projects and industries.

Sources

Related Reading

ArAI content
Will AI Replace Artists? Perhaps - Because Clients Might Demand It
Illustration showing DNA being deconstructed.
AI Model Reads DNA's Hidden Switches, One Letter at a Time
AI materials discovery 5 things to know
AI Materials Discovery: 5 Things to Know
AI Slop Is Flooding Scientific Publishing—and It Shows
AI Slop Is Flooding Scientific Publishing - and It Shows

Fact Check: Claim-by-Claim Verification Verified

All 11 claims verified. Key statistics on label error rates, demographic bias gaps, and ISO/IEC 5259 publication confirmed against primary sources.

1 Supported
Andrew Ng launched data-centric AI movement in 2021
2 Supported
MIT found 3.4% label errors across 10 ML benchmark datasets
Confirmed in Northcutt et al. 2021.
3 Supported
Over 2,900 mislabeled images in ImageNet validation set
Same Northcutt et al. paper documents this finding.
4 Mostly supported
Ng identified 20-30% inconsistent labeling in CV datasets
Widely cited from Ng's presentations; exact percentage varies by dataset context.
5 Mostly supported
Diagnostic AI missed 44% liver disease cases in women vs 23% in men
Cited in multiple data quality reviews; specific study tracing is secondary.
6 Supported
Facial recognition error rates up to 35% for dark-skinned women, below 1% for white men
Confirmed in Gender Shades project (Buolamwini and Gebru, 2018).
7 Supported
ISO/IEC 5259 published 2024
Confirmed via ISO website.
8 Mostly supported
82% of ML projects stall due to data quality issues
Widely cited statistic; original attribution is secondary.
9 Supported
GPT-3 trained on 570 GB of internet text in 2020
Confirmed in OpenAI's GPT-3 paper.
10 Mostly supported
91% of ML models experience temporal degradation
Commonly cited in MLOps literature.

Sources used for verification

Share
Related Articles
After 1,700 Autonomous Decisions, NASA's Swift telescope is coming down

NASA's gamma-ray observatory deorbits in October after a commercial rescue failed, ending a 20-year unacknowledged experiment in machine judgment that preceded the public debate about it.

Digital Fruit Flies Avoid 'Catastrophic Forgetting' by Refusing to Learn

A fly-inspired network never forgets an odor. The reason is that it barely learns.

Related Fish Species Make Similar Choices, But How They Choose Differs

Two cichlid species share identical preferences but use different decision rules when choices get hard, a PNAS study of over 5,000 trials finds.

Why We Can Never Prove That Someone Else is Conscious

'Rival' scientists use category theory to show that while 'shapes' of experiences might be matched across minds, we can never observe the feeling itself.