A new paper explores how practitioners define, assess, and manage data quality in AI-driven systems, moving beyond traditional views of data as mere input. The research, based on interviews with 16 practitioners, identifies six key themes, including shifts in traceability from debugging to attributing model behavior and the introduction of circularity when models themselves assess quality. It also highlights concerns around authenticity with synthetic data and the critical role of lawfulness and representativeness in foundation model training data. AI
IMPACT This research reframes data quality as a critical engineering and organizational concern for AI systems, impacting trust and outcomes.
RANK_REASON The cluster contains a research paper detailing empirical findings on data quality in AI systems. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →