Researchers have identified a significant gap in how current metrics evaluate synthetic tabular data, particularly concerning the preservation of inter-column dependencies. Standard metrics often fail to detect when these crucial dependencies are lost, leading to an inaccurate assessment of data fidelity. The study introduces a new diagnostic tool, XGB-C2ST, which decomposes a classifier's two-sample test into marginal, dependency, and cross-component analyses. Applying this to the TabbyFlow/EF-VFM generator revealed a persistent 'dependency gap' that negatively impacts minority-class utility, a flaw missed by existing evaluation methods. AI
IMPACT Highlights a critical flaw in synthetic data evaluation, potentially impacting the reliability of AI models trained on such data.
RANK_REASON The item is an academic paper detailing a new method for evaluating synthetic data. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- Connected Papers
- DagsHub
- EF-VFM
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- ScienceCast
- scite Smart Citations
- TabbyFlow
- XGB-C2ST
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →