A new benchmark study on arXiv investigates the effectiveness of deep tabular generative models compared to simpler baselines. Across 8 public datasets subsampled from 200 to 20,000 rows, and 4 natively small clinical datasets, the research found that deep models rarely outperformed trivial baselines. The study also observed that model rankings were surprisingly stable at smaller training sizes, contrary to predictions, and that rankings on subsampled large datasets did not strongly correlate with those on natively small clinical datasets. All code, data, and results from the 2,220 runs are publicly available. AI
IMPACT Questions the utility of complex generative models for small tabular datasets, suggesting simpler methods may suffice.
RANK_REASON Academic paper published on arXiv detailing a benchmark study. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →