PulseAugur
实时 09:49:47
English(EN) [D] How do you get preprocessed dataset of a paper [D]

研究人员难以获取预处理数据集以实现论文可复现性

Reddit r/MachineLearning 版块的一名用户正在寻求关于如何从研究论文中获取预处理数据集的建议,因为他们自己尝试复现报告的统计数据未能成功。他们已联系作者但未获回复,不确定是继续使用自己可复现但数据不匹配的数据集、抽样以匹配报告的大小,还是将问题升级给期刊。用户正在寻找从研究人员那里获取预处理文件的有效策略和礼仪。 AI

影响 凸显了AI研究可复现性方面的挑战,影响了已发表研究结果的可靠性。

排序理由 用户生成内容,讨论常见的研究可复现性问题,而非主要公告。

在 r/MachineLearning 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

研究人员难以获取预处理数据集以实现论文可复现性

本文如何被排名

Signal score
3 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Commentary
用户生成内容,讨论常见的研究可复现性问题,而非主要公告。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. r/MachineLearning TIER_1 English(EN) · /u/Individual-Safety906 ·

    [D] 如何获取论文的预处理数据集 [D]

    <!-- SC_OFF --><div class="md"><p>Hi all,</p> <p>I'm trying to reproduce a paper where the reported dataset statistics in Table 1 don't match what I get from the public raw data, even after implementing the preprocessing exactly as described.</p> <p>I've tried all reasonable inte…