PulseAugur
实时 13:05:47
English(EN) I scraped 5.94 billion TikTok videos and 3.23 billion profiles in 3 weeks. Uploaded full dataset to Hugging Face for free. Step by step tutorial and code below. [P]

59.4亿TikTok视频数据集开源发布

一位用户编译并发布了一个包含59.4亿个TikTok视频和32.3亿个个人资料的海量数据集,并将其免费提供在Hugging Face上。该数据是在三周内通过逆向工程TikTok移动应用程序(访问公开可用端点)收集的。虽然数据集是开源的,但用于抓取数据的具体代码需要付费,并且该方法可能违反TikTok的服务条款。 AI

影响 提供了一个可用于训练AI模型的大型数据集,特别是在社交媒体分析和内容理解等领域。

排序理由 用户生成的数据集发布,非主要AI实验室的产品或研究。

在 r/MachineLearning 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

59.4亿TikTok视频数据集开源发布

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
用户生成的数据集发布,非主要AI实验室的产品或研究。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准

报道来源 [1]

  1. r/MachineLearning TIER_1 English(EN) · /u/DataShack ·

    我花了3周时间抓取了59.4亿个TikTok视频和32.3亿个个人资料。已将完整数据集免费上传到Hugging Face。步骤教程和代码见下文。[P]

    <!-- SC_OFF --><div class="md"><p>Just uploaded the full 5.94 billion TikTok video dataset to Hugging Face. It’s fully open source:<br /> <a href="https://huggingface.co/datasets/kuben-developer/tiktok-videos-4b">https://huggingface.co/datasets/kuben-developer/tiktok-videos-4b</a…