Hugging Face has released The Stack v3, a new and expanded dataset for training large language models on code. This dataset is available in two versions: a pre-processed 'train' version optimized for direct use and a 'full' version containing the entire 114 TB corpus for users who wish to apply their own filtering and deduplication methods. The dataset aims to be the largest open collection of code available for AI development. AI
IMPACT Provides a massive, open-source dataset to accelerate the development of code-generating AI models.
RANK_REASON Release of a large dataset for AI model training. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →