PulseAugur
EN
LIVE 17:29:14

Hugging Face releases The Stack v3, largest open code dataset

Hugging Face has released The Stack v3, a new and expanded dataset for training large language models on code. This dataset is available in two versions: a pre-processed 'train' version optimized for direct use and a 'full' version containing the entire 114 TB corpus for users who wish to apply their own filtering and deduplication methods. The dataset aims to be the largest open collection of code available for AI development. AI

IMPACT Provides a massive, open-source dataset to accelerate the development of code-generating AI models.

RANK_REASON Release of a large dataset for AI model training. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Hugging Face releases The Stack v3, largest open code dataset

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Nunki08 ·

    Hugging Face releases The Stack v3 – largest open code dataset yet

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1v59aek/hugging_face_releases_the_stack_v3_largest_open/"> <img alt="Hugging Face releases The Stack v3 – largest open code dataset yet" src="https://preview.redd.it/wfsr7qg426fh1.jpg?width=140&amp;height=136&…