The Stack, a dataset of code used for training AI models, will exclude copyrighted code in its next iteration. Users who wish to remove already-ingested copyrighted code from the dataset can only pursue legal action or express dissatisfaction. AI
IMPACT Raises questions about data provenance and legal recourse for AI training datasets.
RANK_REASON User commentary on a dataset's policy regarding copyrighted code.
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →