PulseAugur
EN
LIVE 14:53:31

The Stack dataset to exclude copyrighted code in future iterations

The Stack, a dataset of code used for training AI models, will exclude copyrighted code in its next iteration. Users who wish to remove already-ingested copyrighted code from the dataset can only pursue legal action or express dissatisfaction. AI

IMPACT Raises questions about data provenance and legal recourse for AI training datasets.

RANK_REASON User commentary on a dataset's policy regarding copyrighted code.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

The Stack dataset to exclude copyrighted code in future iterations

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    @ concretedog > This will exclude those repositories in *the next iteration* of The Stack. So if I want to remove the (copyrighted!) code they already have, I c

    @ concretedog > This will exclude those repositories in *the next iteration* of The Stack. So if I want to remove the (copyrighted!) code they already have, I can only sue and/or weep? # thestack # ai