PulseAugur
中
实时 20:09:42
English(EN) Databricks can't seem to shake authors' copyright claim that could result in 'extraordinary' damages

Databricks 因作者起诉其 LLM 训练数据而面临“巨额”版权赔偿

美国法官已允许针对 Databricks 的集体诉讼继续进行,该诉讼指控其 DBRX 大型语言模型使用了盗版的受版权保护的书籍进行训练。作者声称 Databricks 收购了 MosaicLM,而 MosaicLM 使用了包含约 196,000 种图书(包括他们的作品)的 RedPajama 数据集。Databricks 辩称作者无法证明 DBRX 是使用该特定数据训练的,但法官要求提供更多信息以确定是否发生了版权侵权。 AI

影响 版权侵权案件中可能产生的巨额赔偿,可能会影响 LLM 训练数据的获取策略。

排序理由 关于 LLM 训练数据版权侵权的集体诉讼正在进行中。

在 The Register — AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Databricks 因作者起诉其 LLM 训练数据而面临“巨额”版权赔偿

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
关于 LLM 训练数据版权侵权的集体诉讼正在进行中。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
policy, paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
162 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. The Register — AI TIER_1 English(EN) · O'Ryan Johnson ·

    Databricks 似乎无法摆脱作者的版权主张,该主张可能导致“巨额”赔偿

    <h4>Authors say it acquired an LLM that was trained on their copyrighted data, and judge keeps asking for more info</h4> <p>Databricks cannot shake a class action lawsuit targeting its LLM, which several book authors contend was created with a database that contained pirated vers…