A discussion on Reddit's r/LocalLLaMA subreddit posits that large language models (LLMs) are not built on proprietary technology but rather on vast amounts of proprietary data, which the poster characterizes as stolen intellectual property. The conversation, addressed to a 'Michael,' suggests that the underlying architecture of LLMs is open, but their training datasets are a significant point of contention regarding ownership and legality. AI
IMPACT Raises questions about the ethical and legal sourcing of data used to train large language models.
RANK_REASON The cluster consists of a single Reddit post discussing the nature of LLM training data, which falls under commentary.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →