Researchers have developed a novel method to identify which large language models (LLMs) are trained on data scraped from specific websites. The technique involves hosting dynamic websites that serve unique "canary tokens" to visiting web scrapers. By prompting LLMs and observing if they generate outputs containing these unique tokens, researchers can infer which LLMs have been exposed to data from those sites. This approach was demonstrated to reliably identify scrapers feeding 22 different LLM systems, including some not publicly disclosed by their companies, offering a way for third parties to gain insight into LLM data sourcing. AI
IMPACT This method could enable better control over unwanted web scraping for LLM training data, potentially influencing data acquisition strategies.
RANK_REASON The cluster describes a research paper detailing a novel technical method for identifying AI web scrapers. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- canary tokens
- Emily Wenger
- Hugging Face
- large language models
- robots.txt
- user agent
- web scrapers
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →