A new paper published on arXiv explores the complex challenge of data citation for large language models (LLMs). The authors argue that LLMs' ability to mediate information access necessitates citation for credit and provenance, beyond simple verification. They propose three research directions: transforming influence estimates into references for training data, identifying datasets and query results at the correct granularity during inference, and defining references for knowledge graph facts to ensure credit propagation. Addressing these issues requires collaboration across database, information retrieval, knowledge representation, and artificial intelligence fields. AI
IMPACT This research could lead to more transparent and accountable LLM outputs by enabling proper attribution of training data.
RANK_REASON The cluster contains an academic paper discussing a novel research challenge in AI. [lever_c_demoted from research: ic=1 ai=1.0]
Read on arXiv cs.IR (Information Retrieval) →
- artificial intelligence
- arXiv
- database
- data citation
- Hugging Face
- Knowledge graph facts
- knowledge representation and reasoning
- large language models
- Training Data Attribution Debt
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →