Researchers have introduced a new method called influence-guided response rewriting to improve training data attribution (TDA) for large language models. This technique uses influence functions to identify influential training examples and then rewrites their responses to align with desired behaviors, a method that proved more effective than simply reweighting the same examples. The study demonstrated that response rewriting leads to stronger and more consistent behavioral shifts in LLMs compared to traditional reweighting, even for safety-related tasks. AI
IMPACT This research could lead to more effective methods for understanding and controlling LLM behavior by improving how training data influences model outputs.
RANK_REASON The cluster contains a research paper detailing a new method for training data attribution in LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Functions
- large language models
- ScienceCast
- training data attribution
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →