Researchers have introduced Diff Mining, a novel framework designed to identify the specific objectives and behaviors learned by language models during the finetuning process. This method compares the logits of a finetuned model against its base model to pinpoint salient tokens that indicate learned behaviors, even when these behaviors are unrelated to the finetuning domain. Diff Mining requires only access to output logits, making it scalable to large models, and can be used for tasks such as finetune domain detection and auditing tools to detect injected biases. AI
IMPACT Provides a new method for auditing and understanding the specific behaviors learned by language models post-finetuning.
RANK_REASON The cluster contains a research paper detailing a new method for analyzing language models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Diff Mining
- Gotit.pub
- Hugging Face
- Language Models
- non-negative matrix factorization
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →