A new research paper published on arXiv explores the concept of model coherence, specifically focusing on how narrow finetunes can lead to self-contradiction. The study introduces a set of 175 questions designed to reveal these contradictions, which are difficult to attribute to simple ambiguity or indifference. The findings indicate that even models with high specificity scores exhibit significant incoherence, including issues like identity conflation and introspection failures, suggesting that narrow finetuning may limit the models' ability to exhibit coherent misaligned behavior. AI
IMPACT Highlights potential limitations in current AI finetuning methods and their impact on model reliability.
RANK_REASON Academic paper published on arXiv detailing a new method for evaluating model coherence. [lever_c_demoted from research: ic=1 ai=1.0]
- An Investigation of Model Coherence: Narrow Finetunes Contradict Themselves Under Resampling
- arXiv
- Hugging Face
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →