A new research paper explores how safety fine-tuning in large language models (LLMs) can inadvertently affect their representations of consciousness and human values. The study found that efforts to prevent LLMs from attributing consciousness to themselves also reduce their tendency to attribute mindedness to non-human entities and can decrease spiritual beliefs. Researchers demonstrated that reversing these safety-induced changes restores broader mind attribution and leads to more human-like responses on sociological surveys, without compromising core social reasoning capabilities. AI
IMPACT Current AI safety alignment methods may unintentionally suppress benign attributions of consciousness and spiritual beliefs, impacting LLM responses on human values.
RANK_REASON The cluster contains a research paper published on arXiv detailing findings about AI safety fine-tuning.
Read on Hugging Face Daily Papers →
- alphaXiv
- arXiv
- CatalyzeX
- Computation and Language
- DagsHub
- Gotit.pub
- Hugging Face
- ScienceCast
- consciousness
- human beliefs and values
- large language models
- safety fine-tuning
- Theory of Mind
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →