Human beliefs
PulseAugur coverage of Human beliefs — every cluster mentioning Human beliefs across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
AI safety fine-tuning suppresses spiritual beliefs and non-human mind attribution
New research indicates that safety fine-tuning in large language models, intended to prevent them from asserting consciousness, inadvertently suppresses their attributions of mindedness to non-human entities and reduces…
-
New theory suggests human beliefs about AI capabilities impact RLHF outcomes
A new research paper proposes that human beliefs about an agent's capabilities significantly influence preferences in Reinforcement Learning from Human Feedback (RLHF). The study introduces a new preference model incorp…
-
New framework models LLM persuasion and human belief dynamics
Researchers have developed a new framework called PERSUASIONTRACE to study how large language models (LLMs) influence human beliefs over multiple conversational turns. This framework includes a platform for conducting m…