Researchers have developed TrustMI, a method to causally control how large language model (LLM) assistants decide whether to trust users or third parties. By analyzing 2,000 contrastive conversations, they learned steering matrices that can be applied to frozen models to influence trust decisions. This approach was tested across six instruction-tuned models and showed that trust can be manipulated monotonically in both directions, impacting safety-related agent behaviors like harmful requests and prompt injections. AI
IMPACT This research offers a novel method for controlling LLM trust, potentially enhancing AI safety by mitigating risks associated with harmful requests and prompt injections.
RANK_REASON The cluster describes a research paper detailing a new method for controlling LLM trust. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →