Researchers have introduced the concept of a "semantic bandit" to analyze how large language models (LLMs) explore and exploit options in decision-making tasks. This framework highlights that LLMs' exploration behavior is significantly influenced by semantic priors derived from their pre-training on language, rather than solely by the task's formal structure. The study found that semantically informative labels can lead to reduced exploration and increased exploitation, which can be detrimental if the semantic information is misaligned with the actual reward structure. Furthermore, LLMs tend to explore more in response to negative rewards than positive ones, suggesting biases inherent in their language-based training. AI
IMPACT Reveals inherent biases in LLM decision-making due to language pre-training, impacting reliability in real-world applications.
RANK_REASON Academic paper introducing a new theoretical framework for analyzing LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- large language models
- multi-armed bandit
- ScienceCast
- Semantic Bandits
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →