A new research paper explores the impact of supervised fine-tuning (SFT) on the behavioral diversity of large language models (LLMs) in sequential decision-making tasks. The study, conducted using variants of Tic Tac Toe, found that while SFT can improve action accuracy, it often leads to a premature collapse in action diversity. This phenomenon, termed narrow-support imitation, can hinder exploratory behavior in LLMs. The researchers suggest that action augmentation, which trains models on all optimal actions per state, could help mitigate this diversity collapse. AI
IMPACT Supervised fine-tuning may inadvertently limit LLM exploration and decision-making capabilities, suggesting a need for new training approaches.
RANK_REASON Research paper published on arXiv detailing findings about LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
- action accuracy
- action augmentation
- action diversity
- arXiv
- Large language models
- narrow-support imitation
- policy collapse
- Supervised fine-tuning
- Tic Tac Toe
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →