Researchers have developed BehaviorTrace, an open evaluation harness to assess the attribution of training data in online reinforcement learning (RL) for language models. The study, conducted on Qwen2.5-1.5B using GRPO, found that many apparent attribution signals were actually confounds. A simple gradient-magnitude ranking method performed comparably to targeted estimators, and model fluency also proved to be a strong predictor of learned behavior. While per-rollout results varied significantly across seeds and generation draws, one consistent signal emerged: the gradient of trigger tokens aligned with the behavior's actual occurrence. AI
IMPACT Highlights limitations in attributing learned behaviors to specific training data in RL models, suggesting a need for more robust evaluation methods.
RANK_REASON Research paper detailing a new evaluation harness for attribution methods in RL. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →