Researchers have developed a novel two-stage framework for training compact instruction-following rerankers, combining off-policy teacher optimization with on-policy student distillation. The first stage strengthens a 4B teacher reranker using off-policy GRPO with LLM-judge feedback on 88K examples. The second stage involves a 1B student reranker sampling its own rankings and receiving soft rewards derived from the teacher's policy, improving performance particularly under distribution shift. This method achieved superior results on the MAIR-11 and MAIR-Full benchmarks compared to traditional distillation techniques and even surpassed released RL-trained rerankers. AI
IMPACT This research offers a more efficient way to train compact AI rerankers, potentially improving deployment capabilities and performance under distribution shifts.
RANK_REASON The cluster contains a research paper detailing a new method for training AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →