Researchers have developed SteerablePlex, a new method to improve control over full-duplex conversational models. These models, capable of simultaneous listening and speaking, often struggle with maintaining conversational scenarios as history grows. To address this, a new benchmark, SimIF-Bench, was created to evaluate instruction-following capabilities. The SteerablePlex approach utilizes a Group Reward-Decoupled Normalization Policy Optimization (GDPO) training recipe, enabling the models to adhere to textual instructions while preserving their turn-taking abilities. This enhanced control makes SteerablePlex a more reliable user simulator compared to existing open-source models and GPT-Realtime. AI
IMPACT Enhances controllability of full-duplex models, potentially improving user simulation and conversational AI applications.
RANK_REASON The cluster describes a new research paper introducing a novel method and benchmark for conversational AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- gpt-realtime
- Group Reward-Decoupled Normalization Policy Optimization
- Hugging Face
- SimIF-Bench
- SteerablePlex
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →