FreedomIntelligence has released HuatuoGPT-3-9B, a new medical large language model. This model is built upon Qwen3.5-9B and utilizes a novel training method called One-stage Policy Optimization (OnePO). OnePO allows the model to adapt to medical contexts in a single reinforcement learning stage, bypassing the need for initial supervised fine-tuning. AI
IMPACT Introduces a new training paradigm for domain-specific LLMs that could accelerate adaptation to specialized fields.
RANK_REASON Release of a new model with a novel training methodology. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →