A new research paper investigates whether censorship from Chinese AI models transfers to distilled student models. The study found that a distilled version of GPT-OSS-120B, trained on outputs from the censored DeepSeek V4 Flash model, did not inherit the censorship regarding China-sensitive topics. Despite DeepSeek V4 Flash refusing to discuss topics like Uyghur labor transfer programs, the distilled model provided detailed information on these subjects. The research also detailed a self-distillation process that yielded similar results, with the model achieving high performance on financial reasoning tasks at a significantly lower cost. AI
IMPACT Suggests that censorship may not be an inherent trait transferable through distillation, potentially impacting how models are trained and deployed globally.
RANK_REASON Research paper detailing findings on AI model distillation and censorship transfer. [lever_c_demoted from research: ic=1 ai=1.0]
Read on HN — claude cli stories →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →