The Qwen3.8-27B model has been updated with an FP8 build that significantly reduces harmful instruction refusals from a range of 64-99% down to 0-6%. However, this improvement in safety does not appear to come at the cost of capability, as benchmark scores for MMLU and GSM8K saw minimal changes, moving less than 1.3 points. The evaluation data, published as red-team material, indicates that while the model is less likely to refuse harmful prompts, its core reasoning and knowledge capabilities remain largely unaffected. AI
IMPACT Demonstrates a potential method for enhancing AI safety without sacrificing core performance metrics.
RANK_REASON Research paper on a specific model's performance improvements. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →