Researchers at Hirundo have developed a method to remove political alignment and censorship from the Qwen3.6-35B-A3B large language model. Their technique, which involves training a LoRA adapter on specific examples and merging it with the base weights, reduced politically aligned or censored responses from 89.8% to 2.8% on their custom CCPC-500 benchmark. While general capabilities saw only a minor average change of 0.72 points, the model's ability to refuse to answer or provide biased information on sensitive topics was significantly diminished. The researchers note that this method differs from 'abliteration' by targeting specific behaviors without broadly impairing the model's refusal capabilities. AI
IMPACT Demonstrates a method for fine-tuning LLMs to remove specific ideological biases without significantly impacting general capabilities.
RANK_REASON Research paper detailing a method to remove political alignment from an open-weights LLM. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →