Researchers from ESTS have detailed their submissions to the WMT26 Model Compression Shared Task, focusing on English-to-Simplified Chinese and English-to-Egyptian Arabic translation. Their approach involved pruning experts from GPT-OSS-20B based on task-specific routing mass and cross-lingual routing divergence, then fine-tuning the remaining specialists with synthetic data generated by GPT-5.1. Further compression was achieved using MXFP4 quantization on the retained expert weights, resulting in models ranging from 4.186B to 7.770B parameters. AI
IMPACT This research demonstrates advanced techniques for compressing large language models, potentially enabling more efficient deployment of translation systems.
RANK_REASON The item is a research paper detailing methods for model compression submitted to a shared task. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →