Researchers have developed a new benchmark, ICWBench, to evaluate how well large language models follow in-context watermarking instructions. Their evaluation of 14 LLMs revealed that none could consistently achieve both high detectability and answer quality. To address this, they proposed a two-stage training method, self-distillation with logits perturbation (SDLP) and reinforcement learning, which significantly improved performance on Qwen3-14B and GPT-OSS-20B while maintaining response quality. AI
IMPACT This research introduces a new evaluation method for LLM watermarking and a training technique to improve instruction following, potentially enhancing the security and traceability of AI-generated content.
RANK_REASON The cluster contains an academic paper detailing a new benchmark and training method for LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →