Recent evaluations of Chain-of-Thought (CoT) controllability in large language models reveal that current frontier models, including OpenAI's GPT-5.5 and Anthropic's Fable 5, perform poorly on tasks requiring adherence to specific formatting constraints within their reasoning processes. These models typically score between 0-30% on such evaluations. However, further experimentation suggests that prompt engineering can significantly improve performance, with open-source models showing a 2-3x increase in accuracy. This indicates that the CoT controllability evaluations may be under-elicited, and current metrics might not fully represent models' actual capabilities in obscuring their reasoning. AI
IMPACT Prompt engineering can significantly improve LLM controllability, suggesting current safety evaluations may be too simplistic.
RANK_REASON The item discusses an evaluation methodology for LLMs and suggests improvements, fitting the research category. [lever_c_demoted from research: ic=1 ai=1.0]
- Anthropic
- Claude 3
- Claude Opus 4.6
- Constitutional AI
- CoTControl eval
- Gemini
- GPT-4
- GPT-5.5
- GPT-OSS-120B
- Mythos Preview
- OpenAI
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →