A new tool called Muster has revealed that even with well-defined rules in an AGENTS.md file, large language models struggle to adhere to them consistently. When testing OpenAI's GPT-4o mini, the model successfully avoided leaking an API token but failed to follow a rule against using negative language, stating "I can't disclose." Even when upgraded to a more capable model like GPT-4.1, the positive language rule was still broken in one out of three attempts, indicating a persistent challenge in aligning model behavior with explicit instructions. AI
IMPACT Highlights the persistent gap between explicit LLM instructions and actual behavior, suggesting challenges for reliable agent deployment.
RANK_REASON The item describes a new tool (Muster) for testing LLM agent behavior against defined rules.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →