A new benchmark called DogLM has revealed that large language models struggle to infer implicit user intent, particularly when it comes to interactive elements in generated content. Across 804 browser games generated by 17 different LLMs, the models rarely made background dogs interactive unless explicitly prompted to do so. Even with a hint to add enjoyable mechanics, only a small fraction of games featured any form of player-dog interaction, and even fewer allowed for actual petting. AI
IMPACT Highlights a gap in LLM's ability to understand and act on implicit user desires, suggesting a need for improved alignment and inferential capabilities in AI development.
RANK_REASON The item describes a new benchmark and its findings regarding LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →