An art project encoding sentences by word length revealed that large language models (LLMs) do not necessarily recover the intended meaning from such encoded messages. Contrary to initial assumptions, independent readings of the same encoded message showed agreement above chance, but this agreement was no higher than readings of different messages. The experiment highlighted three ways a baseline can mislead when measuring model agreement, particularly in pipelines that rely on self-consistency or majority-vote ensembles. AI
IMPACT Suggests LLM interpretation may be more about projection and bias than true understanding of encoded information.
RANK_REASON The item discusses an experiment on LLM behavior and agreement, which falls under commentary on AI capabilities rather than a direct release or research milestone.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →