A new study published on arXiv investigated the sublexical sensitivity of large language models (LLMs) by comparing their processing of Italian pseudowords to human behavior. The research found that LLMs showed less sensitivity to sublexical cues in pseudoword-only conditions compared to humans and even a simpler character-n-gram model like fastText. This suggests that current LLMs may not process language at the sublexical level in the same way humans do, with tokenization and training data coverage proposed as potential explanations. AI
IMPACT Suggests current LLMs may not process language at the sublexical level like humans, potentially impacting nuanced language understanding tasks.
RANK_REASON The cluster contains an academic paper detailing research findings on LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →