Researchers have proposed a quadratic term correction to Heaps' Law, which traditionally describes the relationship between word types and tokens using a power-law function. While Heaps' Law is linear on a log-log scale, empirical observations show a slight concavity. The new quadratic model, tested on English novels, fits the type-token data more accurately with a linear coefficient slightly above 1 and a quadratic coefficient around -0.02. This formalism also offers a way to estimate curvature using a 'pseudo-variance' concept, though it may face numerical instability with large token counts. AI
RANK_REASON The cluster contains an academic paper detailing a new mathematical model for linguistic analysis. [lever_c_demoted from research: ic=1 ai=0.4]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →