This research paper introduces modifications to Zipf's and Heaps' laws by developing models for the proportion of words that appear only once (hapaxes). The study assumes a standard urn model for word token sampling and posits that the hapax rate is a function of text length. Four such functions—constant, cancelation, linear, and logistic—were examined, with the logistic model showing the best fit for a sample of English texts. The paper also discusses the need for more complex mixture models for larger text corpora. AI
IMPACT Provides theoretical linguistic insights potentially relevant to natural language processing model development.
RANK_REASON Academic paper detailing theoretical corrections to linguistic laws. [lever_c_demoted from research: ic=1 ai=0.4]
- arXiv
- cancelation model
- constant model
- English
- Hapax Rate Models
- Heaps' law
- linear model
- logistic regression model
- Łukasz Dębowski
- urn problem
- Zipf's law
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →