Two new research papers explore the use of Large Language Models (LLMs) to automate and standardize corpus linguistics tasks. The first paper introduces TACOMORE, a structured prompting framework designed to improve the accuracy and replicability of LLM analysis for keywords, collocates, and concordances, though it notes persistent hallucination issues. The second paper presents a large-scale pipeline for LLM-assisted corpus annotation, demonstrating high accuracy in analyzing grammatical variations and changes over time, which can unlock research questions previously out of reach. AI
IMPACT LLM frameworks and pipelines are emerging to automate complex linguistic analysis, potentially accelerating research and uncovering new patterns in language data.
RANK_REASON The cluster contains two academic papers published on arXiv detailing new methodologies for using LLMs in corpus linguistics research.
- arXiv
- Cameron Morin
- Corpus of Contemporary American English
- Corpus of Historical American English
- Hugging Face
- OpenAI API
- alphaXiv
- Bingru Li
- CatalyzeX
- corpus linguistics
- DagsHub
- Gotit.pub
- LLM
- ScienceCast
- TACOMORE
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →