PulseAugur
EN
LIVE 12:49:48

LLMs streamline corpus linguistics with new analysis frameworks

Two new research papers explore the use of Large Language Models (LLMs) to automate and standardize corpus linguistics tasks. The first paper introduces TACOMORE, a structured prompting framework designed to improve the accuracy and replicability of LLM analysis for keywords, collocates, and concordances, though it notes persistent hallucination issues. The second paper presents a large-scale pipeline for LLM-assisted corpus annotation, demonstrating high accuracy in analyzing grammatical variations and changes over time, which can unlock research questions previously out of reach. AI

IMPACT LLM frameworks and pipelines are emerging to automate complex linguistic analysis, potentially accelerating research and uncovering new patterns in language data.

RANK_REASON The cluster contains two academic papers published on arXiv detailing new methodologies for using LLMs in corpus linguistics research.

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

LLMs streamline corpus linguistics with new analysis frameworks

COVERAGE [2]

  1. arXiv cs.CL TIER_1 English(EN) · Bingru Li, Han Wang, Nicholas Groom ·

    TACOMORE: Exploring a replicable prompting protocol for LLM-assisted corpus analysis

    arXiv:2412.10139v2 Announce Type: replace Abstract: As corpus linguistics continues to scale, researchers are facing a growing methodological bottleneck: while computational tools can easily count billions of words, the qualitative interpretation of these data remains a slow and …

  2. arXiv cs.CL TIER_1 English(EN) · Cameron Morin, Matti Marttinen Larsson ·

    A large-scale pipeline for LLM-assisted corpus annotation: variation and change in the English consider construction

    arXiv:2510.12306v3 Announce Type: replace Abstract: As natural language corpora expand at an unprecedented rate, manual annotation remains a significant methodological bottleneck in corpus linguistic work. We address this challenge by presenting a scalable pipeline for automating…