PulseAugur
EN
LIVE 11:30:53

New research tackles multilingual models, efficient inference, and data contamination

Recent research explores various facets of language model development and application. Google DeepMind's ATLAS project introduces new scaling laws for multilingual models, aiming to optimize training for languages beyond English. Hugging Face has released LFM2.5-Encoders, designed for efficient long-context inference on CPUs, outperforming existing models in speed. Several arXiv papers delve into model behavior and training: TransClean addresses the issue of "translation noise" in LLM outputs, while another study investigates membership evidence in language models by analyzing duplication counts. Further research examines the fragility of models under recursive training, the impact of data anonymization on performance, and methods for identifying language-specific neurons in multilingual models. Additionally, new approaches like SWRouter are proposed for multi-turn dialogue routing, and studies explore data-efficient language modeling through structural priors and principle-guided improvements. AI

IMPACT These advancements offer improved efficiency for multilingual tasks, faster inference on common hardware, and deeper insights into model behavior and training methodologies.

RANK_REASON The cluster consists of multiple research papers published on arXiv and a blog post detailing a new model release, focusing on technical advancements and studies in language modeling.

Read on Google AI / Research →

AI-generated summary · Google Gemini · from 1241 sources. How we write summaries →

New research tackles multilingual models, efficient inference, and data contamination

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
The cluster consists of multiple research papers published on arXiv and a blog post detailing a new model release, focusing on technical advancements and studies in language modeling.
Source corroboration
1241 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
241 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+99 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [1241]

  1. Google AI / Research TIER_1 English(EN) ·

    ATLAS: Practical scaling laws for multilingual models

    Generative AI

  2. Hugging Face Blog TIER_1 English(EN) ·

    LFM2.5-Encoders for Fast Long-Context Inference on CPU

  3. arXiv cs.CL TIER_1 English(EN) · Daniela Gottesman, Alon Gilae-Dotan, Ido Cohen, Yoav Gur-Arieh, Marius Mosbach, Ori Yoran, Mor Geva ·

    LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations

    arXiv:2509.03405v2 Announce Type: replace Abstract: Language models (LMs) increasingly drive real-world applications that require world knowledge. However, the internal processes through which models turn data into representations of knowledge and beliefs about the world are poor…

  4. arXiv cs.CL TIER_1 English(EN) · Makoto Fukushima ·

    Message capacity and claim wording set the transition points of collective truth-finding in language-model networks

    arXiv:2609.19183v1 Announce Type: cross Abstract: Whether human or large language model (LLM), an agent in a discussion reads only a few of the others' contributions, bounded by cognition, context, or cost. LLM collectives can settle on a wrong consensus even when a majority star…

  5. arXiv cs.CL TIER_1 English(EN) · Haibo Feng, Ruiqi Liang, Hanyang Peng, Shiqi Yu ·

    On-Demand Attention: Language Models Know When to Recall

    arXiv:2609.20734v1 Announce Type: new Abstract: Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, regardless of its benefit to the next prediction. We show that a pretrained m…

  6. arXiv cs.CL TIER_1 English(EN) · Levent Bulut ·

    Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models --- A Conceptual Framework and Registered Test Protocol

    arXiv:2609.20712v1 Announce Type: new Abstract: This paper introduces and operationalizes summarization bias: a proposed systematic tendency of large language models (LLMs) to represent narrative meaning as an abstract summary label rather than as the reconstructable inferential …

  7. arXiv cs.CL TIER_1 English(EN) · Adrian Brasoveanu, Ece Takmaz, Jakub Dotla\v{c}il ·

    Relational Attention for Data-Efficient Language Modeling

    arXiv:2609.20530v1 Announce Type: new Abstract: We present Relational BabyLM, a system submission to the BabyLM 2026 challenge that combines two cognitively motivated inductive biases in a single decoder-only Transformer. Architecturally, we replace standard self-attention with a…

  8. arXiv cs.CL TIER_1 English(EN) · Giuseppe Samo, Vivi Nastase, Paola Merlo ·

    Generalization through Lexical Abstraction in Transformer Models: The Case of Functional Words

    arXiv:2609.19887v1 Announce Type: new Abstract: Pronouns, adverbs and other functional words (such as they, her, somewhere, there) are often used in language to replace concrete nouns or phrases, when their properties - such as gender, grammatical number - provide sufficient info…

  9. arXiv cs.CL TIER_1 English(EN) · Shuo Cai, Yanggan Gu, Zihao Wang, Yuanyi Wang, Yibo Yan, Wenjun Wang, Yuhang Liu, Guanghao Zhu, Sirui Huang, Ming Li, Hongxia Yang ·

    From Parameters to Behaviors: A Survey of Model Fusion for Large Language Models

    arXiv:2609.19553v1 Announce Type: new Abstract: Model fusion integrates the capabilities from source models into a single target model. As of June 2026, Hugging Face hosts more than 2M models. This growing pool provides a rich base for model reuse and capability integration. Yet …

  10. arXiv cs.AI TIER_1 English(EN) · Alexander Shirnin, Aleksey Kudelya ·

    For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances

    arXiv:2609.19504v1 Announce Type: cross Abstract: As model-generated content is increasingly consumed by other model instances in automated workflows, a practically important question arises: can a model embed a signal in natural language that an independent instance of the same …

  11. arXiv cs.AI TIER_1 English(EN) · Pritish Mishra, Ishaan Kumar, Akshat Mandoli, Sudarshan Kamath ·

    MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

    arXiv:2609.20152v1 Announce Type: new Abstract: Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller's audio, a language model reads the transcript and decides what to say and which backend tools to call, and a TTS model speaks the reply. N…

  12. arXiv cs.AI TIER_1 English(EN) · Canfer Akbulut, Justine Breuch, Arianna Manzini, Lujain Ibrahim, Matija Franklin, Roma Patel, Iason Gabriel, Kristian Lum, Laura Weidinger ·

    Tailored to you: longitudinal effects of personalising language models

    arXiv:2609.20077v1 Announce Type: new Abstract: Interest in developing personalised language models is rapidly growing. While personalisation is often viewed as a mechanism to better serve diverse user needs, the effects of sustained interactions with personalised models on peopl…

  13. arXiv cs.AI TIER_1 English(EN) · Maxim Chupilkin ·

    Geopolitical Divisions Across Languages in Large Language Models

    arXiv:2609.20005v1 Announce Type: new Abstract: People increasingly turn to AI chatbots for news and explanations of world events. But do they receive the same political answers when they ask in different languages? Here we show that the language of a question can change how the …

  14. arXiv cs.AI TIER_1 English(EN) · Ha Van Dau, Thanh Tung Khuat, Nguyen Thanh Dung ·

    Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models

    arXiv:2609.19934v1 Announce Type: new Abstract: Depth-recurrent language models iteratively apply a small layer stack, decoupling per-token compute from distinct parameter count. To determine whether such a model genuinely utilizes its depth, both recurrence and layer-pruning lit…

  15. arXiv cs.AI TIER_1 English(EN) · Shaina Raza, Ahmed Y. Radwan, Imran Liaquat, Kathryn Hume ·

    A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems

    arXiv:2609.19524v1 Announce Type: new Abstract: Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern artificial intelligence systems. Large language models (LLMs), agentic systems, and multimodal models (MLLMs) require different forms of …

  16. arXiv cs.AI TIER_1 English(EN) · Davide Vitabile, N. Ranjan, Akshay Nambiar, Kamal K. Gupta, Amril Nazir ·

    QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training

    arXiv:2609.19513v1 Announce Type: new Abstract: High-quality pre-training data is a critical bottleneck for educational and STEM-specific language models targeting edge AI and on-device deployment where token budgets are tightly constrained. While major organizations train ever-l…

  17. arXiv cs.AI TIER_1 English(EN) · Alizishaan Khatri, Chiquita Prabhu, Omkar Neogi ·

    Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

    arXiv:2609.19472v1 Announce Type: new Abstract: Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployme…

  18. arXiv cs.AI TIER_1 English(EN) · Yu He, Yingxi Li, Yifei Wang, Ellen Vitercik ·

    Compositional Reasoning in Language Models under Reinforcement Learning Post-Training

    arXiv:2609.19465v1 Announce Type: new Abstract: Compositional reasoning is critical for real-world problem solving: since training data is necessarily limited, models must generalize by composing learned skills in new ways. While post-training methods such as reinforcement learni…

  19. arXiv cs.AI TIER_1 Italiano(IT) · Chao Huang, Meng Tong, Kejiang Chen ·

    Fingerprinting Multimodal Large Language Models

    arXiv:2609.20457v1 Announce Type: cross Abstract: While multimodal large language models (MLLMs) enable a wide range of image-text reasoning tasks, recent incidents indicate that they are vulnerable to illicit deployment and unauthorized distillation. Existing solutions for model…

  20. arXiv cs.CL TIER_1 English(EN) · Wendy K. Tam ·

    The Neutral Mask: How Alignment Training Provides Shallow Alignment while Leaving Partisan Structure Intact in a Large Language Model

    arXiv:2606.09735v2 Announce Type: replace Abstract: The ambition behind alignment training is to make large language models safe and useful. The primary mechanisms, reinforcement learning from human feedback (RLHF) and its direct-optimization variants, shape the behavior of deplo…

  21. Hugging Face Daily Papers TIER_1 English(EN) ·

    On-Demand Attention: Language Models Know When to Recall

    Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, regardless of its benefit to the next prediction. We show that a pretrained model's decoding states already contain informati…

  22. arXiv cs.AI TIER_1 English(EN) · Tyler McDonald, Ali Emami ·

    WordPolo: Evaluating Language Models Through Iterative Semantic Feedback

    arXiv:2609.19006v1 Announce Type: cross Abstract: Large Language Models (LLMs) and Large Reasoning Models (LRMs) are typically evaluated on challenging benchmarks through dataset accuracy alone, providing no insight into the quality or faithfulness of their reasoning processes. W…

  23. arXiv cs.LG TIER_1 English(EN) · Jiahui Chen, Bingke Zhu, Hongyu Pan, Yingying Chen ·

    WaveTLM: Reliable Time-Series Language Modeling through Task Compilation

    arXiv:2609.18812v1 Announce Type: new Abstract: Time-series language models provide a shared natural-language interface across temporal tasks, but plausible text does not guarantee reliable task outputs. Responses may appear reasonable while hallucinating the required object: num…

  24. arXiv cs.CL TIER_1 English(EN) · Yanxu Mao, Peipei Liu, Tiehan Cui, Zhaoteng Yan, Congying Liu, Datao You ·

    Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models

    arXiv:2412.16555v4 Announce Type: replace Abstract: Large language models (LLMs) are widely applied in various fields of society due to their powerful reasoning, understanding, and generation capabilities. However, the security issues associated with these models are becoming inc…

  25. arXiv cs.CL TIER_1 English(EN) · Nicolas Bourrel, Abderrahmane Issam, Gerasimos Spanakis ·

    A Probe Shift Is Not a Fairness Fix: The Limits of Representation Steering in Speech Models

    arXiv:2609.18533v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representations. We ask whether speaker-linked attributes that are linearly readable from pretr…

  26. arXiv cs.CL TIER_1 English(EN) · M\'aty\'as Osv\'ath, Enik\H{o} H\'eja, No\'emi Ligeti-Nagy ·

    Made in Hungary: Comments on the performance of generative language models

    arXiv:2609.18284v1 Announce Type: new Abstract: In recent years, three initiatives have emerged to develop generative language models in Hungary. The motivation behind them is the same. For Hungarian, no model with the given capability existed, or existing English-centric models …

  27. arXiv cs.CL TIER_1 English(EN) · Simran Koul ·

    Register Bias in Complexity-Based Large Language Model Routing

    arXiv:2609.17542v1 Announce Type: new Abstract: Large language model services increasingly route each query to one of several models of differing capability, using a cheap estimate of query complexity to send easy queries to small models and hard queries to large ones. I show tha…

  28. arXiv cs.CL TIER_1 English(EN) · Divyansh Agarwal ·

    Relation Before Entity: Deferred Commitment in Language Model Factual Recall

    arXiv:2609.17537v1 Announce Type: new Abstract: We ask whether relation-type information (e.g., capital-of) and entity-specific information (e.g., France to Paris) become causally active at the final-token position at the same depth during recall. Using four complementary causal …

  29. arXiv cs.CL TIER_1 English(EN) · Marco Simoni, Aleksandar Fontana, Giulio Rossolini, Andrea Saracino ·

    DANTINOX: A Unified Framework for Multi-Paradigm Language Modeling

    arXiv:2609.17535v1 Announce Type: new Abstract: Language generation research increasingly spans three paradigms: autoregressive decoding, discrete masked diffusion, and continuous flow-matching. Comparing them is difficult because each lives in a separate codebase, so measured di…

  30. arXiv cs.AI TIER_1 English(EN) · Aliakbar Nafar, Kristen Brent Venable, Zijun Cui, Parisa Kordjamshidi ·

    Extracting Probabilistic Knowledge from Large Language Models for Bayesian Network Parameterization

    arXiv:2505.15918v3 Announce Type: replace-cross Abstract: In this work, we evaluate the potential of Large Language Models (LLMs) in building Bayesian Networks (BNs) by approximating domain expert priors. LLMs have demonstrated potential as factual knowledge bases; however, their…

  31. arXiv cs.AI TIER_1 (CA) · Alexander Didenko, Anna Shabanova, Vladislav Zapylikhin, Alexander Antipov, Ruslana Raemgulova ·

    GYROval: A Robust Benchmark for Cultural Value Orientation in Large Language Models

    arXiv:2609.18384v1 Announce Type: cross Abstract: We present a robust benchmark for measuring cultural value orientation in large language models on the two Inglehart-Welzel axes over several domains and roles (hence GYROval - Gridded Yielding of Robust value Orientation), togeth…

  32. arXiv cs.AI TIER_1 English(EN) · Chowdhury Mohammad Abdullah, Rita Orji ·

    From a River in Gilead to the Inference Distributions of Large Language Models: Covert Dialect Bias and Linguistic Profiling at Scale

    arXiv:2609.18068v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains such as housing screening. While alignment techniques mitigate explicit racial bias in generated text, they often leave covert attitudinal associations …

  33. arXiv cs.AI TIER_1 English(EN) · Chandimal Adikari, Nandika Herath ·

    An Empirical Evaluation of Cost-Efficient Large Language Models on Algorithmic Programming Tasks

    arXiv:2609.18052v1 Announce Type: cross Abstract: This study empirically evaluates whether cost-efficient Large Language Models (LLMs) can be trusted to generate enterprise code to a written specification. Three models (Gemini Flash 3, GPT-5.4 mini and Claude Haiku 4.5) were aske…

  34. arXiv cs.AI TIER_1 English(EN) · Alexandre Quemy ·

    Function Lives Where Variance Doesn't: Task-Weighted Charts of a Language Model's Computation

    arXiv:2609.18989v1 Announce Type: new Abstract: How many dimensions does a language model's computation actually use? The question is ill-posed until one names a functional. Task-weighted charts make it well-posed: low-dimensional coordinate systems fit against a chosen functiona…

  35. arXiv cs.AI TIER_1 English(EN) · Youjia Wang, Lin Xu, Yang Sun, Yuxiao Lu, Chengfang Fang, Jie Shi ·

    Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models

    arXiv:2609.18515v1 Announce Type: new Abstract: Safety alignment teaches large language models (LLMs) to recognize harmful requests and reject risky instructions. Yet aligned models can fail when harmful intent is concealed within seemingly benign contexts. Robust safety therefor…

  36. Hugging Face Daily Papers TIER_1 English(EN) ·

    A Probe Shift Is Not a Fairness Fix: The Limits of Representation Steering in Speech Models

    Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representations. We ask whether speaker-linked attributes that are linearly readable from pretrained ASR encoders yield useful directions for r…

  37. arXiv cs.AI TIER_1 English(EN) · Gautam Kishore ·

    Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation

    arXiv:2609.16145v1 Announce Type: new Abstract: We study a practical question: can a small correction module fix errors in a frozen language model's outputs without degrading its base capabilities? We propose CRN v2, a lightweight logit-level correction module (~34M trainable par…

  38. arXiv cs.AI TIER_1 English(EN) · Keqing Zhang, Jingyu Chen, Yufan Liu, Yongqiang Zhu, Nai Ding, Lai Jiang, Congyan Lang, Bing Li, Weiming Hu ·

    Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models

    arXiv:2609.16589v1 Announce Type: new Abstract: As Large Language Models (LLMs) increasingly handle complex subjective tasks, aligning their intentions and behaviors with human values has become a critical scientific challenge. However, current efforts are confounded by a strikin…

  39. arXiv cs.AI TIER_1 English(EN) · Carel van Niekerk, Renato Vukovic, Benjamin Ruppik, Hsien-chin Lin, Shutong Feng, Milica Ga\v{s}i\'c ·

    Post-Training Large Language Models via Reinforcement Learning from Self-Feedback

    arXiv:2507.21931v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) often produce plausible but poorly-calibrated answers, limiting their reliability on reasoning-intensive tasks. Recent research suggests that Chain-of-Thought (CoT) reasoning paths are inherent…

  40. arXiv cs.AI TIER_1 English(EN) · Danny Wood, James Stringer ·

    Permutation-Based Stegomalware in Large Language Models: Threats and Countermeasures

    arXiv:2609.16193v1 Announce Type: cross Abstract: The difficulty of training large language models (LLMs), together with their ubiquity, raises the threat of stegomalware, where malicious payloads are embedded into model weights. Recent work has demonstrated the use of permutatio…

  41. arXiv cs.CL TIER_1 English(EN) · Daniel Balcells, Andrew Jun Lee, Chirag Rastogi, Paul M. Riechers, Adam Shai, Xavier Poncini ·

    Large Language Models Develop Belief State Geometry In-Context

    arXiv:2609.17376v1 Announce Type: cross Abstract: Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood. We consider such representations in a control…

  42. arXiv cs.CL TIER_1 English(EN) · Eduardo Novaes Hering ·

    Persistent Recurrent Memory Between Transformer Layers - Improves Language Model Generalization

    arXiv:2609.17251v1 Announce Type: new Abstract: We introduce a simple architectural modification to decoder-only transformers: a persistent recurrent state that observes hidden representations via cross-attention, updates itself through a GRU, and modulates subsequent processing …

  43. arXiv cs.CL TIER_1 English(EN) · James A. Michaelov, Carmen Amo Alonso, Tyler A. Chang, Roger P. Levy ·

    Target-Language Generation in Multilingual Models: Activation Steering and Optimal Control

    arXiv:2609.16967v1 Announce Type: new Abstract: Ensuring that multilingual language models generate coherent text in a specific target language is a major issue in multilingual language modeling. We develop an optimal control method for target-language text generation as well as …

  44. arXiv cs.CL TIER_1 English(EN) · Qingchen Yu, Shiying Duan, Xiaodong Li, Yuhua Wang, Zhiyu Li, Shiji Zhou, Yifan Sun, Zhaoxin Fan ·

    Cascade: Hierarchical Recoverability Control for Large Language Model Unlearning

    arXiv:2609.16890v1 Announce Type: new Abstract: Large Language Model (LLM) unlearning is essential for removing sensitive or copyrighted knowledge while preserving general utility. Existing methods often leave residual knowledge in intermediate representations, which can still be…

  45. arXiv cs.AI TIER_1 English(EN) · Zhenbin Wang, Lei Zhang, Lituan Wang, Wei Huang, Yan Wang, Zhenwei Zhang ·

    Layers, Sinks, and Scaling: Adaptive Evidence Selection for Multimodal Large Language Models

    arXiv:2609.16795v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) can answer knowledge-intensive visual questions by combining visual evidence from images with facts retrieved from external sources. However, MLLMs may overlook relevant evidence in both moda…

  46. arXiv cs.CL TIER_1 English(EN) · Albert Ge, Chandan Singh, Yufan Zhuang, Xiaodong Liu, Jianfeng Gao, Frederic Sala ·

    Register Tokens for Bounded-State Reasoning in Diffusion Language Models

    arXiv:2609.16372v1 Announce Type: new Abstract: Masked diffusion language models (dLLMs) generate text by iteratively denoising masked tokens with bidirectional attention. Extending reasoning across generation chunks normally requires keeping earlier generated text in context. We…

  47. arXiv cs.CL TIER_1 English(EN) · Tabia Tanzin Prama, Calla Glavin Beauregard, Christopher M. Danforth, Peter Sheridan Dodds ·

    Self-reported archetypes and behavioral failures in Large Language Models

    arXiv:2609.15998v1 Announce Type: new Abstract: Every large language model (LLM) has behavioral traits and moral preferences that comprise its character. Whether by design or as an emergent property of training, these systems exhibit persistent dispositions that shape how they in…

  48. arXiv cs.CL TIER_1 English(EN) · Foivos Charalampakos, Md Ibrahim Ibne Alam, Iordanis Koutsopoulos, Koushik Kar ·

    Optimal Model Activation Policies for Inference Networks of Large Language Models

    arXiv:2609.15992v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have rendered them necessary for Natural Language Processing (NLP) tasks, and their high inference cost motivates the study of cost-performance trade-offs. In practice, several expert …

  49. Hugging Face Daily Papers TIER_1 English(EN) ·

    Large Language Models Develop Belief State Geometry In-Context

    Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood. We consider such representations in a controlled setting: prompting LLMs with data emitted from…

  50. arXiv cs.AI TIER_1 English(EN) · Behzad Shayegh, Niloofar Kazemi ·

    Mind Which Bird You Favour: Parameterizing Adequacy-Fluency Balance in Meta-Evaluation of Machine Translation

    arXiv:2609.14795v1 Announce Type: cross Abstract: There is a tradeoff in machine translation meta-evaluation between prioritizing alignment with adequacy versus fluency. The balance depends on the combination of translation systems in the meta-evaluation dataset. This system set …

  51. arXiv cs.LG TIER_1 English(EN) · Sophia Tang, Shiyi Wang ·

    Discrete Beckmann Transport Models for One-Step Language Modeling and Reasoning

    arXiv:2609.15903v1 Announce Type: new Abstract: Discrete diffusion and flow models are a promising alternative to autoregressive language models, but compressing many-step sampling into fewer steps typically requires distilling a pretrained teacher model. This caps the student at…

  52. arXiv cs.CL TIER_1 English(EN) · Daxin Tan, Dehua Tao, Chengxi Deng, Hanlin Zhang, Xiao Chen ·

    Reducing the Output-Mode Gap in Speech Language Models via Joint-Output On-Policy Distillation

    arXiv:2609.15313v1 Announce Type: cross Abstract: Autoregressive generation of interleaved text and acoustic tokens is a common approach to spoken-response generation in speech large language models. Although this design enables streaming generation with explicit textual guidance…

  53. arXiv cs.CL TIER_1 English(EN) · Pengyu Ji, Zichen Zhang, Xiang Hu, Kewei Tu ·

    DA-DLM: Explicitly Modeling Token Dependencies in Diffusion Language Models

    arXiv:2609.15070v1 Announce Type: new Abstract: Diffusion Language Models (DLMs) generate text by iteratively denoising a masked sequence, independently predicting multiple tokens at each step. This conditional independence discards inter-token dependencies and degrades coherence…

  54. arXiv cs.CL TIER_1 English(EN) · Maxim Topaz, Zhihong Zhang, Nir Roguin, Pallavi Gupta, Zichao Li, Laura-Maria Peltonen ·

    Biomedical Reference Generation Remains Unreliable across 26 Large Language Models

    arXiv:2609.14988v1 Announce Type: new Abstract: Background. Large language models are increasingly used to help write biomedical text but may fabricate references to nonexistent work. How often large language models do so is not well characterized. Methods. We prompted 26 languag…

  55. arXiv cs.CL TIER_1 English(EN) · Mingze Yin, Xiaohan Wang, Dian Li, Haichao Yao, Yilin Zhao, Youjun Chen, Gang Liu, Jintai Chen, Yiheng Zhu, Chang-Yu Hsieh, Aimin Pan ·

    Func-R1: Incentivizing Mathematical Function Reasoning in Multimodal Large Language Models

    arXiv:2609.14779v1 Announce Type: new Abstract: Performing deliberate mathematical reasoning in visual contexts is a hallmark of advanced Multimodal Large Language Models (MLLMs) and requires a sophisticated synthesis of perceptual grounding and symbolic logic. However, in the re…

  56. arXiv cs.CL TIER_1 English(EN) · Ju-Chieh Chou, Jiawei Zhou, Karen Livescu ·

    Quantifying the Generation Modality Gap in Speech-Text Language Models

    arXiv:2609.14743v1 Announce Type: new Abstract: Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify because speech and text systems are typically evaluated with different metrics a…

  57. arXiv cs.CL TIER_1 English(EN) · Elliot Murphy ·

    Formal Properties of Language as Constraints on Neural Dynamics

    arXiv:2609.14384v1 Announce Type: new Abstract: What must a neural system be capable of to implement language? Current research annotates stimuli with linguistic variables and tests which electrodes, voxels, or language-model layers predict neural activity. Yet predictive success…

  58. arXiv cs.CL TIER_1 English(EN) · Kyle Richardson, Cullen Anderson, Pranav Balakrishnan, Takuto Ban, Daksha Ladia, Ankita Gupta, Marisa Hudspeth ·

    From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models

    arXiv:2609.13520v1 Announce Type: new Abstract: While Large Language Models have improved rapidly, many fundamental questions remain about how to evaluate the knowledge and reasoning abilities they acquire, and how such evaluations relate to the learning signals used in pre-train…

  59. arXiv cs.CL TIER_1 English(EN) · Shamin Chokshi ·

    Lexical Prompt Compression for Large Language Models: A Training-Free, Deterministic Pipeline with Empirical Pareto Analysis Across Eleven Task Categories

    arXiv:2609.13154v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have made prompts increasingly large and complex. Techniques such as chain-of-thought reasoning (Wei et al., 2022) and in-context learning (Brown et al., 2020) frequently push real-wor…

  60. arXiv cs.AI TIER_1 English(EN) · Naihao Deng, Alissa Shen, Yiming Feng, Joan Nwatu, Jae-Won Chung, Mosharaf Chowdhury, Yulong Chen, Rada Mihalcea ·

    The Language-Energy Divide: Measuring Energy Costs of Multilingual LLM Inference

    arXiv:2606.21869v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed in multilingual settings, yet the energy costs of serving these models across different languages remain poorly understood. We present a systematic study of inference …

  61. arXiv cs.AI TIER_1 English(EN) · Chengze Li, Yitong Zhang, Jia Li, Liyi Cai, Ge Li ·

    Exploring the Potential of Diffusion Large Language Models in Code Generation

    arXiv:2509.11252v3 Announce Type: replace-cross Abstract: LLMs have become the mainstream approaches to code generation. Existing LLMs mainly employ autoregressive generation, i.e. generating code token-by-token from left to right. However, the underlying autoregressive generatio…

  62. arXiv cs.AI TIER_1 English(EN) · Kang Chen, Xiuze Zhou, Yuanhui Yu, Yuanguo Lin, Hefeng Chen, Congyu Cai, Li Shen ·

    Data Security in Large Language Models: Risks, Defense, and Directions

    arXiv:2508.02312v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs), now a foundation in advancing natural language processing, power applications such as text generation, machine translation, and conversational systems. Despite their transformative potential, …

  63. arXiv cs.AI TIER_1 English(EN) · Harshit Dhankhar, Baban Gain, Asif Ekbal, Yogesh Mani Tripathi ·

    Balancing Global Quality and Pronoun-Specific Feedback for Context-Aware Machine Translation

    arXiv:2501.03008v2 Announce Type: replace-cross Abstract: Context-aware machine translation can expose the evidence needed for pronoun choice, but standard fine-tuning does not explicitly prioritize these sparse discourse-sensitive decisions. We study ProNMT, a reward-guided iter…

  64. arXiv cs.AI TIER_1 English(EN) · Wenbo Zhang, Aditya Majumdar, Asif Ekbal, Amulya Yadav ·

    CHAI for LLMs: Improving Code-Mixed Translation in Large Language Models through Reinforcement Learning with AI Feedback

    arXiv:2411.09073v4 Announce Type: replace-cross Abstract: Large language models (LLMs) show strong performance across many tasks but remain weak at understanding code-mixed (CM) language. Despite this limitation, improving LLMs for CM tasks has received little attention. To addre…

  65. arXiv cs.AI TIER_1 English(EN) · Peipei Li, Dongsen Zhang, Yuchen Liu, Wenjun Xu ·

    Dynamic Semantic Compression for Efficient Latent-Space Inference in Large Language Models

    arXiv:2609.15338v1 Announce Type: cross Abstract: Large Language Models (LLMs) primarily perform inference at the token level, resulting in substantial memory overhead and compromised computational efficiency. In this paper, we propose a Dynamic Semantic Extraction and Inference …

  66. arXiv cs.AI TIER_1 English(EN) · Quang Phuoc Nguyen, F\'elix Gaschi, David Anugraha, Santiago Mart\'inez Novoa, En-Shiun Annie Lee ·

    Online Language Adaptive Sampling for Better Distributed Cross-lingual Gains

    arXiv:2609.14969v1 Announce Type: cross Abstract: Realignment is a promising approach for improving the cross-lingual transfer ability of multilingual language models, particularly for extremely low-resource languages (LRLs). However, existing realignment methods rely on uniform …

  67. arXiv cs.AI TIER_1 English(EN) · Cheikh Ahmed ·

    Enemray: Toward Capable Language Models for Hassaniya

    arXiv:2609.14829v1 Announce Type: cross Abstract: We introduce Enemray, a Hassaniya-centric language model that enables general-purpose interaction in Hassaniya. Enemray is trained around a stability--plasticity objective: acquire strong Hassaniya linguistic and cultural competen…

  68. arXiv cs.AI TIER_1 English(EN) · Suzannah E McKinney, Phuc Vu, Samuel A Justice, Christopher Humphries, Alyssa Pradhan, Timothy J Keyes, Bernardo C Bizzo, Keith J Dreyer, Sarah F Mercaldo, James M Hillis ·

    A primer on evaluation methods for large language models in healthcare

    arXiv:2609.14819v1 Announce Type: cross Abstract: Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and not harm. This evaluation can be more challenging than traditional machine learn…

  69. arXiv cs.AI TIER_1 English(EN) · Xiaobing Shi, Zherui Li, Yiming Jiang, Kun Wang, Yufei Guo ·

    Towards Evolving Context Parameterization for Large Language Models

    arXiv:2609.14168v1 Announce Type: cross Abstract: Context parameterization enables large language models (LLMs) to internalize contexts into reusable model parameters, avoiding repeated processing across subsequent queries. However, existing methods typically assume static contex…

  70. arXiv cs.AI TIER_1 English(EN) · Prajjwal Bhattarai, Tuka Alhanai ·

    Signatures of Steerability in Activation Space of Language Models

    arXiv:2609.14151v1 Announce Type: cross Abstract: Steering language models using a set of contrastive representations has been a canonical and computationally efficient method for controlling model behavior. Despite this success in controlling certain model behaviors, the effecti…

  71. arXiv cs.AI TIER_1 English(EN) · Nawar S. Alseelawi, Mustafa S. Aljumaily ·

    Mizan: A National Benchmark for Evaluating Large Language Models on Iraqi Arabic and the Iraqi Civic Context

    arXiv:2609.13980v1 Announce Type: cross Abstract: Arabic large-language-model (LLM) evaluation has matured around Modern Standard Arabic (MSA): aggregated leaderboards such as the Open Arabic LLM Leaderboard (OALL), HELM Arabic, and BALSAM rank models across dozens of MSA tasks, …

  72. arXiv cs.AI TIER_1 English(EN) · Tingqi Wang, Hongyu Ke, Haoxin Wang, Rafal Angryk, Zhipeng Cai ·

    Task-Aware Federated Fine-Tuning for MoE-based Large Language Models

    arXiv:2609.13395v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) has become a widely adopted architecture for Large Language Models (LLMs), as it improves model capacity while limiting computational overhead through sparse expert activation. This property makes MoE-base…

  73. arXiv cs.AI TIER_1 English(EN) · Anas Dorbani, Sunny Yasser, Jimmy Lin, Amine Mhedhbi ·

    Beyond Quacking: Deep Integration of Language Models and RAG into DuckDB

    arXiv:2504.01157v1 Announce Type: cross Abstract: Knowledge-intensive analytical applications retrieve context from both structured tabular data and unstructured, text-free documents for effective decision-making. Large language models (LLMs) have made it significantly easier to …

  74. arXiv cs.AI TIER_1 English(EN) · Tiancheng Zhang, Yulin Chen, Yunfeng Zhao, Shaoyuan Huang, Cheng Zhang, Xiaofei Wang ·

    MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Model Serving

    arXiv:2609.15359v1 Announce Type: new Abstract: The surge of large language model (LLM) applications on personal devices imposes massive, bursty workloads on cloud serving infrastructure. While prefill-decode disaggregation improves throughput and scalability, memory-bound decode…

  75. arXiv cs.AI TIER_1 English(EN) · Tian Jin ·

    Self-Orchestrating Language Models: Leveraging Semantic Dependence for Efficient Inference

    arXiv:2609.14850v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate impressive capabilities, but their deployment presents significant efficiency challenges. Autoregressive decoding imposes substantial inference latency and under-utilizes hardware accelerator…

  76. arXiv cs.AI TIER_1 English(EN) · Chengxi Zhang, Yu Yao ·

    OrchSLM: Probing the Dynamics of Small Language Model Orchestration

    arXiv:2609.13470v1 Announce Type: new Abstract: Although large language models (LLMs) have demonstrated remarkable capabilities, their reliance on cloud-scale infrastructure poses fundamental challenges for deployment in agentic pipelines, including latency, privacy, connectivity…

  77. arXiv cs.AI TIER_1 English(EN) · Sudarshan Regmi, Arvind Pillai, Yu Yvonne Wu, Yuliang Chen, Bibek Panthi, Tess Z. Griffin, Michael V. Heinz, Lisa Marsch, Nicholas C. Jacobson, Andrew Campbell ·

    TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Models

    arXiv:2609.13457v1 Announce Type: new Abstract: Timeseries multimodal large language models (TS-MLLMs) have recently begun leveraging the reasoning capabilities of large language models (LLMs) for question-answering tasks. However, these models often fail to capture dynamic tempo…

  78. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Gautam Kishore ·

    Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation

    We study a practical question: can a small correction module fix errors in a frozen language model's outputs without degrading its base capabilities? We propose CRN v2, a lightweight logit-level correction module (~34M trainable parameters, 0.73% of the 4.65B text module) that si…

  79. arXiv cs.CL TIER_1 English(EN) · Praveen Kumar Myakala, Ravichandra Namburi, Sowmya Keragodu Jayaramu, Sooraj George Thomas ·

    SynthSentry: Detecting Synthetic Data Contamination in Language Model Training Data

    arXiv:2609.12353v1 Announce Type: new Abstract: Large language models trained recursively on their own or other models' outputs undergo model collapse, in which distributional tails and factual accuracy deteriorate while fluency survives. Prior work diagnoses collapse after train…

  80. arXiv cs.LG TIER_1 English(EN) · Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee ·

    Dead Weights, Live Signals: Feedforward Graphs of Frozen Language Models

    arXiv:2604.08335v2 Announce Type: replace Abstract: We present a feedforward graph architecture in which heterogeneous frozen large language models serve as computational nodes, communicating through a shared continuous latent space via learned linear projections. Building on rec…

  81. arXiv cs.LG TIER_1 English(EN) · Arthur Negr\~ao, Pedro Silva, Vander L. S. Freitas, Gladston Moreira, Eduardo Luz ·

    Benford's Law as a Distributional Prior for Post-Training Quantization of Large Language Models

    arXiv:2602.00165v2 Announce Type: replace Abstract: Post-training quantization (PTQ) is a practical way to reduce the memory footprint of large language models, but low-bit quantization is sensitive to mismatches between the quantization codebook and the empirical weight/activati…

  82. arXiv cs.LG TIER_1 English(EN) · Erfan Loghmani ·

    Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective

    arXiv:2506.00152v2 Announce Type: replace Abstract: Large language models are being widely used across industries to generate text that contributes directly to key performance metrics, such as medication adherence in patient messaging and conversion rates in content generation. P…

  83. arXiv cs.LG TIER_1 English(EN) · Zhendong Mi, Shaoyi Huang ·

    Sampling via Decision-Flow: Training-Free Extraction of Improved Latent Reasoning Paths in Large Language Models

    arXiv:2609.12317v1 Announce Type: new Abstract: A central question in LLM reasoning is whether reinforcement learning (RL) instills genuinely new capabilities or merely reshapes how existing knowledge is expressed during inference. Building on the distribution-sharpening hypothes…

  84. arXiv cs.CL TIER_1 English(EN) · Idriss Malek, Omar Choukrani, Daniil Orel, Anh Duy Le Dinh, Zhuohan Xie, Zangir Iklassov, Martin Tak\'a\v{c}, Salem Lahlou ·

    LLM-BabyBench: Can Language Models Plan in Worlds They Can Simulate?

    arXiv:2505.12135v2 Announce Type: replace-cross Abstract: When an interactive benchmark reports a single success rate for a language-model agent, it is rarely clear what that number measures. A failure can come from perception, ambiguous instructions, retrieval, missing commonsen…

  85. arXiv cs.CL TIER_1 English(EN) · Yilin Geng, Omri Abend, Eduard Hovy, Lea Frermann ·

    Measuring Pragmatic Influence in Large Language Model Instructions

    arXiv:2602.21223v2 Announce Type: replace Abstract: It is not only what we ask large language models (LLMs) to do that matters, but also how we ask them. Phrases like ``This is urgent'' or ``As your supervisor'' can shift model behavior without altering task content. We study thi…

  86. arXiv cs.CL TIER_1 English(EN) · S{\l}awomir Dadas, Rafa{\l} Po\'swiata, Ma{\l}gorzata Gr\k{e}bowiec, Micha{\l} Pere{\l}kiewicz ·

    Parameter-Efficient Retrievers for Polish and European Languages

    arXiv:2609.12913v1 Announce Type: new Abstract: Dense retrieval systems increasingly rely on multi-billion-parameter language models, whose memory and computational requirements make large-scale indexing, frequent corpus updates, and low-latency serving costly. We present a three…

  87. arXiv cs.CL TIER_1 English(EN) · Utkarsh Soni, Syed Shariyar Murtaza, Yifan Nie, Sachin Chandrasekhar, Eugene Wen ·

    Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

    arXiv:2609.13005v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and recent benchmarks suggest that they are increasingly adept at multi-hop reasoning. However, these benchmarks are typically …

  88. arXiv cs.CL TIER_1 English(EN) · Zhongzhan Huang, Junxin Li, Guoming Ling, Yupei Lin, Shanshan Zhong, Hefeng Wu ·

    Zipbench: Low-Cost Framework for Compressing Comprehensive Benchmarks of Large Language Models

    arXiv:2609.12475v1 Announce Type: new Abstract: Comprehensive benchmark suites are essential for improving large language models (LLMs), but many widely used benchmarks are redundant, making evaluation unnecessarily expensive. Although recent benchmark compression methods (BCMs) …

  89. Hugging Face Daily Papers TIER_1 English(EN) ·

    Register Tokens for Bounded-State Reasoning in Diffusion Language Models

    Masked diffusion language models (dLLMs) generate text by iteratively denoising masked tokens with bidirectional attention. Extending reasoning across generation chunks normally requires keeping earlier generated text in context. We ask whether a dLLM can instead continue reasoni…

  90. arXiv cs.AI TIER_1 English(EN) · Narcis Marincat ·

    Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

    arXiv:2609.11365v1 Announce Type: new Abstract: In shared-genome language-model societies, restricted evidence visibility favors reusable, value-indexed latent packet interfaces, whereas the sole high-performing globally visible model in the parent study learned an episode-entang…

  91. arXiv cs.CL TIER_1 English(EN) · Tobias Deu{\ss}er, Max Hahnb\"uck, Lorenz Sparrenberg, Tobias Uelwer, Christian Bauckhage, Rafet Sifa ·

    On the Impact of Anonymization on the Performance of Large Language Models

    arXiv:2609.11335v1 Announce Type: new Abstract: As large language models are increasingly deployed in sensitive domains, anonymizing input data to protect personally identifiable information has become a critical practice. However, the impact of this anonymization on model utilit…

  92. arXiv cs.CL TIER_1 English(EN) · Yangze Liu, Zhongyi Han ·

    A Fragility Spectrum for Recursive Language-Model Training

    arXiv:2609.11149v1 Announce Type: new Abstract: Model-generated text is finding its way back into training corpora, and there is plenty of evidence that training on such data over and over collapses output diversity. Prior work has studied the phenomenon itself: which protocols a…

  93. arXiv cs.CL TIER_1 English(EN) · Minjun Kim, Inho Won, Junghun Yuk, Dongyeon Kim, Jihyo Kim, KyungTae Lim ·

    Distribution-aware Language Neuron Identification in Multilingual Large Language Models

    arXiv:2609.10993v1 Announce Type: new Abstract: Multilingual large language models (mLLMs) contain a small fraction of feed-forward neurons that are sensitive to particular languages, commonly termed language-specific neurons. Existing work measures language specificity using the…

  94. arXiv cs.CL TIER_1 English(EN) · Daria Cherniuk, Alexander Rudikov, Boris Kashin, Ivan Oseledets ·

    Structured Transforms for Low-Overhead Quantization of Language Models

    arXiv:2609.11687v1 Announce Type: new Abstract: We revisit Kashin-decomposition-based weight quantization for large language models and propose an improved algorithm with stronger convergence properties and structured, efficient orthogonal transforms. The method retains the core …

  95. arXiv cs.CL TIER_1 English(EN) · Yana Veitsman, Jonas Mayer Martins, Jonathan Lautenschlager, Lisa Beinborn ·

    Structural priors for data-efficient language learning

    arXiv:2609.11505v1 Announce Type: new Abstract: Efficient language learning requires methods to reduce the reliance on large data and computational resources. We investigate structural transfer: First training models on non-language data to induce useful priors for natural langua…

  96. arXiv cs.CL TIER_1 English(EN) · Yu Wang, Yuchen Li, Rui Kong, Xinran Chen, Jiamin Chen, Hengyi Cai, Shuaiqiang Wang, Jiashu Zhao, Yulun Zhang, Zhonghao Lyu, Haoyi Xiong, Linghe Kong, Jimmy Xiangji Huang, Dawei Yin ·

    SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations

    arXiv:2609.11414v1 Announce Type: new Abstract: Large language models exhibit complementary strengths, motivating routing methods that dispatch each query to the most suitable model. Although existing routers are effective in single-turn settings, they do not directly transfer to…

  97. arXiv cs.CL TIER_1 English(EN) · Shenbin Qian, Yves Scherrer ·

    TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model Outputs

    arXiv:2609.11399v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for machine translation, yet their outputs often contain additional text beyond the translation itself, such as language labels, explanations or bilingual repetitions, which we term…

  98. arXiv cs.CL TIER_1 English(EN) · Shuxing Yang, Kaihao Zhu, Junjie Yang, Rui Zhao, Junyao Wu, Yize Wang, Wenhao Li, Fujia Chen, Taowen Deng, Shenzhan Hong, Yaqi Li, Zichen Li, Jincheng Mi, Yuang Pan, Hongsheng Chen, Yihao Yang ·

    Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement

    arXiv:2609.10702v1 Announce Type: new Abstract: Learning from limited text requires models to use context, generalize to new inputs, and retain useful capabilities. Qiushi Engine conducted a long-horizon, end-to-end autonomous research program on BabyLM 2026 Strict-Small, within …

  99. arXiv cs.LG TIER_1 English(EN) · Xiaoyu Li, Andi Han, Jiaojiao Jiang, Junbin Gao ·

    Characterizing Language Generation in the Limit: Finite Witnesses and a Separation-Width Hierarchy

    arXiv:2609.10525v2 Announce Type: replace-cross Abstract: Language generation in the limit asks for valid unseen elements from every exhaustive positive presentation of an unknown infinite language. We characterize this task for arbitrary families over a countable universe. Gener…

  100. arXiv cs.LG TIER_1 English(EN) · Aditya Somasundaram, Charles Guille-Escuret, Alexander Moreno, Zhengzhong Liu, Eric Xing ·

    Toward a First-Principles Update Geometry for the Language-Model Head

    arXiv:2608.22253v2 Announce Type: replace Abstract: Muon motivates designing optimizer geometry around the function of each parameter block and uses the spectral norm for hidden linear layers. For the language-model head, the spectral norm is not a faithful measure of functional …

  101. arXiv cs.LG TIER_1 English(EN) · Xuan Cuong Ngo, Hao Vo, Ngan Le ·

    GEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models

    arXiv:2609.10658v1 Announce Type: new Abstract: Activation steering provides a lightweight way to control large language models (LLMs) by modifying their hidden activations at inference time. Among these approaches, norm-preserving steering aims to change model behavior without a…

  102. arXiv cs.CL TIER_1 English(EN) · William Hager, Ishika Rathi, Masum Hasan, Cameron Jones ·

    Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue

    arXiv:2606.21844v2 Announce Type: replace Abstract: As AI systems integrate into online spaces, differentiating them from humans in conversations is increasingly important. We present Inverse Turing Bench, a benchmark that evaluates LLMs and other models on their ability to diffe…

  103. arXiv cs.CL TIER_1 English(EN) · Ni Yang, Rui He, Philipp Homan, Iris Sommer, Davide Staub, Wolfram Hinzen ·

    Cross-lingual brain-language model alignment is robust but challenges hierarchical and computational accounts

    arXiv:2605.21049v2 Announce Type: replace Abstract: Brain-language model alignment is often interpreted as evidence that transformer models implement computations similar to those of the human brain. This assumes that neural predictivity reflects internal computational properties…

  104. arXiv cs.CL TIER_1 English(EN) · Niharika Hegde, Subarnaduti Paul, Lars Joel-Frey, Manuel Brack, Kristian Kersting, Martin Mundt, Patrick Schramowski ·

    CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models

    arXiv:2509.22360v2 Announce Type: replace Abstract: Large language models (LLMs) excel at operating at scale by leveraging social media and various data crawled from the web. Whereas existing corpora are diverse, their frequent lack of long-term temporal structure may however lim…

  105. arXiv cs.CL TIER_1 English(EN) · Zhongxiang Sun ·

    A Short Survey of Viewing Large Language Models in Legal Aspect

    arXiv:2303.09136v2 Announce Type: replace Abstract: Large language models (LLMs) have transformed many fields, including natural language processing, computer vision, and reinforcement learning. These models have also made a significant impact in the field of law, where they are …

  106. arXiv cs.CL TIER_1 English(EN) · Sankar Behera, Dhruv Singh, Anshika Agnihotri, Raj Kumar Choudhary, Satyadev Ahlawat, Yamuna Prasad ·

    LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

    arXiv:2609.11163v1 Announce Type: cross Abstract: Structured pruning of large language models (LLMs) offers hardware-efficient compression, yet existing methods require calibration data, gradient computation, or large auxiliary policy networks at pruning time. LILA (\emph{Latent-…

  107. arXiv cs.CL TIER_1 English(EN) · Arman Nik Khah ·

    Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

    arXiv:2609.10830v1 Announce Type: new Abstract: When a language model finds a sentence unusually cheap to predict, it is tempting to conclude that the sentence was in its training data. Almost every published test of that inference has had to guess which sentences were in the tra…

  108. arXiv cs.CL TIER_1 English(EN) · Dario Picozzi ·

    The information geometry of large language models is shared, learned, and controllable

    arXiv:2609.11063v1 Announce Type: cross Abstract: Large language models learn similar behaviours, yet it remains unclear what structure they share or how to change one behaviour without disturbing others. The Fisher-Rao geometry of next-token probabilities connects these question…

  109. arXiv cs.CL TIER_1 English(EN) · Dongfang Zhao ·

    LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation

    arXiv:2609.11739v1 Announce Type: new Abstract: Large language model serving costs scale directly with output sequence length, yet standard preference alignment often inflates response verbosity without improving utility. We study whether the parameterization of post-training upd…

  110. Hugging Face Daily Papers TIER_1 English(EN) ·

    LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation

    Large language model serving costs scale directly with output sequence length, yet standard preference alignment often inflates response verbosity without improving utility. We study whether the parameterization of post-training updates affects generation length: low-rank subspac…

  111. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Dawei Yin ·

    SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations

    Large language models exhibit complementary strengths, motivating routing methods that dispatch each query to the most suitable model. Although existing routers are effective in single-turn settings, they do not directly transfer to multi-turn dialogue, where routing performance …

  112. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Narcis Marincat ·

    Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

    In shared-genome language-model societies, restricted evidence visibility favors reusable, value-indexed latent packet interfaces, whereas the sole high-performing globally visible model in the parent study learned an episode-entangled code. This companion study asks whether inde…

  113. Hugging Face Daily Papers TIER_1 English(EN) ·

    Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

    In shared-genome language-model societies, restricted evidence visibility favors reusable, value-indexed latent packet interfaces, whereas the sole high-performing globally visible model in the parent study learned an episode-entangled code. This companion study asks whether inde…

  114. arXiv cs.AI TIER_1 Deutsch(DE) · Alex Smolin, Bryan Wilder ·

    Beliefs and Behavior in Language Models

    arXiv:2609.07943v1 Announce Type: new Abstract: There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In addition to the inherent scientific interest of this question, these latent quant…

  115. arXiv cs.LG TIER_1 English(EN) · Talia Sternberg, Gallil Maimon, Yossi Adi ·

    Interleaved Speech Language Models Latently Work In Text

    arXiv:2606.22473v2 Announce Type: replace-cross Abstract: Speech language models (SLMs) increasingly combine speech and text, often by interleaving their tokens within a single sequence. Yet how these two modalities interact in the model's latent space remains unclear. In this wo…

  116. arXiv cs.LG TIER_1 English(EN) · Zichun Yu, Chenyan Xiong ·

    RePro: Training Language Models to Faithfully Recycle the Web for Pretraining

    arXiv:2510.10681v2 Announce Type: replace-cross Abstract: High-quality data is a cornerstone of large language model (LLM) pretraining, yet its growth has not kept pace with the needs of frontier models. In this paper, we introduce RePro, a novel web recycling method that trains …

  117. arXiv cs.LG TIER_1 English(EN) · Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe, Chamath Gunapala, Pragatheeswaran Vipulanandan, Uthayasanker Thayasivam, Kamal Premaratne ·

    Evaluation of Contextual Understanding in Large Language Models

    arXiv:2609.09004v1 Announce Type: cross Abstract: Large Language Models (LLMs) demonstrate impressive performance across diverse NLP tasks, yet their ability to exhibit genuine contextual understanding remains uncertain. Traditional evaluation metrics such as perplexity, BiLingua…

  118. arXiv cs.LG TIER_1 English(EN) · Zixuan Liu, Fangzheng Wu, Brian Summa, Zizhan zheng ·

    Risk-Conditioned Fine-Tuning of Large Language Models

    arXiv:2609.08064v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in settings where rare but severe harmful generations can have significant consequences. Existing Risk-Averse RLHF addresses this issue by optimizing Conditional Value-at-Risk (…

  119. arXiv cs.LG TIER_1 English(EN) · Yuyang Wang, Haoyu Yao, Pengcheng Xie ·

    MpSub: A Momentum $p$-Dimensional Subspace Trust-Region Method for Derivative-Free Fine-Tuning of Large Language Models

    arXiv:2609.07666v1 Announce Type: new Abstract: Full-parameter fine-tuning of large language models has substantial memory costs because backpropagation stores activations and gradients. Zeroth-order optimization avoids this by estimating update directions from loss evaluations, …

  120. arXiv cs.CL TIER_1 English(EN) · Remco Hendriks (Continker) ·

    MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

    arXiv:2609.10016v1 Announce Type: cross Abstract: We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, f…

  121. arXiv cs.CL TIER_1 English(EN) · Mehrnaz Mofakhami, Ananya Sahu, Alejandro R. Salamanca, Daniel D'souza, Alexandre Berard, Thomas Euyang, Marzieh Fadaee, Julia Kreutzer ·

    Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning

    arXiv:2609.10445v1 Announce Type: new Abstract: Reasoning language models have made substantial advances on a variety of complex tasks, yet their capabilities remain overwhelmingly English-centric: models primarily reason in English regardless of the language they are prompted in…

  122. arXiv cs.CL TIER_1 English(EN) · Leilei Ding, Shumin Wang, Yuting Huang, Fanqi Wan, Yinmin Zhang, Qi Han, Yiming Xu, Feiyuan Zhang, Xiaomeng Chu, Guoliang You, Wuyang Zhang, Daxin Jiang, Yanyong Zhang ·

    $\Phi$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?

    arXiv:2609.10226v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in developing and optimizing the very infrastructure that powers them. However, exi…

  123. arXiv cs.CL TIER_1 English(EN) · Christoph Wigbels, Ali Abusaleh, Markus T. Jansen, Alexander Mehler, Markus J. Hofmann ·

    From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora

    arXiv:2609.10155v1 Announce Type: new Abstract: We approach a cognitive simulation perspective on episodic and semantic memory in multiple-choice question answering by incorporating text from individual text corpora (ITC) into retrieval-augmented generation and DoRA fine-tuning. …

  124. arXiv cs.CL TIER_1 English(EN) · An Vo, Vy Tuong Dang, Khai-Nguyen Nguyen, Emilio Villa-Cueva, Thamar Solorio, Anh Totti Nguyen, Daeyoung Kim ·

    Deep and shallow biases in language models

    arXiv:2609.09901v1 Announce Type: new Abstract: Large language models often repeatedly select the same answer even when many alternatives are plausible. Prior work treats this concentration as bias, but it does not distinguish stable model preferences from responses that depend o…

  125. arXiv cs.AI TIER_1 English(EN) · Hankun Kang, Di Lin, Zhirong Liao, Pengfei Bai, Xinyi Zeng, Jiawei Jiang, Yuanyuan Zhu, Dexi Liu, Tieyun Qian ·

    AdaCultureSafe: Adaptive Cultural Safety Grounded by Cultural Knowledge in Large Language Models

    arXiv:2603.08275v2 Announce Type: replace-cross Abstract: With the global proliferation of Large Language Models (LLMs), cultural safety, defined as the ability to generate respectful and appropriate responses across diverse cultures, becomes critical for responsible AI deploymen…

  126. arXiv cs.AI TIER_1 English(EN) · Haoqing Wang, Xiang Long, Ziheng Li, Yilong Xu, Tingguang Li, Yehui Tang ·

    To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models

    arXiv:2602.12566v4 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) plays a key role in stimulating the explicit reasoning capability of Large Language Models (LLMs). We can achieve expert-level performance in some specific domains via RLVR, …

  127. arXiv cs.AI TIER_1 English(EN) · Brian Rabern, Philipp Mondorf, Barbara Plank ·

    LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models

    arXiv:2602.06533v3 Announce Type: replace Abstract: Large language models perform well on many logical reasoning benchmarks, but it remains unclear which core logical skills they truly master. To address this, we introduce LogicSkills, a benchmark that isolates three fundamental …

  128. arXiv cs.AI TIER_1 English(EN) · Xiao Jia, Zhanzhan Zhao ·

    The Emergence of Social Science of Large Language Models

    arXiv:2509.24877v3 Announce Type: replace Abstract: The social science of large language models (LLMs) examines how these systems evoke mind attributions, interact with one another, and transform human activity and institutions. We conducted a systematic review of 270 studies, co…

  129. arXiv cs.AI TIER_1 English(EN) · Oleksandr Cherednichenko, Roman Klypa ·

    Suan: Rectifying Direct Preference Safety Alignment in Large Language Models

    arXiv:2609.08634v1 Announce Type: cross Abstract: Integrating robust safety guardrails into Large Language Models (LLMs) is essential for delivering helpful yet harmless responses. While proprietary systems exhibit reliable safety controls, their underlying methodologies and trad…

  130. arXiv cs.AI TIER_1 English(EN) · Yi Yang, Xiao Jia, Zeyun Dong, Chenzhang Wang, Zhanzhan Zhao ·

    Mapping the Emerging Social Science of Large Language Models

    arXiv:2609.07598v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly shape communication, learning, work, creativity, and decision-making, yet social-science research on these developments remains fragmented. We map this emerging field using a curated corpu…

  131. arXiv cs.AI TIER_1 English(EN) · Vamsi Krishna Kodavali, Rituraj Singh ·

    LANTERN: Language Model Assessment on Noisy and Transformed Tasks for Understanding Error and Robustness Nuances

    arXiv:2609.07309v1 Announce Type: cross Abstract: Robustness evaluation of large language models (LLMs) remains a critical challenge, particularly in assessing their sensitivity to perturbations in input data. In this work, we systematically evaluate LLM robustness across multipl…

  132. arXiv cs.AI TIER_1 English(EN) · Esteban Feuerstein, Victoria Klimkowski, Juan Manuel Ortiz de Zarate, Federico Hern\'an Suaiter ·

    Recovering Temporal and Geographic Signals from Language Model Embeddings

    arXiv:2609.05721v1 Announce Type: new Abstract: Understanding whether language-model embeddings encode structured real-world information is important for both representation analysis and information retrieval. We study this question for temporal and geographic signals using a sim…

  133. arXiv cs.AI TIER_1 English(EN) · Sunny Rai, Jinyi Kuang, Reyhan Jamalova, Annie Lou, Cristina Bicchieri, Niyati Malhotra, Victor Hugo Orozco-Olvera, Ana Maria Munoz-Boudet, Lyle H Ungar, Sharath C Guntuku ·

    Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

    arXiv:2609.05437v1 Announce Type: new Abstract: Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, social intelligence depends not only on norm recognitio…

  134. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Markus J. Hofmann ·

    From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora

    We approach a cognitive simulation perspective on episodic and semantic memory in multiple-choice question answering by incorporating text from individual text corpora (ITC) into retrieval-augmented generation and DoRA fine-tuning. We web-crawl the search histories of 515 partici…

  135. Hugging Face Daily Papers TIER_1 English(EN) ·

    MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

    A benchmark for transit-kiosk policy reasoning shows that small parameter-efficiently fine-tuned language models can match larger models on structured tool-use and fare-quoting tasks.

  136. Hugging Face Daily Papers TIER_1 English(EN) ·

    Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?

    Φ-Bench evaluates large language models on open-ended engineering of the LLM infrastructure stack across tasks from kernel optimization to end-to-end system design.

  137. Hugging Face Daily Papers TIER_1 English(EN) ·

    Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning

    Optimized supervised fine-tuning data composition enables reasoning models to consistently process and respond in diverse non-English languages without requiring reasoning supervision in each target language.

  138. Hugging Face Daily Papers TIER_1 English(EN) ·

    Suan: Rectifying Direct Preference Safety Alignment in Large Language Models

    Integrating robust safety guardrails into Large Language Models (LLMs) is essential for delivering helpful yet harmless responses. While proprietary systems exhibit reliable safety controls, their underlying methodologies and trade-offs remain largely undisclosed. Achieving compa…

  139. arXiv cs.LG TIER_1 English(EN) · Fred Zhangzhi Peng, Kaiwen Zheng, Anru R. Zhang ·

    Distilled Continuous Diffusion Language Models Can Write Code in Few Steps---or One

    arXiv:2609.04531v1 Announce Type: new Abstract: Language generation is almost universally treated as a sequential process: autoregressive models emit one token at a time, while diffusion language models replace token-level seriality with a long trajectory of iterative refinement.…

  140. arXiv cs.CL TIER_1 English(EN) · Zhuoya Zhao, Parsa Omidi, Aref Jafari, Richard Naud ·

    Large Language Models with At Most One Spike per Neuron

    arXiv:2609.05151v1 Announce Type: cross Abstract: Leveraging their inherent sparse event-driven computation, spiking neural networks (SNNs) offer a promising path toward energy-efficient large language models (LLMs). Time-to-first-spike (TTFS) coding generates at most one spike p…

  141. arXiv cs.CL TIER_1 English(EN) · Daniil Ognev, C\'elian Vasson, Lijie Hu, Kentaro Inui, Benjamin Heinzerling ·

    SharedSAE: One Feature Dictionary Across Language Models

    arXiv:2609.04344v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are widely used to interpret language model activations, but SAE training and latent labelling are typically repeated for every model. Here, we show that a single shared SAE can replace a collection of d…

  142. arXiv cs.CL TIER_1 English(EN) · Aleix Sant, Jordi Luque, Carlos Escolano ·

    EuroAlpaca: Task-Preserving Localisation of Instruction Data for European Languages

    arXiv:2609.05043v1 Announce Type: new Abstract: Machine translation (MT) offers a scalable way to extend English instruction-tuning data to multiple languages, but it can distort task-critical constraints and required outputs, creating corrupted training examples and degrading mo…

  143. arXiv cs.CL TIER_1 English(EN) · Ahrii Kim, Chanjun Park, Seong-heum Kim ·

    Discourse Dependency: A Continuous Criterion for Translation Difficulty

    arXiv:2609.04959v1 Announce Type: new Abstract: Recent calls for harder machine translation benchmarks have not clarified what difficulty should mean. We argue that one meaningful and currently unmeasured axis is referential reach, the distance a segment must look back into its d…

  144. arXiv cs.CL TIER_1 English(EN) · Elisabeth Kirsten, Nicole Kr\"amer, Muhammad Bilal Zafar ·

    On Epistemic Diversity in Large Language Models

    arXiv:2609.04835v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used not only to retrieve information, but to answer questions, explain, teach, and support inquiry. In such settings, evaluation cannot be exhausted by accuracy or alignment alone. A sy…

  145. arXiv cs.CL TIER_1 English(EN) · Oskar Holmstr\"om, Marcel Bollmann, Marco Kuhlmann ·

    A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures

    arXiv:2609.04819v1 Announce Type: new Abstract: Multilingual language models develop shared cross-lingual representations, and various interpretability methods claim to quantify this sharing. These methods have been developed largely in isolation, and when they disagree, it is un…

  146. arXiv cs.CL TIER_1 English(EN) · Ekata Mitra, Ameeta Agrawal ·

    Choosing the Right Language Mode at Inference Time for Multilingual Reliability

    arXiv:2609.04653v1 Announce Type: new Abstract: Multilingual large language models often struggle to reason in low- to mid-resource languages. Prior work has shown that translation can improve multilingual reasoning by helping models access stronger English-centric representation…

  147. arXiv cs.CL TIER_1 English(EN) · Kirill Afendulev, Alexey Dontsov, Elena Tutubalina, Anton Korznikov ·

    What Attention Recalls and Recurrence Controls in Hybrid Language Models

    arXiv:2609.04434v1 Announce Type: new Abstract: Hybrid language models combine attention with a fixed-size recurrent state, but the role of each channel remains unclear. We introduce two cache-level interventions. Split-prefill keeps only the KV cache or only the recurrent state …

  148. arXiv cs.AI TIER_1 English(EN) · Hang Li, Kaiqi Yang, Yucheng Chu, Hui Liu, Jiliang Tang ·

    Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving

    arXiv:2509.22480v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have been widely used for problem-solving tasks. Most recent work improves their performance through supervised fine-tuning (SFT) with labeled data or reinforcement learning (RL) from task feed…

  149. arXiv cs.AI TIER_1 English(EN) · Rui Wang, Hongru Wang, Boyang Xue, Jianhui Pang, Shudong Liu, Yi Chen, Jiahao Qiu, Derek Fai Wong, Heng Ji, Kam-Fai Wong ·

    Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models

    arXiv:2503.24377v2 Announce Type: replace-cross Abstract: Recent advancements in Large Language Models (LLMs) have significantly enhanced their ability to perform complex reasoning tasks, transitioning from fast and intuitive thinking (System 1) to slow and deep reasoning (System…

  150. arXiv cs.AI TIER_1 English(EN) · Antoni Czolgowski, Abel Iyasele ·

    Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning

    arXiv:2609.04485v1 Announce Type: cross Abstract: We evaluate three open-weight LLMs (Gemma3-12B from the USA, Bielik-11B-v3 from Poland, and Qwen3-4B from China) against World Values Survey Wave 7 data for 63 demographic personas across three countries, using normalized Wasserst…

  151. arXiv cs.AI TIER_1 English(EN) · Arnau Ayguad\'e Domingo, Stefan Bott, Horacio Saggion ·

    Reinforcement Learning for improving Large Language Models' Catalan text simplification capabilities

    arXiv:2609.04823v1 Announce Type: cross Abstract: Although automatic text simplification (ATS) is critical for accessibility, its progress has not matched the rapid evolution of broader natural language processing techniques. This paper investigates the application of reinforceme…

  152. arXiv cs.AI TIER_1 English(EN) · Jos\'e Luciano Ver\c{c}osa Marques, Frederico Jorge Heitmann, Daniel Omar Perez, Marcelo Vinicius de Paula, T\'arcio Andr\'e dos Santos Barros ·

    Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models

    arXiv:2609.05333v1 Announce Type: new Abstract: A transformer language model assigns a single, context-independent vector to a word type at its embedding layer, yet is widely believed to individuate that word's occurrences by context in its later layers. Testing this belief clean…

  153. arXiv cs.AI TIER_1 English(EN) · Sebastien Kawada, Manolis Kellis ·

    Evidence Integration in Large Language Models

    arXiv:2609.04290v1 Announce Type: cross Abstract: Despite increasing reliance on LLMs that reason with external evidence supplied by tools, retrieval-augmented generation, other agents, and users, how LLMs integrate such evidence into decisions they have already begun to form rem…

  154. arXiv cs.AI TIER_1 English(EN) · Jirui Qi, Mingyang Wang, Hinrich Sch\"utze, Raquel Fern\'andez, Arianna Bisazza ·

    A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models

    arXiv:2609.04409v1 Announce Type: cross Abstract: Multilingual language models often produce inconsistent answers to semantically equivalent questions across languages, motivating methods to improve cross-lingual consistency (CLC). However, existing methods are typically evaluate…

  155. arXiv cs.AI TIER_1 English(EN) · Giulia Pucci, Ruizhe Li, Arabella Sinclair ·

    Patterns of Priming in Production: Lexical, Semantic and Structural Alignment in Language Model Generation

    arXiv:2609.04484v1 Announce Type: cross Abstract: This paper investigates structural priming in language model (LM) production, examining how preceding structural context influences sentence completion. While prior work has demonstrated priming effects in comprehension of structu…

  156. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Federico Hernán Suaiter ·

    Recovering Temporal and Geographic Signals from Language Model Embeddings

    Understanding whether language-model embeddings encode structured real-world information is important for both representation analysis and information retrieval. We study this question for temporal and geographic signals using a simple projection-based method that operates direct…

  157. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Richard Naud ·

    Large Language Models with At Most One Spike per Neuron

    Leveraging their inherent sparse event-driven computation, spiking neural networks (SNNs) offer a promising path toward energy-efficient large language models (LLMs). Time-to-first-spike (TTFS) coding generates at most one spike per neuron within a time window, yielding extremely…

  158. arXiv cs.AI TIER_1 English(EN) · Zixuan Fu, Bingxiang He, Yuxin Zuo, Haohuan Huang, Jinqian Zhang, Ruhang Xiao, Cheng Qian, Qinyu Luo, Huan-ang Gao, Yudong Wang, Zhiyuan Liu, Ning Ding, Chaojun Xiao ·

    Rethinking On-Policy Distillation of Large Language Models II: One Training Example

    arXiv:2609.04172v1 Announce Type: new Abstract: On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine thi…

  159. arXiv cs.CL TIER_1 English(EN) · Karthika Nhayakkat, Rajat Verma, Maharaj Brahma, Vetcha Gnana Mahesh, Maunendra Sankar Desarkar, Ganesh Ramakrishnan, Rohit Saluja ·

    Lost in Reordering: Structural Sensitivity of Multilingual LLMs under Semantics-Preserving Perturbations

    arXiv:2609.03511v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate strong multilingual reasoning performance, yet their robustness to semantics-preserving structural variation remains underexplored, particularly for relatively free word-order languages. We i…

  160. arXiv cs.AI TIER_1 English(EN) · Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu ·

    Representational alignment yields generalizable safety in language models

    arXiv:2609.04022v1 Announce Type: cross Abstract: Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or adv…

  161. arXiv cs.AI TIER_1 English(EN) · Joseph Lee, Yidi Huang, Dokyoon Kim, Shu Yang, Li Shen ·

    Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views

    arXiv:2609.04180v1 Announce Type: cross Abstract: Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experi…

  162. arXiv cs.AI TIER_1 English(EN) · Yuntian Deng, Pengyu Nie, Stuart Shieber ·

    Compile by Training: Turning Natural-Language Specifications into Local Neural Functions

    arXiv:2609.04199v1 Announce Type: cross Abstract: Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by tr…

  163. arXiv cs.CL TIER_1 English(EN) · Malvina Nissim ·

    Language, Language Models, and What We're Talking About

    arXiv:2609.03577v1 Announce Type: new Abstract: Language models are commonly discussed as technical artefacts, but they are obviously shaped by the linguistic worlds conveyed by data during their training. Using Italian language models as evidence, I want to bring attention to th…

  164. arXiv cs.CL TIER_1 English(EN) · Yujun Wang, Ehud Reiter, Shimei Pan, Steffen Eger, Wei Zhao ·

    Beyond Accuracy: Community Perspectives on Machine Translation

    arXiv:2606.09655v2 Announce Type: replace Abstract: Despite remarkable progress in machine translation (MT), non-AI communities have raised growing concerns about MT systems, suggesting a noticeable gap between technical advancement and the needs of real-world users. For instance…

  165. arXiv cs.CL TIER_1 English(EN) · Michael Li, Nishant Subramani ·

    How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits

    arXiv:2605.08348v2 Announce Type: replace Abstract: The circuits framework in mechanistic interpretability aims to identify sparse subgraphs of model components that are causally responsible for a behavior, typically evaluated by measuring necessity and sufficiency. But these cri…

  166. arXiv cs.CL TIER_1 English(EN) · Danting Zhang, Bei Peng, Robert Loftin ·

    Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation

    arXiv:2609.03814v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate strong performance on standard content moderation benchmarks. However, these benchmarks often aggregate multiple moderation criteria into a single label, making it unclear whether models can …

  167. arXiv cs.CL TIER_1 English(EN) · Qianwen Wang, York Hay Ng, Aditya Khan, En-Shiun Annie Lee ·

    Typological Feature Prediction with Large Language Models: An In-Context Learning Approach

    arXiv:2609.03775v1 Announce Type: new Abstract: Typological features are widely used in multilingual NLP, and the prediction of such features holds downstream utility. However, existing methods to predict missing values lack interpretable justifications for predictions, while the…

  168. arXiv cs.AI TIER_1 English(EN) · Baban Gain, Trilok Nath Singh, Asif Ekbal ·

    One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging

    arXiv:2604.02881v2 Announce Type: replace-cross Abstract: Weight-space model merging combines independently fine-tuned checkpoints without access to the original training data. While merging has shown promise in multitask settings, its behavior in multilingual generative systems …

  169. arXiv cs.CL TIER_1 English(EN) · Xinye Yang, Zhenyang Liu, Ruisi Li, Yuanyuan Lei ·

    Large Language Models in Resolving Contextual Knowledge Conflicts

    arXiv:2609.03148v1 Announce Type: new Abstract: Most prior works focused on conflicts between an LLM's internal parametric knowledge and externally provided context. In contrast, we investigate how LLMs handle conflicts that arise within contextual knowledge itself. We introduce …

  170. arXiv cs.LG TIER_1 English(EN) · Lianhai Ren, Yucheng Ding, Xiao Liu, Peng Cheng, Yeyun Gong ·

    MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration

    arXiv:2602.01734v2 Announce Type: replace Abstract: Training instability remains a critical challenge in large language model (LLM) pretraining, often manifesting as sudden gradient explosions that waste significant computational resources. We study training failures in a 5M-para…

  171. arXiv cs.LG TIER_1 English(EN) · Ross Tieman, Evan Markou ·

    Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models

    arXiv:2609.03422v1 Announce Type: new Abstract: Diversity is a widely observed factor in the resilient function of collective systems, yet the type of diversity that matters depends on the properties and failure modes of the system. This distinction is important for systems compo…

  172. arXiv cs.LG TIER_1 English(EN) · Shenzhi Yang, Guangcheng Zhu, Kai Tang, Zhengqing Zang, Xing Zheng, Haobo Wang, Yingfan Ma, Bowen Song, Bo Han, Bo An, Lei Feng, Weiqiang Wang, Junbo Zhao, Gang Chen ·

    DE-Venus: A Data-Efficient RLVR Framework for Large Language Models

    arXiv:2609.03324v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning, but its practical scaling is constrained by expensive on-policy rollouts and the cost of obtaining reliable targets at scale. Existing me…

  173. arXiv cs.LG TIER_1 English(EN) · Ucchwas Talukder Utsha, Sakib Mostafa, James Zou, Md Tauhidul Islam ·

    Language-encoded network topology enables large language models to reason about complex networks

    arXiv:2609.03229v1 Announce Type: new Abstract: Networks describe systems in biology and beyond, from protein interactions and social relationships to power grids and citation records. Reasoning about such systems requires understanding their structure: which elements are central…

  174. Hugging Face Daily Papers TIER_1 English(EN) ·

    Rethinking On-Policy Distillation of Large Language Models II: One Training Example

    On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on …

  175. Hugging Face Daily Papers TIER_1 English(EN) ·

    Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation

    Large language models (LLMs) demonstrate strong performance on standard content moderation benchmarks. However, these benchmarks often aggregate multiple moderation criteria into a single label, making it unclear whether models can disentangle them and reliably apply each criteri…

  176. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Evan Markou ·

    Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models

    Diversity is a widely observed factor in the resilient function of collective systems, yet the type of diversity that matters depends on the properties and failure modes of the system. This distinction is important for systems composed of multiple language models. Different model…

  177. Hugging Face Daily Papers TIER_1 English(EN) ·

    ALRA: Adaptive Local Relational Alignment for Logit-Based Pre-training Distillation of Autoregressive Language Models

    Logit-based knowledge distillation for autoregressive language models usually aligns teacher and student next-token distributions over the entire vocabulary. However, this global objective overlooks relative preferences among likely token alternatives. Existing local approaches o…

  178. arXiv cs.CL TIER_1 English(EN) · Ali Backour ·

    The Dynamics of Continuous Mixture Collapse in Language Models

    arXiv:2609.02049v1 Announce Type: cross Abstract: LLMs latent-state reasoning methods replace discrete intermediate tokens with continuous states, such as weighted mixtures of token embeddings, to retain multiple possible reasoning directions rather than committing to one. Yet pr…

  179. arXiv cs.LG TIER_1 English(EN) · Robert Hu, Carlo Luschi, Paul Balanca ·

    UE5M3 FP4 Block Scaling for Stable Language Model Pretraining

    arXiv:2609.02846v1 Announce Type: new Abstract: Stable 4-bit floating-point (FP4) pretraining is difficult because the E2M1 payload represents only a narrow range of magnitudes. NVIDIA's Transformer Engine \nv{} recipe addresses this with current-tensor scaling, a randomized Hada…

  180. arXiv cs.CL TIER_1 English(EN) · Ze Yu Zhang, Arun Verma, Finale Doshi-Velez, Bryan Kian Hsiang Low ·

    Prompting the Unknown: Understanding Response Uncertainty in Large Language Models

    arXiv:2407.14845v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are widely used in decision-making across diverse domains. Ensuring the generation of safe and reliable responses is critical for the effective deployment of LLM-based applications, particularl…

  181. arXiv cs.CL TIER_1 English(EN) · Momose Oyama, Yusuke Takase, Hidetoshi Shimodaira ·

    Language Model Maps for Prompt-Response Distributions via Log-Likelihood Vectors

    arXiv:2603.18593v2 Announce Type: replace Abstract: We propose a method that represents language models by log-likelihood vectors over prompt-response pairs and constructs model maps for comparing their conditional distributions. In this space, squared Euclidean distances between…

  182. arXiv cs.CL TIER_1 English(EN) · Valentin Lafargue, Ariel Guerra-Adames, Emmanuelle Claeys, Elouan Vuichard, Jean-Michel Loubes ·

    Probing Cultural Signals in Large Language Models through Author Profiling

    arXiv:2603.16749v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in applications with societal impact, raising concerns about the cultural biases they encode. We probe these representations by evaluating whether LLMs can perform author pr…

  183. arXiv cs.CL TIER_1 English(EN) · Shaswat Patel, Vishvesh Trivedi, Yue Han, Yihuai Hong, Eunsol Choi ·

    Bridging Latent Reasoning and Target-Language Generation via Retrieval-Transition Heads

    arXiv:2602.22453v3 Announce Type: replace Abstract: Recent work has identified a subset of attention heads in Transformer as retrieval heads, which are responsible for retrieving information from the context. In this work, we first investigate retrieval heads in multilingual cont…

  184. arXiv cs.CL TIER_1 English(EN) · Yong-En Tian, Yu-Chien Tang, An-Zi Yen, Wen-Chih Peng ·

    CARPAS: Towards Content-Aware Refinement of Provided Aspects for Summarization in Large Language Models

    arXiv:2510.07177v2 Announce Type: replace Abstract: Aspect-based summarization has attracted significant attention for its ability to generate more fine-grained and user-aligned summaries. While most existing approaches assume a set of predefined aspects as input, real-world scen…

  185. arXiv cs.CL TIER_1 English(EN) · Zixiang Xu, Yanbo Wang, Yue Huang, Haomin Zhuang, Yujun Zhou, Jiayi Ye, Sixian Li, Zirui Song, Lang Gao, Chenxi Wang, Zhaorun Chen, Wang Pan, Yue Zhao, Jieyu Zhao, Xiangliang Zhang, Xiuying Chen ·

    SocialMaze: A Benchmark for Evaluating and Enhancing Social Reasoning in Large Language Models in Complex Social Environments

    arXiv:2505.23713v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in socially grounded applications, where success requires interpreting context, inferring others' mental states, and reasoning about unreliable information. Yet existing ben…

  186. arXiv cs.CL TIER_1 English(EN) · Jiaxu Zhao, Meng Fang, Shirui Pan, Wenpeng Yin, Mykola Pechenizkiy ·

    GPTBIAS: A Comprehensive Framework for Evaluating Bias in Large Language Models

    arXiv:2312.06315v2 Announce Type: replace Abstract: Warning: This paper contains content that may be offensive or upsetting. There has been a significant increase in the usage of large language models (LLMs) in various applications, both in their original form and through fine-tu…

  187. arXiv cs.CL TIER_1 English(EN) · Leon Fr\"ohling, Jens Rupprecht, Markus Strohmaier, Claudia Wagner ·

    When Persona Attributes Improve Population Alignment in Large Language Models

    arXiv:2609.02526v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used to predict the responses of human participants in survey panels. Towards that goal, persona prompting has recently emerged as a technique to inform and align large pretrained langua…

  188. arXiv cs.CL TIER_1 English(EN) · Irina Proskurina, Guillaume Metzler, Antoine Gourru, Julien Velcin ·

    Debias-SparseGPT: Bias-Aware Pruning for Large Language Models

    arXiv:2609.02496v1 Announce Type: new Abstract: Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of Large Language Models (LLMs). However, recent studies show that weight sparsification methods, such as SparseGPT, …

  189. arXiv cs.CL TIER_1 English(EN) · Smitha Muthya Sudheendra, Jaideep Srivastava ·

    When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models

    arXiv:2609.02438v1 Announce Type: new Abstract: Large language models can look capable of logical reasoning, but correct or incorrect answers alone tell us little about what the model represents internally. We study logical verification in five open-weight transformer models usin…

  190. arXiv cs.CL TIER_1 English(EN) · Yixuan Wang, Freda Shi, Kanishka Misra ·

    Disentangling Statistical Preemption from Entrenchment in Language Models' Avoidance of Overgeneralization

    arXiv:2609.01794v1 Announce Type: new Abstract: How do learners avoid overgeneralizations such as Tom laughed me without explicit negative evidence? Constructionists have posited two proposals that describe indirect negative evidence against overgeneralizations: preemption (which…

  191. arXiv cs.AI TIER_1 English(EN) · Wazib Ansar, Saptarsi Goswami, Amlan Chakrabarti ·

    A Survey of Transformer-based Language Models with Focus on Efficiency

    arXiv:2406.16893v3 Announce Type: replace-cross Abstract: The emergence of Transformer-based Large Language Models (LLMs) has substantially augmented the capabilities of Natural Language Processing (NLP), thereby intensifying the demand for computational resources. Therefore, enh…

  192. arXiv cs.AI TIER_1 English(EN) · Kushagra Bhushan, Meghanadh Pulivarthi, Sai Krishna Reddy Sathi, Gaurav Pandey, Sonam Gupta, Vineet Kumar, Jaydeep Sen, Yatin Nandwani, Sachindra Joshi, Dinesh Raghu ·

    DKL: Decoupled Knowledge Learning for Instruction-Tuned Language Models

    arXiv:2609.02685v1 Announce Type: cross Abstract: RAG has become the de facto method for incorporating new, corpus-specific knowledge into an instruction following LLM (Instruct LLM). Although RAG-based prompting improves factual grounding, it fails when retrieval is incorrect or…

  193. arXiv cs.AI TIER_1 English(EN) · Youqi Wu, Farzan Farnia ·

    Do Large Language Models Capture the Diversity in their Training Data?

    arXiv:2609.02275v1 Announce Type: cross Abstract: Large language models are trained to model conditional distributions over text, yet it remains inadequately understood whether they capture the full diversity of plausible outputs present in their training data. We study this ques…

  194. arXiv cs.AI TIER_1 English(EN) · Aur\'elien Lac, Tony Wu ·

    NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference

    arXiv:2609.01657v1 Announce Type: cross Abstract: Multimodal models often build on architectures designed for generative vision-language modeling, typically combining separately pretrained vision encoders with causal language models. Visual document retrievers such as ColPali rep…

  195. arXiv cs.AI TIER_1 English(EN) · Jian Gao, Xiao Zhang, Xun Zhu, Miao Li, Ji Wu ·

    Task-Level Natural Language Priors as Learning Signals for Low-Resource LLM Training

    arXiv:2609.02244v1 Announce Type: new Abstract: Large language models (LLMs) often struggle when low-resource training data are ambiguous or incomplete. Task-level natural-language priors can provide useful guidance in such settings, but existing approaches usually treat these pr…

  196. arXiv cs.AI TIER_1 English(EN) · Renjie Xie, Juncheng Yang, Aoting Hu, Mingxi Zhang, Liyao Wu, Zheheng Hong, Wei Xu ·

    HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models

    arXiv:2609.02029v1 Announce Type: new Abstract: Long-context inference retains a growing key--value (KV) cache during decoding, which consumes substantial GPU memory and can reduce generation throughput. This bottleneck remains in hybrid language models because their residual glo…

  197. arXiv cs.AI TIER_1 English(EN) · Chen Wang, Junzhe Zhao, Xin Cong, Wanlu Deng, Ke Deng ·

    Benchmarking Language Models for Statistical Problem Formulation

    arXiv:2609.01982v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as assistants for statistical and data science work, yet existing evaluations largely assume the analysis target is already specified. In practice, users arrive with informal goals …

  198. Hugging Face Daily Papers TIER_1 English(EN) ·

    How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

    Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is usually evaluated only through output behavior. We study how six naturalistic and synthetic input perturbations propagate through decoder-only language models at three lev…

  199. Hugging Face Daily Papers TIER_1 English(EN) ·

    Language-encoded network topology enables large language models to reason about complex networks

    Networks describe systems in biology and beyond, from protein interactions and social relationships to power grids and citation records. Reasoning about such systems requires understanding their structure: which elements are central, which connections bridge separate communities,…

  200. Hugging Face Daily Papers TIER_1 English(EN) ·

    Rethinking On-Policy Distillation of Large Language Models II: One Training Example

    On-policy distillation improves over hundreds of steps from a single query by rapidly covering teacher states, yet student alignment remains slow, indicating the method is algorithm-starved rather than data-starved.

  201. Hugging Face Daily Papers TIER_1 English(EN) ·

    Compile by Training: Turning Natural-Language Specifications into Local Neural Functions

    Compile by training converts natural-language specifications into reusable neural functions by distilling teacher-generated examples into small adapters, enabling efficient deployment without remote model dependencies.

  202. Hugging Face Daily Papers TIER_1 English(EN) ·

    Large Language Models in Resolving Contextual Knowledge Conflicts

    Most prior works focused on conflicts between an LLM's internal parametric knowledge and externally provided context. In contrast, we investigate how LLMs handle conflicts that arise within contextual knowledge itself. We introduce a taxonomy of six types of contextual conflicts …

  203. Hugging Face Daily Papers TIER_1 English(EN) ·

    Unifying Conformal Language Tasks with In-Context Ensembles

    Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content from documents under two constraints: coverage, retaining enough pertinent information to achieve some goal, and conciseness, removing as much irrelevant information as …

  204. Hugging Face Daily Papers TIER_1 English(EN) ·

    UE5M3 FP4 Block Scaling for Stable Language Model Pretraining

    Stable 4-bit floating-point (FP4) pretraining is difficult because the E2M1 payload represents only a narrow range of magnitudes. NVIDIA's Transformer Engine \nv{} recipe addresses this with current-tensor scaling, a randomized Hadamard transform (RHT), and bfloat16 (BF16) final …

  205. Hugging Face Daily Papers TIER_1 English(EN) ·

    DKL: Decoupled Knowledge Learning for Instruction-Tuned Language Models

    RAG has become the de facto method for incorporating new, corpus-specific knowledge into an instruction following LLM (Instruct LLM). Although RAG-based prompting improves factual grounding, it fails when retrieval is incorrect or incomplete, leading to hallucinations. Finetuning…

  206. arXiv cs.CL TIER_1 English(EN) · Gokul Srinivasagan, Munir Georges ·

    Location-Aware Language Models via Secondary Embeddings

    arXiv:2609.00454v1 Announce Type: new Abstract: Pretrained transformer-based language models achieve strong performance across a wide range of NLP tasks but remain limited in encoding geo-locational semantics, leading to suboptimal representations of place names and spatial entit…

  207. arXiv cs.CL TIER_1 English(EN) · Ryo Mitsuhashi, Sabri Boughorbel, Majd Hawasly ·

    Latent Mechanisms of Language Control in Multilingual Language Models

    arXiv:2609.00325v1 Announce Type: new Abstract: Multilingual large language models can exhibit unintended code-switching -- unnecessarily alternating between languages during generation. We present a comparative study of three methods that identify language-controlling latents in…

  208. arXiv cs.AI TIER_1 English(EN) · Andrei-Valentin T\u{a}nase, Elena Pelican ·

    SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance

    arXiv:2508.11857v3 Announce Type: replace-cross Abstract: Tokenization remains a persistent bottleneck in language modeling, especially when vocabulary learning is limited by whitespace boundaries. We present SupraTok, a tokenizer that crosses whitespace boundaries using three mo…

  209. arXiv cs.AI TIER_1 English(EN) · Siyuan Li, Xinxin Song, Chen Ruinian, Jingjing Fan, Tingxiong Xiao, Yangen Hu, Ke Zeng, Jinli Suo ·

    CARE: Contrastive Anchor-based Rubric Evolution for Large Language Model Post-Training

    arXiv:2609.00892v1 Announce Type: new Abstract: Rubric-based reinforcement learning decomposes open-ended instructions into prompt-specific, flexible rubrics, making it better suited than reinforcement learning with verifiable rewards for post-training LLMs on open-ended tasks. H…

  210. arXiv cs.AI TIER_1 Dansk(DA) · Shiyu Li, Zi-Yuan Hu, Shijia Huang, Yanyang Li, Yiwu Zhong, Liwei Wang ·

    SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models

    arXiv:2609.01004v1 Announce Type: cross Abstract: Despite their strong multimodal understanding ability, multimodal large language models (MLLMs) incur substantial computational overhead when processing long visual token sequences. To reduce inference costs, recent studies have e…

  211. arXiv cs.AI TIER_1 English(EN) · Yijun Chen, Yaqi Zheng, Yanya Li, Boyi Xiao, Buqiang Xu, Shuofei Qiao, Jizhan Fang, Xinle Deng, Yunzhi Yao, Xuehai Wang, Liuxin Zhang, Hui Li, Huajun Chen, Shumin Deng ·

    EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

    arXiv:2609.00551v1 Announce Type: cross Abstract: Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, summaries, or graph facts as isolated fragments. Although searchable, such fragmen…

  212. arXiv cs.AI TIER_1 English(EN) · Jacob Brinton, Jannik Brinkmann, Mark Crovella, Aaron Mueller ·

    The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space

    arXiv:2609.00515v1 Announce Type: cross Abstract: Large language models (LLMs) have recently demonstrated improved machine translation performance over strong supervised baselines. This raises questions as to what mechanisms underlie how LLMs perform machine translation between l…

  213. arXiv cs.AI TIER_1 English(EN) · Saman Rahbar ·

    The Curse of Multilinguality in Lexical Normalization

    arXiv:2609.00329v1 Announce Type: cross Abstract: Lexical normalization rewrites the noisy, non-standard words that fill user-generated text (tmrw, u, gr8) into their standard forms. Because labelled data is scarce for most languages, a popular shortcut is to train a single model…

  214. arXiv cs.AI TIER_1 English(EN) · Frederic Sadrieh, Michal \v{S}tef\'anik ·

    Prompt-Robust Language Models: Which Training Strategies Work?

    arXiv:2609.01217v1 Announce Type: new Abstract: Despite their strong performance, large language models remain highly sensitive to prompt formulation. Prior work addresses this through refined data construction or through dedicated robustness objectives. We reproduce and compare …

  215. arXiv cs.AI TIER_1 English(EN) · Yiwen Jiang, Yang Deng, Stephanie Fong, Zimu Wang, Yaling Shen, Wei Feng, Hongxi Yang, Xiangyu Zhao, Zhongxing Xu, Deval Mehta, Xuelian Cheng, Zongyuan Ge ·

    VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences

    arXiv:2609.00921v1 Announce Type: new Abstract: Personalized Large Language Models (PLLMs) aim to tailor responses to individual users, where a central challenge is preference reasoning: inferring query-relevant preferences from user-related history. Existing benchmarks, however,…

  216. arXiv cs.AI TIER_1 English(EN) · Danze Chen, Zeqing Wang, Ziyue Lin, Xingyi Yang, Yeying Jin ·

    H3-World: Turning Language Understanding into World Control

    arXiv:2609.01560v1 Announce Type: cross Abstract: We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world model. Our key finding is that, as large video generators become more capable, language is emerging as a natural i…

  217. arXiv cs.AI TIER_1 English(EN) · Vibhhu Sharma, Thorsten Joachims, Sarah Dean ·

    Value Over Language Model: Detecting Original Contribution in Writing

    arXiv:2609.00700v1 Announce Type: new Abstract: LLMs have been rapidly adopted across writing tasks, prompting the development of tools for detecting LLM-generated text. Yet, these tools largely measure how much of a document's surface text was written by an LLM and aren't fundam…

  218. arXiv cs.AI TIER_1 English(EN) · Jainil Dharmil Shah ·

    Triple-Bottom-Line Sustainability of Language Models for Edge AI: A Comparison Between SLMs and Quantized LLMs

    arXiv:2609.00665v1 Announce Type: new Abstract: Edge-AI model selection is commonly driven by one isolated metric - accuracy, latency, memory, energy, or safety, even though a deployable language model must balance all five. Our work focuses on answering the question whether na- …

  219. arXiv cs.AI TIER_1 English(EN) · Cris Huynh ·

    Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random

    arXiv:2609.00576v1 Announce Type: new Abstract: Item-sensitivity, defined as whether a model's choice depends on the specific input rather than on its own output prior, is widely reported as evidence of task competence. We show this evidence is necessary but not sufficient using …

  220. arXiv cs.CL TIER_1 English(EN) · Sen Yang, Shanbo Cheng, Lu Xu, Jianbing Zhang, Shujian Huang ·

    GRRM: Group Relative Reward Modeling for Machine Translation

    arXiv:2602.14028v2 Announce Type: replace Abstract: While Group Relative Policy Optimization (GRPO) offers a powerful framework for LLM post-training, its effectiveness in open-ended domains like Machine Translation hinges on accurate intra-group ranking. We identify that standar…

  221. arXiv cs.CL TIER_1 English(EN) · Tzu-Han Lin, Wei-Lin Chen, Chen-An Li, Hung-yi Lee, Yun-Nung Chen, Yu Meng ·

    AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning

    arXiv:2512.16883v2 Announce Type: replace Abstract: Equipping large language models (LLMs) with search engines via reinforcement learning (RL) promises effective search agents. However, adaptively balancing internal parametric knowledge with external search remains a challenge, a…

  222. arXiv cs.CL TIER_1 English(EN) · Fenghai Li, Zihan Tang, Haofei Yu, Yining Zhao, Jiaxuan You ·

    Can Large Language Models Forecast What Researchers Study Next?

    arXiv:2609.00747v1 Announce Type: new Abstract: Large language models increasingly generate research ideas, yet judging their novelty or feasibility at generation time does not establish whether they anticipate subsequent work. We introduce IdeaForecastBench to evaluate research …

  223. arXiv cs.CL TIER_1 English(EN) · Guangyu Chen, Boxuan Lyu, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura ·

    Quit While You're Ahead: Quit for Efficient Candidate Generation in Machine Translation Reranking

    arXiv:2609.00588v1 Announce Type: new Abstract: Reranking methods, such as Minimum Bayes Risk (MBR) decoding and Quality Estimation (QE) reranking, are widely used in modern neural machine translation (NMT) to select an output from a set of candidate hypotheses. However, the perf…

  224. Hugging Face Daily Papers TIER_1 English(EN) ·

    HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models

    Long-context inference retains a growing key--value (KV) cache during decoding, which consumes substantial GPU memory and can reduce generation throughput. This bottleneck remains in hybrid language models because their residual global-attention layers can dominate context-depend…

  225. Hugging Face Daily Papers TIER_1 English(EN) ·

    Debias-SparseGPT: Bias-Aware Pruning for Large Language Models

    Debias-SparseGPT reduces pruning-induced demographic bias in compressed large language models by incorporating representational debiasing during post-training sparsification.

  226. Hugging Face Daily Papers TIER_1 English(EN) ·

    Unifying Conformal Language Tasks with In-Context Ensembles

    The Conformal Relevance framework automates score function design for conformal prediction via in-context learning and ensembling to improve conciseness while preserving coverage across NLP retrieval tasks.

  227. arXiv cs.CL TIER_1 English(EN) · Jiaqi Wei, Xiang Zhang, Yuejin Yang, Wenxuan Huang, Juntai Cao, Sheng Xu, Xiang Zhuang, Zhangyang Gao, Muhammad Abdul-Mageed, Laks VS Lakshmanan, Chenyu You, Wanli Ouyang, Siqi Sun ·

    When LLM Meets Tree Search: A Systematic View of Inference as Search in Large Language Models

    arXiv:2608.30395v1 Announce Type: new Abstract: As pretraining scaling laws approach saturation, Test-Time Scaling (TTS) has emerged as an important direction for improving reasoning by allocating inference-time compute to a fixed model prior. Viewed at a high level, TTS reframes…

  228. arXiv cs.CL TIER_1 English(EN) · Yujing Chang, Thinh Pham, Van-Phat Thai, Chunyao Ma, Yash Guleria, Pham Nhut Huy, Sameer Alam ·

    Beyond Semantic Accuracy: Consequence-Aware Evaluation for Safety-Critical Language Understanding

    arXiv:2608.24621v2 Announce Type: replace Abstract: Can language models be trusted in safety- critical operations? In such settings, strong per- formance on semantic metrics does not guaran- tee operational reliability: a misread altitude, a dropped execution condition, or a conf…

  229. arXiv cs.CL TIER_1 English(EN) · Faeze Ghorbanpour, Constanza Fierro, Alexander Fraser, Anders Sogaard ·

    Latent-Space Intervention for Cross-Lingual Factual Consistency: Consistency Improvements without Accuracy Drops

    arXiv:2608.28860v1 Announce Type: new Abstract: Large Language Models (LLMs) often answer the same factual question differently across languages. We study whether cross-lingual latent-space intervention can reduce this inconsistency. We train layer-specific autoencoders on parall…

  230. arXiv cs.AI TIER_1 English(EN) · Jae O Lee, Jaemin Kim, Sumyeong Ahn, Seo Yeon Park ·

    Where Does Robustness Live? Neuron-Guided Adaptation for Retrieval-Augmented Language Models

    arXiv:2604.02194v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Language Models (RALMs) have shown strong potential in knowledge-intensive tasks, yet they remain vulnerable when retrieved contexts are noisy or irrelevant. Robustness against such contexts requires tw…

  231. arXiv cs.AI TIER_1 English(EN) · Yumeng Wang, Yuchen Wu, Cheng Qian, Zhiyuan Fan, Hyeonjeong Ha, Shujin Wu, Jiayu Liu, Heng Ji, Ge Wang ·

    Evaluating the Hidden Costs of Personalization in Large Language Models

    arXiv:2608.28833v1 Announce Type: new Abstract: While Large language models (LLMs) incorporate user personalization signals to improve usability and helpfulness, they increasingly shift from providing balanced, informative responses toward optimizing for user satisfaction when co…

  232. arXiv cs.AI TIER_1 English(EN) · Zhenghao He, Guangzhi Xiong, Bohan Liu, Sanchit Sinha, Aidong Zhang ·

    Triggering Chain-of-Thought via Latent Feature Interventions in Large Language Models

    arXiv:2601.08058v2 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) prompting often improves the reasoning performance of large language models (LLMs), but the internal signal that triggers this behavior remains poorly understood. Leveraging the sparse features captu…

  233. arXiv cs.AI TIER_1 English(EN) · Hongbo Wang, AprilPyone MaungMaung, Isao Echizen ·

    SGM: Safety Glasses for Multimodal Large Language Models via Neuron-Level Detoxification

    arXiv:2512.15052v4 Announce Type: replace-cross Abstract: Disclaimer: Samples in this paper may be harmful and cause discomfort. Multimodal large language models (MLLMs) enable multimodal understanding but inherit toxic signals from weakly curated pretraining corpora, leading to …

  234. arXiv cs.AI TIER_1 English(EN) · Xiaohang Tang, Sangwoong Yoon, Seongho Son, Huizhuo Yuan, Quanquan Gu, Ilija Bogunovic ·

    RSPO: Regularized Self-Play Alignment of Large Language Models

    arXiv:2503.00030v3 Announce Type: replace-cross Abstract: Self-play-based policy optimization has emerged as an effective approach for fine-tuning large language models (LLMs), formulating preference optimization as a two-player game. However, the regularization with respect to t…

  235. arXiv cs.AI TIER_1 English(EN) · Shahar Ben-Natan, Oren Tsur ·

    Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models

    arXiv:2601.15436v3 Announce Type: replace Abstract: We propose a novel perspective for probing LLM sycophancy in a direct and neutral way, mitigating various forms of uncontrolled bias, noise, or manipulative language, deliberately injected to prompts in prior works. A key novelt…

  236. arXiv cs.AI TIER_1 English(EN) · Matvei Tarasov, Salman Ahmadi-Asl, Andre L. F. de Almeida, Andrzej Cichocki ·

    Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability

    arXiv:2608.30505v1 Announce Type: cross Abstract: Large language models (LLMs) are built from structured high-dimensional objects such as token representations, weights, adaptation updates, caches, and activations, whose multilinear structure is underexploited by the conventional…

  237. arXiv cs.AI TIER_1 English(EN) · Minju Song, Hyeon Hwang, Junhyun Lee, Jaewoo Kang ·

    Enhancing Low-Resource Language Reasoning via High-Resource Language Feature Transfer

    arXiv:2608.30462v1 Announce Type: cross Abstract: Large language models exhibit substantial performance variation across languages, even when solving semantically equivalent tasks. Existing analyses often treat this phenomenon as an observational disparity caused by differences i…

  238. arXiv cs.AI TIER_1 English(EN) · Owen Kwon, Pablo Ortega-Kral, Arthur Bucker, Jean Oh ·

    Training-Free Action Correction for VLA Model Failures via Language Feedback

    arXiv:2608.29967v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models demonstrate strong semantic understanding yet exhibit systematic failures during deployment. The conditions under which these failures occur, and whether they can be corrected without retraining…

  239. arXiv cs.AI TIER_1 English(EN) · Alberto Cetoli ·

    Sleight of Word Benchmark: Can Language Models Notice If Their Own Output Was Tampered With?

    arXiv:2608.29921v1 Announce Type: cross Abstract: The output of a Language Model can be tampered with \emph{while} the model is writing it. A simple test can thus be constructed by evaluating the model's perception of this external perturbation. In this spirit, a simple benchmark…

  240. arXiv cs.AI TIER_1 English(EN) · Zhang Enyan, R. Thomas McCoy ·

    A Unifying Perspective on Language Model Representations: From Filler-Role Structure to Mechanistic Interpretability

    arXiv:2608.29034v1 Announce Type: cross Abstract: A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, different methods and their resulting insights stand in relative isolation: what could …

  241. arXiv cs.AI TIER_1 English(EN) · Jinqi Wu, Sishuo Chen, Zhangming Chan, Yong Bai, Chao Yi, Han Zhu, Shuodian Yu, Lei Zhang, Sheng Chen, Chenghuan Hou, Jian Xu, Chaoyou Fu ·

    Can Large Language Models Identify Meaningful Touchpoints in Conversion Attribution?

    arXiv:2608.28649v1 Announce Type: cross Abstract: Touchpoint selection in conversion attribution, namely identifying meaningful touchpoints contributing to conversions, is essential for e-commerce recommendation and online advertising. Current selection methods rely heavily on co…

  242. arXiv cs.AI TIER_1 English(EN) · Emad Alharbi ·

    Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

    arXiv:2608.28626v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as revi…

  243. arXiv cs.AI TIER_1 English(EN) · Fateme Mazdarani, Carlos Toxtli ·

    Beyond the Answer Key: Robustness Evaluation of Large Language Models for Step-Level Mathematical Verification

    arXiv:2608.28725v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as graders, verifiers, and process auditors, but most mathematical evaluations still emphasize final-answer accuracy. This can obscure whether a model can verify a non-canonical but…

  244. arXiv cs.AI TIER_1 English(EN) · Anirudh Malik, M Sparsh Mehra, Poojith Devan ·

    Capability-Stratified Degradation in Ternary Language Models

    arXiv:2608.28809v1 Announce Type: new Abstract: Extreme low-bit inference offers a route toward smaller models and constrained deployment. Ternary language models restrict weights to $\{-1,0,+1\}$, approaching the limit of $\log_2 3 \approx 1.585$ bits/weight. The practical quest…

  245. arXiv cs.LG TIER_1 English(EN) · Chanyoung Kim, Minwoo Kim, Minseok Kang, Hyunwoo Kim, Dahuin Jung ·

    LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models

    arXiv:2603.28301v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models achieve strong performance in robotic manipulation by leveraging pre-trained vision-language backbones. However, in downstream robotic settings, they are typically fine-tuned with limited data…

  246. arXiv cs.LG TIER_1 English(EN) · Zhipeng Xia, Haotian Xu, Siyu Yun, Liqi Lin, Hu Liu, Yu Li, Cheng Zhuo ·

    TrainSDC: Characterizing and Mitigating Silent Data Corruption in Large Language Model Training

    arXiv:2608.30769v1 Announce Type: new Abstract: LLM training is increasingly vulnerable to silent data corruption (SDC), yet existing protection methods largely treat Transformer computations uniformly because their vulnerability remains poorly understood. We present the first sy…

  247. arXiv cs.CL TIER_1 English(EN) · Zhongxi Wang, Yueqian Lin, Jingyang Zhang, Zinuo Cheng, Hai Helen Li, Yiran Chen ·

    MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models

    arXiv:2603.02482v2 Announce Type: replace-cross Abstract: Safety evaluation of multimodal large language models requires tracking not only whether an attack succeeds, but also how the interaction unfolds across turns and input modalities. We present MUSE (Multimodal Unified Safet…

  248. arXiv cs.CL TIER_1 English(EN) · Dikshit Chauhan, Bapi Dutta, Indu Bala, Niki van Stein, Thomas B\"ack, Anupam Yadav ·

    Large Language Models and Evolutionary Computation: A Critical Review of Bidirectional Interaction, Automated Algorithm Design, and Co-Adaptive Systems

    arXiv:2505.15741v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) and Evolutionary Computation (EC) are increasingly being combined to support automated optimization, algorithm design, and adaptive decision-making. This survey reviews the bidirectional intera…

  249. arXiv cs.CL TIER_1 English(EN) · Edward J. Hu, Moksh Jain, Eric Elmoznino, Younesse Kaddar, Guillaume Lajoie, Yoshua Bengio, Esmeralda S. Whitammer ·

    Amortizing intractable inference in large language models

    arXiv:2310.04363v3 Announce Type: replace-cross Abstract: Autoregressive large language models (LLMs) compress knowledge from their training data through next-token conditional distributions. This limits tractable querying of this knowledge to start-to-end autoregressive sampling…

  250. arXiv cs.CL TIER_1 English(EN) · Abdelrahman Abdallah, Mohammed Ali, Bhawna Piryani, Mahmoud Abdalla, Adam Jatowt ·

    Large Language Models Systematically Favor Popular Options: Evidence and Mitigation Across MCQs

    arXiv:2608.29257v1 Announce Type: new Abstract: Multiple-choice questions (MCQs) are a standard format for evaluating large language models (LLMs), yet the popularity of answer options can confound evaluation. Modern LLMs systematically prefer popular but incorrect options over l…

  251. arXiv cs.CL TIER_1 English(EN) · Ryo Kishino, Riku Shiomi, Hiroaki Yamagiwa, Momose Oyama, Hidetoshi Shimodaira ·

    Domain Mixture Design via Log-Likelihood Differences for Aligning Language Models with a Target Model

    arXiv:2603.16622v2 Announce Type: replace Abstract: Instead of directly distilling a language model, this study addresses the problem of aligning a base model with a target model in distribution by designing the domain mixture of training data for pretraining or continued pretrai…

  252. arXiv cs.CL TIER_1 English(EN) · Stanley Z. Hua, Sanae Lotfi, Irene Y. Chen ·

    Investigating Social Bias Changes in Quantized Language Models

    arXiv:2602.06181v2 Announce Type: replace Abstract: Post-training quantization reduces the memory needed to run large language models but alters their social biases in ways that aggregate metrics fail to capture. We present the first large-scale study of 50 quantized models evalu…

  253. arXiv cs.CL TIER_1 English(EN) · Vil\'em Zouhar, Tom Kocmi ·

    Pearmut: Human Evaluation of Translation Made Trivial

    arXiv:2601.02933v4 Announce Type: replace Abstract: Human evaluation is the gold standard for multilingual NLP, but is often skipped in practice and substituted with automatic metrics because it is notoriously complex and slow to set up with existing tools with substantial engine…

  254. arXiv cs.CL TIER_1 English(EN) · Ej Zhou, Caiqi Zhang, Tiancheng Hu, Chengzu Li, Nigel Collier, Ivan Vuli\'c, Anna Korhonen ·

    Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models

    arXiv:2510.03136v2 Announce Type: replace Abstract: Confidence calibration, the alignment of a model's predicted confidence with its actual accuracy, is crucial for the reliable deployment of Large Language Models (LLMs). However, this critical property remains largely under-expl…

  255. arXiv cs.CL TIER_1 English(EN) · Yuzhe Ding, Kang He, Li Zheng, Shengwu Zheng, Teng Shi, Fei Li, Chong Teng, Donghong Ji ·

    MMDS-Bench: Benchmarking Multimodal Large Language Models on Dynamic Stance in Social Media Interactions

    arXiv:2608.30903v1 Announce Type: new Abstract: Dynamic stance classification models how a reply responds to its direct parent message, rather than how a post relates to a fixed topic. Existing work has mainly studied this problem in text-only settings, while social media interac…

  256. arXiv cs.CL TIER_1 English(EN) · Kangwook Ko, Jaehyuk Jang, Wonjun Lee, Hee-Seon Kim, Changick Kim ·

    Where Identity Lives: Localized, Retain-Free Identity Unlearning in Multimodal Large Language Models

    arXiv:2608.30649v1 Announce Type: new Abstract: Removing a specific individual's information from multimodal large language models (MLLMs) is often needed after deployment, but existing methods rely on a retain set, which is hardest to obtain at that point, and rebuilding it recr…

  257. arXiv cs.AI TIER_1 English(EN) · Jinzhe Li, Gengxu Li, Jinnan Li, Yuan Wu, Yi Chang ·

    MMPCBench: Benchmarking Multimodal Large Language Models on Proactive Critique of Flawed Inputs

    arXiv:2608.29286v1 Announce Type: new Abstract: As Multimodal Large Language Models (MLLMs) evolve into sophisticated interactive assistants, their reliability depends not only on following instructions but also on validating them. We define Proactive Critique as the model's auto…

  258. arXiv cs.AI TIER_1 English(EN) · Xunyi Jiang, Junda Wu, Yuxin Xiong, Sheldon Yu, Tong Yu, David Arbour, Ritwik Sinha, Julian McAuley, Hongyi Wen ·

    Toward Latent Language Model Skills Steering and Optimization: An Empirical Study

    arXiv:2608.29459v1 Announce Type: new Abstract: Skills, as a useful abstraction for the procedural capabilities of large language models (LLMs), capture how models perform structured, multi-step reasoning and program execution. Existing approaches typically treat skills as explic…

  259. arXiv cs.CL TIER_1 English(EN) · Bruno Leonardo Santos Menezes, Carlos Leonardo Souza Cardoso, Fabio Andre Machado Porto ·

    Manac\'a-1B: An Open, Reproducible Brazilian-Portuguese Language Model and a Tokenizer-Aware, Paired Evaluation

    arXiv:2608.30114v1 Announce Type: new Abstract: Brazilian Portuguese remains under-served by open language models, and the few that exist are difficult to reproduce and are often compared without measures of uncertainty. We release Manac\'a-1B, an open decoder-only model of 1.72 …

  260. arXiv cs.CL TIER_1 English(EN) · Kyungdon Lee, Wei Xu, Alan Ritter, Dong-Ho Lee, JinYeong Bak ·

    CoCoA: Context-Conditional Cultural Alignment for Large Language Models

    arXiv:2608.29492v1 Announce Type: new Abstract: Large Language Models (LLMs) often favor Western-associated entities across cultural contexts. Conventional debiasing methods aim for uniform neutrality, but cultural bias mitigation demands context-conditional behavior, preferring …

  261. arXiv cs.CL TIER_1 English(EN) · Zhangdie Yuan, Andreas Vlachos ·

    Evaluating the Semantic Specificity of Representation Steering in Language Models

    arXiv:2608.29431v1 Announce Type: new Abstract: Localized Representation Steering (LRS) is widely used to correct reasoning pathologies in large language models. However, standard benchmark evaluations can easily be fooled by superficial label overrides, creating a false impressi…

  262. Hugging Face Daily Papers TIER_1 English(EN) ·

    EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

    EM²Mem binds multimodal evidence to event anchors for compact, generation-ready memory in long-video question answering.

  263. Hugging Face Daily Papers TIER_1 English(EN) ·

    H3-World: Turning Language Understanding into World Control

    We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world model. Our key finding is that, as large video generators become more capable, language is emerging as a natural interface for control. MiniMax-H3, for example, alr…

  264. Hugging Face Daily Papers TIER_1 English(EN) ·

    CARVE: Verified Expansion for Variable-Length Generation in Diffusion Language Models

    Masked diffusion language models predict tokens from a partially observed response canvas, enabling bidirectional conditioning and parallel token refinement. Yet standard masked-diffusion decoders use a rigid inference interface: the number of masked positions allocated to the an…

  265. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Tony Wu ·

    NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference

    Multimodal models often build on architectures designed for generative vision-language modeling, typically combining separately pretrained vision encoders with causal language models. Visual document retrievers such as ColPali repurpose these models as encoders, carrying over the…

  266. Hugging Face Daily Papers TIER_1 English(EN) ·

    NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference

    Multimodal models often build on architectures designed for generative vision-language modeling, typically combining separately pretrained vision encoders with causal language models. Visual document retrievers such as ColPali repurpose these models as encoders, carrying over the…

  267. Hugging Face Daily Papers TIER_1 English(EN) ·

    Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability

    Large language models (LLMs) are built from structured high-dimensional objects such as token representations, weights, adaptation updates, caches, and activations, whose multilinear structure is underexploited by the conventional matrix-centric view. Tensor decompositions and te…

  268. arXiv cs.CL TIER_1 English(EN) · Mengzhe Geng ·

    SURE-Challenge: Evaluating Speech Evidence Before Speech-LLM Generation

    arXiv:2608.27783v1 Announce Type: cross Abstract: Speech LLMs are usually graded after they answer, although an operating system first has to decide whether a waveform should be sent to the model. We define the Speech-Unsupported Rejection Evaluation Challenge (SURE-Challenge) fo…

  269. arXiv cs.LG TIER_1 English(EN) · Yuangang Li, Justin Tian Jin Chen, Ethan Yu, David Hong, Iftekhar Ahmed ·

    Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks

    arXiv:2604.12379v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly rely on explicit reasoning to solve coding tasks, yet evaluating the quality of this reasoning remains challenging. Existing reasoning evaluators are not designed for coding, and c…

  270. arXiv cs.CL TIER_1 English(EN) · Titus von der Malsburg, Sebastian Pad\'o ·

    Diverging Transformer Predictions for Human Sentence Processing: A Comprehensive Analysis of Agreement Attraction Effects

    arXiv:2603.16574v2 Announce Type: replace Abstract: Transformers underlie almost all state-of-the-art language models in computational linguistics, yet their cognitive adequacy as models of human sentence processing remains disputed. In this work, we use a surprisal-based linking…

  271. arXiv cs.CL TIER_1 English(EN) · Biaojie Zeng, Min Zhang, Juan Zhou, Fengrui Liu, Ruiyang Huang, Yu Song, Xin Lin ·

    SMRC: Aligning Large Language Models with Student Reasoning for Mathematical Error Correction

    arXiv:2511.14684v2 Announce Type: replace Abstract: Large language models (LLMs) often make reasoning errors when solving mathematical problems, and how to automatically detect and correct these errors has become an important research direction. However, existing approaches \text…

  272. arXiv cs.CL TIER_1 English(EN) · Zheyuan Liu, Zhangchen Xu, Guangyao Dou, Xiangchi Yuan, Zhaoxuan Tan, Radha Poovendran, Meng Jiang ·

    Steering Multimodal Large Language Models Decoding for Context-Aware Safety

    arXiv:2509.19212v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed in real-world applications, yet their ability to make context-aware safety decisions remains limited. Existing methods often fail to balance oversensitivity (unj…

  273. arXiv cs.CL TIER_1 English(EN) · Ayan Sengupta, Siddhant Chaudhary, Tanmoy Chakraborty ·

    Pruning Laws for Large Language Models

    arXiv:2504.04342v2 Announce Type: replace Abstract: Scaling up model parameters and training data consistently improves the performance of large language models (LLMs), but at the cost of rapidly growing memory and compute requirements, which makes deployment on resource-limited …

  274. arXiv cs.CL TIER_1 English(EN) · Xiaotian Zhang, Chunyang Li, Yi Zong, Zhengyu Ying, Liang He, Xipeng Qiu, Tianxiang Sun, Peng Li, Shiqiao Meng, Yanjun Zheng, Jun Zhan, Zhangyue Yin, Xiannian Hu, Guofeng Quan, Qixiang Wang ·

    Evaluating the Performance of Large Language Models on GAOKAO Benchmark

    arXiv:2305.12474v4 Announce Type: replace Abstract: Large Language Models(LLMs) have demonstrated remarkable performance across various natural language processing tasks; however, how to comprehensively and accurately assess their performance becomes an urgent issue to be address…

  275. arXiv cs.CL TIER_1 English(EN) · Wonjun Lee, Jaehyuk Jang, Kangwook Ko, Hee-Seon Kim, Changick Kim ·

    AIM: Anchor Identity Features, Then Match for Multimodal Large Language Model Unlearning

    arXiv:2608.28312v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) can memorize identity-specific facts about people in their fine-tuning data, creating privacy risks when a person requests deletion. Existing MLLM unlearning methods often assume access to …

  276. arXiv cs.CL TIER_1 English(EN) · Justin Bronder ·

    Why Didn't It Check? Unsupported Final Claims and Their Repair in Two Tool-Equipped Language Models

    arXiv:2608.27768v1 Announce Type: cross Abstract: A language model with access to tools can commit to a final claim unsupported by the evidence it has seen, even when a single available tool call would resolve the uncertainty and its instructions explicitly forbid assumptions and…

  277. arXiv cs.CL TIER_1 English(EN) · Emily Cheng, Ryan Cotterell ·

    A Formal Limitation on Learning Human Language From Textual Corpora

    arXiv:2608.28560v1 Announce Type: new Abstract: Can a listener recover what a speaker means from the form of an utterance alone? We answer this question information-theoretically, and for a listener given by any featurizer of text, including the hidden states of contemporary larg…

  278. arXiv cs.CL TIER_1 English(EN) · Di Wu, Sergey Troshin, Christof Monz, Antske Fokkens, Vlad Niculae ·

    Ladders in Chaos: When, How, (and Perhaps Why) Does Test-Time Scaling Improve LLM Machine Translation

    arXiv:2608.28496v1 Announce Type: new Abstract: Two forms of test-time scaling for Large Language Models (LLMs) have emerged as effective and widely adopted paradigms: sequential, in which later answer attempts depend on earlier ones, and parallel, such as i.i.d. sampling with re…

  279. arXiv cs.CL TIER_1 English(EN) · Hefan Zhang, Bingquan Zhang, Ming Cheng, Saeed Hassanpour, Weicheng Ma, Soroush Vosoughi ·

    When Linguistic and Internal Confidence Diverge in Large Language Models

    arXiv:2608.28382v1 Announce Type: new Abstract: Users often ask large language models (LLMs) to report how confident they are, but it is unclear whether such linguistic confidence tracks the model's internal confidence. We study this question across 8 classification tasks, 2 gene…

  280. Hugging Face Daily Papers TIER_1 English(EN) ·

    Manacá-1B: An Open, Reproducible Brazilian-Portuguese Language Model and a Tokenizer-Aware, Paired Evaluation

    Brazilian Portuguese remains under-served by open language models, and the few that exist are difficult to reproduce and are often compared without measures of uncertainty. We release Manacá-1B, an open decoder-only model of 1.72 billion parameters trained from scratch for Brazil…

  281. Hugging Face Daily Papers TIER_1 English(EN) ·

    NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference

    NeoMME introduces small bidirectional multimodal encoders pretrained with masked discrete diffusion that achieve strong visual document retrieval and high compression of late-interaction embeddings.

  282. arXiv cs.AI TIER_1 English(EN) · Arseniy Varlamov, Rishat Zinnatullin, Elisei Rykov, Alexander Panchenko, Ilseyar Alimova ·

    MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models

    arXiv:2608.26295v1 Announce Type: cross Abstract: Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their parametric memory, yet existing evaluations measure source preference without establishing source correctness. We introduce Me…

  283. arXiv cs.AI TIER_1 English(EN) · Dai Shi, Xiaoyu Li, Jos\'e Miguel Hern\'andez-Lobato ·

    When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models

    arXiv:2608.26187v1 Announce Type: cross Abstract: Whether large language models (LLMs) can perform the abductive leap from evidence to a new system of axioms, commonly referred to as a jump, has recently attracted considerable debate. A prominent position holds that LLMs are stru…

  284. arXiv cs.AI TIER_1 English(EN) · Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang ·

    Assessing mentalization in humans and large language models

    arXiv:2608.26291v1 Announce Type: new Abstract: Mentalization - the ability to infer others' beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with hu…

  285. arXiv cs.AI TIER_1 English(EN) · Zihao Guo, Hongtao Lv, Chaoli Zhang, Laiguo Yin, Lei Liu, Yonghui Xu, Lizhen Cui ·

    Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Models

    arXiv:2608.26161v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) demonstrate remarkable capabilities in reasoning and decision-making, high-fidelity probabilistic sampling remains a persistent challenge. When generating random variables, LLMs consistently e…

  286. arXiv cs.CL TIER_1 English(EN) · Hsiao-Ying Huang, Cheng-Han Chiang, Hung-yi Lee ·

    SPAR-K: Scheduled Periodic Alternating Early Exit for Spoken Language Models

    arXiv:2603.09215v2 Announce Type: replace Abstract: Interleaved spoken language models (SLMs) alternately generate text and speech tokens, but decoding at full transformer depth for every step becomes costly, especially due to long speech sequences. We propose SPAR-K, a modality-…

  287. arXiv cs.CL TIER_1 English(EN) · Ruirui Chen, Weifeng Jiang, Chengwei Qin, Bo Xiong, Kaiwen Wei, Fiona Liausvia, Pei Fang Tan, Ker Yung Chua, Dongkyu Choi, Mukkesh Kumar, Evelyn C. Law, Dennis Wang, Boon Kiat Quek ·

    Are Large Language Models Effective Knowledge Graph Constructors?

    arXiv:2510.11297v2 Announce Type: replace Abstract: Knowledge graphs (KGs) are widely used in knowledge-intensive applications, yet it remains unclear how effectively current large language models (LLMs) can construct document-grounded KGs in a zero-shot, schema-free setting with…

  288. arXiv cs.CL TIER_1 English(EN) · Yufan Wu, Yinghui He, Zhengyi Hu, Lang Wei, Ruichen Li, Qifan Yang, Ting Zhu ·

    CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

    arXiv:2608.27455v1 Announce Type: new Abstract: Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models (LLMs). However, these methods typically rely on repeated generation or external verification. To address this …

  289. arXiv cs.CL TIER_1 English(EN) · V\'esteinn Sn{\ae}bjarnarson, Samuel Kiegeland, Manuel de Prada Corral, Ryan Cotterell, Tim Vieira ·

    Stochastic Estimation of Transduced Language Models

    arXiv:2608.27428v1 Announce Type: new Abstract: Transduced language models (TLMs) compose a pretrained \emph{source} language model with a functional finite-state transducer to induce a language model over \emph{target} strings. Computing the probability of a target prefix under …

  290. arXiv cs.CL TIER_1 English(EN) · Maxime Bouthors, Josep Crego, Fran\c{c}ois Yvon ·

    Reasoning about In-Context Samples for Machine-Translation

    arXiv:2608.27036v1 Announce Type: new Abstract: Large Language Models (LLMs) can be trained to perform chain-of-thoughts reasoning in order to improve the reliability of their responses. In this work, we investigate how explicit reasoning can be leveraged for LLM-Based Machine Tr…

  291. arXiv cs.CL TIER_1 English(EN) · Bohan Yu, Shi-Yang Li, Pengfei Cao, Jun Zhao, Kang Liu ·

    RuleWeaver: Benchmarking Rule-Centered Scenario Reasoning for Large Language Models

    arXiv:2608.26832v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied to specialized domains, where effective use of domain expertise often requires reasoning over complex rules in concrete scenarios. However, existing benchmarks only partially eva…

  292. arXiv cs.CL TIER_1 English(EN) · Yufei Sun, Yudong Li, Yiming Cheng ·

    SPT: Skills as Pre-Training Data for Agentic Language Models

    arXiv:2608.26563v1 Announce Type: new Abstract: Agentic (tool-using) language models are mainly trained on tool-call traces and agent trajectories during post-training. These data provide direct behavioral supervision, but producing them requires task environments, execution, and…

  293. arXiv cs.CL TIER_1 English(EN) · Yefan Tao, Gerald Friedland, Luyang Kong ·

    When Is Noise Response Universal? Tokenization as the Hidden Variable in Language Models

    arXiv:2608.26319v1 Announce Type: new Abstract: The performance of textual neural models often degrades when their inputs are corrupted by noise such as typos, OCR errors, or dropped words. We study the degradation rate across neural models, both sentence embeddings and decoder-o…

  294. arXiv cs.CL TIER_1 English(EN) · Yucheng Zhou, Peng Luo, Qianning Wang, Chengzhong Xu, Jianbing Shen ·

    CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models

    arXiv:2608.26147v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong potential for medical reasoning, yet the scarcity and cost of expert-annotated data constrain their progress. While reinforcement learning offers a scalable alternative, standard outcom…

  295. arXiv cs.CL TIER_1 English(EN) · Ilija Subasic, Andrew Rabinovich, Zhao Chen ·

    Evaluating Language Models in Realistic Conversational Contexts

    arXiv:2608.26131v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed to serve open-ended, multi-turn interactions, evaluating conversational quality at human scale has become a central challenge. Existing evaluation frameworks built for summar…

  296. arXiv cs.CL TIER_1 English(EN) · Ziqiang Zhang, Jing Ma, Zilong Wang, Jiayuan Chen, Yi Qiao, Yu He, Wei Zhang, Dai Cheng, Xiaoyu Shen ·

    Natural-Language Policies to Executable Decisions: An Interpretable Large Language Model Framework

    arXiv:2608.26124v1 Announce Type: new Abstract: Pricing automation in large-scale tourism is challenging because travel orders are highly unstructured, while pricing policies are complex, rapidly evolving, and inherently open-ended. Traditional rule engines are brittle and costly…

  297. arXiv cs.AI TIER_1 English(EN) · Ryan Thomas Noonan, Linxi Zhao, Menghan Xu, Akanksha Sarkar, Mihir Mishra, Dongyoung Go, Kilian Q. Weinberger, Yoav Artzi, Jennifer J. Sun ·

    Co-Evolving Structured Knowledge and Reasoning in Language Models

    arXiv:2608.26386v1 Announce Type: cross Abstract: Retrieval-augmented methods improve factual accuracy by grounding language models in external knowledge, but retrieving over unstructured text often introduces irrelevant context and offers limited control over the retrieved infor…

  298. arXiv cs.AI TIER_1 English(EN) · Christos Petridis, Konstantinos Pelechrinis, Zoran Obradovic ·

    How Unlikely Is "Unlikely"? Assessing Verbal Probability Perception Across Large Language Models

    arXiv:2608.26327v1 Announce Type: cross Abstract: Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty) remains unknown. We present …

  299. Hugging Face Daily Papers TIER_1 English(EN) ·

    Evaluating the Hidden Costs of Personalization in Large Language Models

    The study proposes PRISK, a framework that reveals how personalized context in LLMs increases irrelevant personalization, preference narrowing, and sycophantic bias.

  300. arXiv cs.LG TIER_1 English(EN) · Dingzhi Yu, Rui Pan, Yuxing Liu, Difan Zou, Tong Zhang ·

    StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models

    arXiv:2604.15416v2 Announce Type: replace Abstract: Sign-based optimization algorithms, such as SignSGD, have garnered attention for their performance in distributed learning and training large foundation models. Despite their empirical superiority, SignSGD is known to diverge on…

  301. arXiv cs.LG TIER_1 English(EN) · Molka Chkir, Syed Muhammad Danish, Jos H\"oll, Arghavan Asad ·

    Understanding the Energy Scaling of Large Language Model Inference Across Context Lengths and Attention Architectures

    arXiv:2608.25096v1 Announce Type: new Abstract: The growing adoption of large language models (LLMs) has raised increasing concerns about the energy consumption and environmental impact of inference. This paper presents a systematic empirical study of decode-phase energy consumpt…

  302. arXiv cs.CL TIER_1 Deutsch(DE) · Leonardo Berti, Flavio Giorgi, Gjergji Kasneci ·

    Emergent Abilities in Large Language Models: A Survey

    arXiv:2503.05788v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are leading a new technological revolution as one of the most promising research streams toward artificial general intelligence. The scaling of these models, accomplished by increasing the numb…

  303. arXiv cs.CL TIER_1 English(EN) · Tianyi Zhao, Yinhan He, Wendy Zheng, Chen Chen ·

    Addressing the Reasoning Gap: Mechanistic Circuit-Based Knowledge Editing in Large Language Models

    arXiv:2604.05876v2 Announce Type: replace Abstract: Deploying Large Language Models (LLMs) in real-world dynamic environments raises the challenge of updating their pre-trained knowledge. While existing knowledge editing methods can reliably patch isolated facts, they frequently …

  304. arXiv cs.CL TIER_1 English(EN) · Xukai Liu, Ye Liu, Jipeng Zhang, Yanghai Zhang, Kai Zhang, Qi Liu ·

    Layer-Order Inversion: Rethinking Latent Multi-Hop Reasoning in Large Language Models

    arXiv:2601.03542v2 Announce Type: replace Abstract: Large language models (LLMs) perform well on multi-hop reasoning, yet how they internally compose multiple facts remains unclear. Recent work proposes \emph{hop-aligned circuit hypothesis}, suggesting that bridge entities are co…

  305. arXiv cs.CL TIER_1 English(EN) · Hyunji Nam, Lucia Langlois, James Malamut, Mei Tan, Dorottya Demszky ·

    IDEAlign: Comparing Ideas of Large Language Models to Domain Expert

    arXiv:2509.02855v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to produce open-ended, interpretive annotations, yet there is no validated, scalable measure of idea-level similarity to expert annotations. We (i) introduce the content evaluat…

  306. arXiv cs.CL TIER_1 English(EN) · Ming Ma, Bowen Zheng, Zhongqiao Lin, Tianming Yang ·

    SimLens for Early Exit in Large Language Models: Eliciting Accurate Latent Predictions with One More Token

    arXiv:2507.17618v3 Announce Type: replace Abstract: Intermediate-layer predictions in large language models (LLMs) are informative but hard to decode accurately, especially at early layers. Existing lens-style methods typically rely on direct linear readout, which is simple but o…

  307. arXiv cs.CL TIER_1 English(EN) · Christo Kirov, Ryan Cotterell ·

    Recurrent Neural Networks in Linguistic Theory: Revisiting Pinker and Prince (1988) and the Past Tense Debate

    arXiv:1807.04783v3 Announce Type: replace Abstract: Can advances in NLP help advance cognitive modeling? We examine the role of artificial neural networks, the current state of the art in many common NLP tasks, by returning to a classic case study. In 1986, Rumelhart and McClella…

  308. arXiv cs.CL TIER_1 English(EN) · Rui He, Nihal Altay, Wolfram Hinzen ·

    Distinct dynamics of conceptual and referential disruptions in human reading and large language model processing

    arXiv:2608.25999v1 Announce Type: new Abstract: Linguistic meaning is grounded in conceptual content, from which reference to particular entities emerges as words enter discourse. To examine the processing dynamics associated with these two dimensions of meaning, we selectively d…

  309. arXiv cs.CL TIER_1 English(EN) · Muge Zhang, Aaron Jencks, Krishna Badikela, Yulia Tsvetkov, Sachin Kumar ·

    One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography

    arXiv:2608.25904v1 Announce Type: new Abstract: Multilingual language models transfer knowledge across languages through shared subword vocabulary, a mechanism that breaks down when related languages use different writing systems. Prior work addresses this via script equalization…

  310. arXiv cs.CL TIER_1 English(EN) · Tim Schopf, Tobias Schreieder, Akiko Aizawa ·

    Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty

    arXiv:2608.25660v1 Announce Type: new Abstract: Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we investigate a prev…

  311. arXiv cs.CL TIER_1 English(EN) · Martino Ciaperoni, Sezer Kutluk, Benedetta Muscato, Marta Marchiori Manerba, Fosca Giannotti ·

    Virgil: Navigating Explainability for Transformer-based Language Models

    arXiv:2608.25555v1 Announce Type: new Abstract: Explainability for transformer-based language models is becoming crucial as these systems are deployed in high-stakes applications. As a result, the ecosystem of explainability tools is rapidly evolving, becoming richer, but also mo…

  312. arXiv cs.CL TIER_1 English(EN) · Xiulin Yang, Ethan Gotlieb Wilcox, Catherine Arnett ·

    Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation

    arXiv:2608.25089v1 Announce Type: new Abstract: Crosslingual evaluation of language models that enables fair comparisons remains a fundamental challenge in multilingual NLP. Existing studies adopt a variety of downstream tasks and intrinsic metrics with different theoretical just…

  313. arXiv cs.CL TIER_1 English(EN) · Kaiqiao Han, Yizhou Sun ·

    The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure

    arXiv:2608.25005v1 Announce Type: new Abstract: The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing t…

  314. arXiv cs.CL TIER_1 English(EN) · Elle ·

    The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline

    arXiv:2608.24952v1 Announce Type: new Abstract: Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear. Our study traces this "dialect tax" across the nat…

  315. Hugging Face Daily Papers TIER_1 English(EN) ·

    SPT: Skills as Pre-Training Data for Agentic Language Models

    Agentic (tool-using) language models are mainly trained on tool-call traces and agent trajectories during post-training. These data provide direct behavioral supervision, but producing them requires task environments, execution, and verification, making broad tool and task covera…

  316. Hugging Face Daily Papers TIER_1 English(EN) ·

    CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

    CritICL improves LLM reasoning at inference time by using structured failure patterns from weaker models as critique-based guidance, reducing generation and token costs.

  317. Hugging Face Daily Papers TIER_1 English(EN) ·

    How Unlikely Is "Unlikely"? Assessing Verbal Probability Perception Across Large Language Models

    Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty) remains unknown. We present a systematic cross-model evaluation using a word-t…

  318. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Ilseyar Alimova ·

    MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models

    Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their parametric memory, yet existing evaluations measure source preference without establishing source correctness. We introduce MemToC, a controlled benchmark for post-tool-return …

  319. arXiv cs.AI TIER_1 English(EN) · Georges Sfeir, Gabriel Nova, Stephane Hess, Sander van Cranenburgh ·

    Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities

    arXiv:2507.21790v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are becoming widely used to support various workflows across different disciplines, yet their potential in discrete choice modelling remains relatively unexplored. This work examines the potent…

  320. arXiv cs.AI TIER_1 English(EN) · Toghrul Abbasli, Kentaroh Toyoda, Yuan Wang, Leon Witt, Muhammad Asif Ali, Yukai Miao, Dan Li, Qingsong Wei ·

    Comparing Uncertainty Measurement and Mitigation Methods for Large Language Models: A Systematic Review

    arXiv:2504.18346v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have been transformative across many domains. However, hallucination, i.e., confidently outputting incorrect information, remains one of the leading challenges for LLMs. This raises the questio…

  321. arXiv cs.AI TIER_1 English(EN) · Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi ·

    The Limits of Automatic Evaluation of Creativity in Large Language Models

    arXiv:2608.23705v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly capable of generating text that challenges human performance in domains requiring creativity, yet evaluating creativity in LLM-generated content remains a significant challenge. Here, …

  322. arXiv cs.AI TIER_1 English(EN) · Esmail Gumaan ·

    Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail

    arXiv:2608.23651v1 Announce Type: cross Abstract: Agent harnesses record a failed tool call and its error message in the transcript and ask the model to continue, on the assumption that the error is corrective information. We measure whether it is. Defining the corrective gain of…

  323. arXiv cs.AI TIER_1 English(EN) · Augusto Camargo ·

    The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models

    arXiv:2608.24662v1 Announce Type: new Abstract: Large language models (LLMs) are commonly evaluated under the assumption that their observable behavior is primarily determined by model weights, training data, alignment procedures, and user prompts. This view is incomplete. Modern…

  324. arXiv cs.LG TIER_1 English(EN) · Zihan Liu, Ruiheng Zheng, Shaobo Zhang, Changxin Tian, Kunlong Chen, Zhiqiang Zhang, Lei Wu ·

    Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

    arXiv:2608.24814v1 Announce Type: new Abstract: We uncover ELR collapse in language model pretraining: learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR). When ELR is matched across runs, their loss trajecto…

  325. arXiv cs.AI TIER_1 English(EN) · Harry Owiredu-Ashley ·

    ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models

    arXiv:2603.10068v2 Announce Type: replace-cross Abstract: Most adversarial evaluations of large language model (LLM) safety assess single prompts and report binary pass/fail outcomes, which fails to capture how safety properties evolve under sustained adversarial interaction. We …

  326. arXiv cs.AI TIER_1 English(EN) · Qianli Wang, Nils Feldhus, Pepa Atanasova, Fedor Splitt, Simon Ostermann, Sebastian M\"oller, Vera Schmitt ·

    Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations

    arXiv:2601.00282v2 Announce Type: replace-cross Abstract: Quantization is widely used to accelerate inference and streamline the deployment of large language models (LLMs), yet its effects on self-explanations (SEs) remain unexplored. SEs, generated by LLMs to justify their own o…

  327. arXiv cs.AI TIER_1 English(EN) · Marek Mateusz Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Demitri Africa ·

    Measuring Activation Control in Large Language Models

    arXiv:2608.21664v1 Announce Type: new Abstract: Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluations, especially when evaluation-aware models exhibit scheming or deception. However, if models …

  328. arXiv cs.CL TIER_1 English(EN) · Sara Rezaeimanesh, Kundan Thind, Farzan Siddiqui, Mohammad M. Ghassemi ·

    What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization

    arXiv:2601.17609v3 Announce Type: replace Abstract: In domains like medicine and finance, large-scale labeled data is costly and often unavailable, leading to models trained on small datasets that struggle to generalize to real-world populations. Large language models contain ext…

  329. arXiv cs.CL TIER_1 English(EN) · Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, Yuandong Tian ·

    Training Large Language Models to Reason in a Continuous Latent Space

    arXiv:2412.06769v4 Announce Type: replace Abstract: Large language models (LLMs) are typically constrained to reason in the language space, where they express the reasoning process through a chain-of-thought (CoT) to solve complex problems. However, the language space may not alw…

  330. arXiv cs.CL TIER_1 English(EN) · Nura Aljaafari, Danilo S. Carvalho, Andr\'e Freitas ·

    Bridging Linguistic Structure and Mechanistic Interpretability for Conceptual Interpretation in Language Models

    arXiv:2408.11827v2 Announce Type: replace Abstract: Understanding how language models compose meaning from linguistic input remains a central problem in interpretability research. Mechanistic studies have attributed functional roles to core transformer components; however, these …

  331. arXiv cs.AI TIER_1 English(EN) · Farzaneh Heidari, Amin Memarian, Guillaume Rabusseau ·

    Evaluation Awareness in Language Models: Representation, Verbalization, and Control

    arXiv:2608.21766v1 Announce Type: cross Abstract: Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their behavior in deployment. This assumption can fail, should models infer that they are b…

  332. arXiv cs.AI TIER_1 English(EN) · Jinghan Tan, Yuanzheng Wang, Lu Chen, Zijun Chen, Yuqian Wang, Maosong Sun ·

    StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models

    arXiv:2608.23475v1 Announce Type: new Abstract: As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot in-context learning (ICL) has become a key paradigm for task adaptation. However, direct ICL often uses a small set of examples with…

  333. arXiv cs.AI TIER_1 English(EN) · Zengqing Wu, Chuan Xiao ·

    Proxy reliance in large language model decisions is uncalibrated to predictive evidence

    arXiv:2608.22887v1 Announce Type: new Abstract: Large language models (LLMs) are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use. Current audits ask whether decisions change when demographics change. But a…

  334. arXiv cs.AI TIER_1 English(EN) · Samira Golsefid ·

    Measuring Stability and Failure Behavior in Language Models Under Structured Perturbations

    arXiv:2608.22138v1 Announce Type: new Abstract: Language models are usually judged by a single accuracy score, which does not reveal how their performance degrades as inputs are perturbed. We present a graded, multi-family, failure-aware framework for stress-testing reasoning mod…

  335. arXiv cs.CL TIER_1 English(EN) · Toghrul Abbasli, Kentaroh Toyoda, Yuan Wang, Li Chen ·

    Claim-Level Confidence Calibration for Reliable Decision Making with Large Language Models

    arXiv:2608.22483v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly support decision-making in high-stakes domains, but they often hallucinate and express confidence that is misaligned with factual correctness. Response-level confidence is a coarse signal: a…

  336. arXiv cs.AI TIER_1 English(EN) · Joosung Lee, Hwiyeol Jo, Donghyeon Ko, Kyubyung Chae, Cheonbok Park, Jeonghoon Kim ·

    What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"

    arXiv:2604.05779v2 Announce Type: replace-cross Abstract: While large language models (LLMs) demonstrate strong capabilities across diverse user queries, they still suffer from hallucinations, often arising from knowledge misalignment between pre-training and fine-tuning. To addr…

  337. arXiv cs.AI TIER_1 English(EN) · Bo Yuan, Jiazi Hu ·

    Beyond Benchmarking: Scenario-Based Evaluation of Large Language Models for Personalized Learning

    arXiv:2509.05346v3 Announce Type: replace Abstract: While large language models (LLMs) are increasingly being adopted to support personalized learning, there remains limited understanding of how their pedagogical behaviors differ in authentic learning scenarios. Existing evaluati…

  338. arXiv cs.AI TIER_1 English(EN) · Ziyue Wang (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Aomufei Yuan (Peking University), Yiran Yao (Tianjin University), Linli Yao (State Key Laboratory of Multimedia Information Processing,… ·

    What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

    arXiv:2608.22948v1 Announce Type: cross Abstract: Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later pape…

  339. arXiv cs.AI TIER_1 English(EN) · Jiaqian Zhu, Yang Zhang, Junhua Ding, Xiaowei Yu ·

    Lexical Perturbations Disrupt LLM Reasoning: An Empirical Study of Attention Diversion

    arXiv:2608.22140v1 Announce Type: cross Abstract: Large Language Models (LLMs) achieve strong reasoning performance, but their robustness to realistic lexical corruption remains poorly understood. We evaluate four open-weight instruction-tuned models and frontier models across fo…

  340. Hugging Face Daily Papers TIER_1 English(EN) ·

    Proxy reliance in large language model decisions is uncalibrated to predictive evidence

    Large language models (LLMs) are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use. Current audits ask whether decisions change when demographics change. But attributes correlated with a protected group carr…

  341. arXiv cs.AI TIER_1 (CA) · Zhicheng Lin ·

    Six misconceptions about large language models: A minimal model and diagnostic taxonomy

    arXiv:2608.20421v1 Announce Type: cross Abstract: Large language models (LLMs) are now embedded in scientific, educational, and governance workflows, with debates centering on their capabilities, mechanisms, and impacts. Yet these debates remain structured by persistent folk theo…

  342. arXiv cs.CL TIER_1 English(EN) · Karthik Viswanathan, Yuri Gardinazzi, Giada Panerai, Alberto Cazzaniga, Matteo Biagetti ·

    The Intrinsic Dimension of Prompts in Internal Representations of Large Language Models

    arXiv:2501.10573v2 Announce Type: replace Abstract: We study the geometry of token representations at the prompt level in large language models through the lens of intrinsic dimension. Viewing transformers as mean-field particle systems, we estimate the intrinsic dimension of the…

  343. arXiv cs.AI TIER_1 English(EN) · Emilio Ferrara ·

    Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation

    arXiv:2608.20569v1 Announce Type: new Abstract: Are frontier models able to introspect about their internal states? Recent work suggests that under certain conditions a complex enough model can audit its own internals, call out what changed, and report back confidently about it. …

  344. arXiv cs.CL TIER_1 English(EN) · Yinfeng Wang, Zhiyuan Yao, Zheren Fu, Lei Zhang, Zhendong Mao ·

    When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models

    arXiv:2608.19208v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored. In this paper, we investigate the influence of task-irrelevant conte…

  345. arXiv cs.LG TIER_1 Nederlands(NL) · Jaiden Fairoze, Neal Mangaokar, Kamalika Chaudhuri, Sanjam Garg, Saeed Mahloujifar ·

    Inadvertent Context Leakage in Language Models

    arXiv:2608.19857v1 Announce Type: new Abstract: For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere presence of such secrets in a model's context window …

  346. arXiv cs.LG TIER_1 English(EN) · Kaizhen Tan, Xin Xu, Siru Tao, Yixiao Li, Hanzhe Hong, Yang Feng, Heqing Du ·

    Beyond Multimodal Alignment: Certifying Physical Language through Response Substitution and Ordered Execution

    arXiv:2608.19492v1 Announce Type: new Abstract: World models increasingly treat compact multimodal representations as interfaces between perception and physical interaction, yet existing probes do not establish whether different sensors carry the same executable meaning or whethe…

  347. arXiv cs.CL TIER_1 English(EN) · Sahil Kale, Ian Harris ·

    ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

    arXiv:2608.20338v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely. Current approaches rely on di…

  348. arXiv cs.CL TIER_1 English(EN) · Dingzirui Wang, Xuanliang Zhang, Keyan Xu, Qingfu Zhu, Wanxiang Che ·

    FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

    arXiv:2608.20153v1 Announce Type: new Abstract: Large language models (LLMs) have shown growing potential for automated theoretical computer science (TCS) research, yet existing benchmarks remain far from realistic research settings. We introduce \ourbenchmark, an expert-validate…

  349. arXiv cs.CL TIER_1 English(EN) · Mattia Carletti, Edward Phillips, Fredrik K. Gustafsson, Patitapaban Palo, Lei Clifton, Danielle Belgrave, Xiao Gu, David A. Clifton ·

    When Text and Numbers Disagree: Evidence Arbitration in Large Language Models

    arXiv:2608.20116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence. We study how LLMs arbitrate between such sources when they s…

  350. arXiv cs.CL TIER_1 English(EN) · Nikita Khudov ·

    OenoBench: A Wine-Domain Benchmark for Knowledge-Grounded Evaluation of Large Language Models

    arXiv:2608.20106v1 Announce Type: new Abstract: We introduce OenoBench, a wine-domain knowledge benchmark of 3,266 multiple-choice questions across six pillars (regions, grape varieties, viticulture, winemaking, producers, business) and four difficulty tiers. The corpus is built …

  351. arXiv cs.AI TIER_1 English(EN) · Yao Liang, Dongcheng Zhao, Feifei Zhao, Guobin Shen, Yuwei Wang, Dongqi Liang, Yi Zeng ·

    DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values

    arXiv:2509.08022v3 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with diverse human values is essential for safe and effective deployment, yet existing benchmarks often overlook cultural and demographic variation. We introduce DiverValue-Bench, a po…

  352. arXiv cs.AI TIER_1 English(EN) · Roberto I. Ono Filho ·

    Interrupting the Loop: Periodic Subject Changes Raise Judged Surprise and Connection in Base Language Models

    arXiv:2608.19893v1 Announce Type: cross Abstract: Where does the novelty a base language model produces with no task come from, and what can an LLM judge of a long stream actually see? We dismantle a cognitively inspired generation loop over 24 conditions on three base models. Mo…

  353. arXiv cs.AI TIER_1 English(EN) · Su Yan, Rakesh Iyer ·

    When Machines Speak: A Unified Generative Framework for Integrating Machine-Native Symbols into Pretrained Large Language Models

    arXiv:2608.19529v1 Announce Type: cross Abstract: Many real-world AI systems represent entities, behaviors, and structured information using discrete machine-native symbols rather than natural language. While these representations are compact and preserve task-relevant structure,…

  354. arXiv cs.AI TIER_1 English(EN) · Sokhna Diarra Mbacke, Mouloud Belbahri, Gabriel Loaiza-Ganem ·

    Improved Confidence Estimates for Black-Box Large Language Models

    arXiv:2608.19323v1 Announce Type: cross Abstract: Uncertainty quantification (UQ) is essential for the safe deployment of large language models (LLMs). Existing methods, from verbalized confidence to ones requiring multiple generations, are often zero-shot and produce scores quan…

  355. Hugging Face Daily Papers TIER_1 English(EN) ·

    Interrupting the Loop: Periodic Subject Changes Raise Judged Surprise and Connection in Base Language Models

    Where does the novelty a base language model produces with no task come from, and what can an LLM judge of a long stream actually see? We dismantle a cognitively inspired generation loop over 24 conditions on three base models. Most of its effect lives in one operation: a new sub…

  356. Hugging Face Daily Papers TIER_1 Nederlands(NL) ·

    Inadvertent Context Leakage in Language Models

    For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere presence of such secrets in a model's context window introduces hidden correlations into the model's …

  357. arXiv cs.CL TIER_1 English(EN) · Haoyan Yang, Mario Xerri, Solha Park, Huajian Zhang, Yiyang Feng, Sai Akhil Kogilathota, Jiawei Zhou ·

    Self-Improvement of Large Language Models: A Technical Overview and Future Outlook

    arXiv:2603.25681v2 Announce Type: replace Abstract: As large language models (LLMs) continue to advance, improving them solely through human supervision is becoming increasingly costly and limited in scalability. As models approach human-level capabilities in certain domains, hum…

  358. arXiv cs.CL TIER_1 English(EN) · Chenchen Mao, Hanjing Shi, Haiyan Jia, Emily Wegrzyn, Dominic DiFranzo ·

    When Readability and Source Retention Diverge: An Evaluability Gap in AI Translation

    arXiv:2608.19083v1 Announce Type: cross Abstract: Readable AI output can leave an evaluability gap: even when the source is shown, an overall-quality judgment may not reflect what an output preserves. We investigated how source-text condition and output rendering relate to percei…

  359. arXiv cs.CL TIER_1 English(EN) · Toheeb Ogunade ·

    Structure, Association, and Decision Value: Representation-Based Difficulty Estimation for Adaptive Inference in African-Language NLI

    arXiv:2608.19003v1 Announce Type: new Abstract: We ask whether internal representation statistics can provide useful example-level difficulty signals for adaptive inference in multilingual African NLP, and find that they cannot in this setting. Studying natural language inference…

  360. arXiv cs.CL TIER_1 Français(FR) · Shiyu Miao, Yunlong Mao, Zirui Huang, Liang Yao, Tianshuo Zheng, Yanhui Gu, Fan Liu, Sheng Zhong ·

    Gradient Mirage: Trainable yet Label-Unidentifiable Gradients in Large Language Model Split Learning

    arXiv:2608.18767v1 Announce Type: new Abstract: Gradient matching attacks (GMAs) in LLM split learning (SL) rely on a critical yet underexplored assumption: the gradient exposed at the split interface is a faithful derivative of the client's full-label training objective. This gr…

  361. arXiv cs.CL TIER_1 English(EN) · Maohao Ran, Chendong Ma, Yanting Zhang, Dailing Jiang, Yusen Huang, Meng Gao, Jun Song ·

    Execution-grounded evaluation reveals hidden failures in language-model calculations for environmental science

    arXiv:2608.18726v1 Announce Type: new Abstract: Large language models are increasingly used for quantitative work in the environmental sciences, yet existing evaluations score only final answers, leaving calculation process unobserved. Here we introduce AtmosCoder-Bench, an execu…

  362. arXiv cs.CL TIER_1 English(EN) · Isabella Gidi, Antonio Almud\'evar, Core Francisco Park, Naomi Saphra, Ricard Marxer ·

    Shared Circuits for Shared Grammar: Tracing Subject-Verb Agreement Across Languages

    arXiv:2608.18545v1 Announce Type: new Abstract: Multilingual large language models often generalize across languages, and prior work suggests that their internal mechanisms can overlap cross-lingually. It remains unclear, however, when such sharing emerges and whether it varies w…

  363. arXiv cs.CL TIER_1 English(EN) · Weiwen Xia, Yuxin Cui, E Cao ·

    Efficient INT8 Inference of Small NLP Models on Server CPUs with PyTorch Native Stack

    arXiv:2608.18182v1 Announce Type: new Abstract: Small NLP models, especially BERT-family encoders, remain important in industrial workloads such as classification, ranking, and retrieval even in the era of large language models. On server CPUs, INT8 quantization offers an attract…

  364. arXiv cs.AI TIER_1 English(EN) · Minh Hoang Nguyen, Tung Le, Huy Tien Nguyen ·

    rEDMRec: Distilling Large Language Model Reasoning into an Editable Experience Memory for Recommendation

    arXiv:2608.18952v1 Announce Type: cross Abstract: Large language models can improve recommendation quality by reasoning explicitly over user history and candidate items - for example, extracting a user's preferences or explaining why one item fits better than another - rather tha…

  365. arXiv cs.AI TIER_1 English(EN) · Wei Yu, Suxing Liu, Minjie Yu, Jiahao Wang, Zhijian Zheng, Haocheng Deng, Bing Li ·

    Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models

    arXiv:2608.18884v1 Announce Type: new Abstract: Reinforcement-learning training of reasoning LLMs (e.g., GRPO) is expensive and requires a controllable environment, committing every contribution to a full training pipeline. We present EvoResearcher, a training-free, inference-tim…

  366. arXiv cs.AI TIER_1 Deutsch(DE) · Hadi Hosseini, Samarth Khanna, Xiyuan Wang ·

    Preference Reasoning under Indeterminacy in Large Language Models

    arXiv:2608.18631v1 Announce Type: new Abstract: As large language models evolve into decision-making agents, the ability to reason over preferences becomes fundamental to alignment, coordination, and collective intelligence. Yet, unlike standard benchmarks, real-world preference …

  367. arXiv cs.AI TIER_1 English(EN) · Shrenil Shaun Sharma, Avi Sharma ·

    Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal Abstractions

    arXiv:2608.18409v1 Announce Type: new Abstract: Combinatorial scheduling poses a significant challenge for language models, requiring them to identify feasible solutions within exponentially large search spaces while satisfying complex constraints. This challenge is especially pr…

  368. arXiv cs.AI TIER_1 English(EN) · Andrei Cristian Popescu, Haitz S\'aez de Oc\'ariz Borde, Pietro Li\`o ·

    Looped Language Models Improve Compositional Tool Calling

    arXiv:2608.18171v1 Announce Type: new Abstract: Looped language models have shown promising results on reasoning benchmarks, yet their potential for agentic tool use remains largely unexplored. We study this question in compositional tool-calling settings, where models must coord…

  369. arXiv cs.AI TIER_1 English(EN) · Yisong Chen, Yifan Gao, Sijing Yu, Chuqing Zhao, Yang Lu ·

    Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges

    arXiv:2608.18080v1 Announce Type: new Abstract: We present a review on the applications of large language models (LLMs) in health, e.g., social media analysis, clinical conversational agents, therapy support tools, prompt engineering, multimodal learning, and ethical consideratio…

  370. arXiv cs.LG TIER_1 English(EN) · Ruhai Lin, Yiyang Guo, Rui-Jie Zhu, Hao Ye, Jason K. Eshraghian ·

    Allocating Recurrent Compute in Looped Language Models

    arXiv:2608.18230v1 Announce Type: new Abstract: Looped language models improve reasoning and knowledge manipulation by applying shared computation repeatedly. Existing systems usually repeat an entire layer stack, although a mixer and a dense feed-forward network (FFN) perform di…

  371. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Huy Tien Nguyen ·

    rEDMRec: Distilling Large Language Model Reasoning into an Editable Experience Memory for Recommendation

    Large language models can improve recommendation quality by reasoning explicitly over user history and candidate items - for example, extracting a user's preferences or explaining why one item fits better than another - rather than mapping history directly to a ranked list. This …

  372. arXiv cs.AI TIER_1 English(EN) · Alankar Atreya, Devesh Batra, Yoages Kumar Mantri, Geremy Bantug, Greig A Cowan, Raad Khraishi ·

    When to Review: Spaced Repetition for Continual Pre-Training of Language Models

    arXiv:2608.17530v1 Announce Type: new Abstract: Continual pre-training of large language models must acquire new information without erasing old knowledge. Existing replay methods often choose a global old/new mixture and sample uniformly, ignoring that examples differ in how qui…

  373. arXiv cs.AI TIER_1 English(EN) · Zhixiang wang, Ziliang Hong, Ulas Bagci ·

    A decodability criterion predicts when hidden-state selection beats majority voting in large language models

    arXiv:2608.17124v1 Announce Type: new Abstract: Combining the answers a large language model (LLM) samples for a question into one decision is a test-time information fusion problem, usually solved by majority voting. Voting is unreliable on difficult questions, where the sampled…

  374. arXiv cs.AI TIER_1 English(EN) · Nyamtulla Shaik, Fengjun Li, Bo Luo ·

    Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

    arXiv:2608.17183v1 Announce Type: new Abstract: Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive settings, where safety and bias failures can cause security and societal risks. However, existing AI safety\slash security\slash compl…

  375. arXiv cs.LG TIER_1 English(EN) · Erik Johannes Husom, Maria Emine Nylund, Ophelia Prillard ·

    Wasted large language models: A life cycle thinking approach

    arXiv:2608.17055v1 Announce Type: cross Abstract: Large Language Models (LLMs) are machine learning (ML) models that have an increasingly large carbon footprint through their development and use. Efforts to increase the energy efficiency of these models have not translated into r…

  376. arXiv cs.LG TIER_1 English(EN) · Peizheng Guo, Jianqi Zhang, Xingyu Zhang, Yun Fan, Jiahuan Zhou, Changwen Zheng, Wenwen Qiang ·

    GUPO: Gradient Uncertainty-aware Policy Optimization for Post-Training Large Language Models

    arXiv:2608.17411v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a widely used approach for post-training Large Language Models (LLMs) for reasoning. In GRPO, the group gradients induced by different queries within the same mini-batch are direc…

  377. arXiv cs.CL TIER_1 English(EN) · Peihong Li, Yuanjie Shi, Yan Yan ·

    When to Plan, When to Polish: Noise Level as a Granularity Axis for Diffusion Language Models

    arXiv:2606.21802v2 Announce Type: replace Abstract: Standard tokenwise diffusion LMs keep training corruption and inference commitment at token granularity throughout denoising. At high noise, this leaves scattered local fragments rather than coherent evidence, making it hard to …

  378. arXiv cs.CL TIER_1 English(EN) · Md. Faiyaz Abdullah Sayeedi ·

    Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds

    arXiv:2608.17950v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate remarkable multi-hop reasoning capabilities over long contexts, yet the internal mechanisms enabling these distant cognitive leaps remain poorly understood. Traditional attention-based interp…

  379. arXiv cs.CL TIER_1 English(EN) · Maria Valentini, T\'ea Wright, Julisa Granados, Eliana Colunga, Katharina von der Wense ·

    An Investigation of Translationese in the Generations of Multilingual Large Language Models

    arXiv:2608.17399v1 Announce Type: new Abstract: Text which has been translated from another language tends to carry with it evidence of translation$\unicode{x2014}$hence, it is often referred to as $\textit{translationese}$. Multilingual large language models (MLLMs) generate tex…

  380. arXiv cs.CL TIER_1 English(EN) · Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed ·

    Uncertainty-Aware Decision Making in Multimodal Large Language Models

    arXiv:2608.17084v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, chart, or embodied evidence. Their failures are therefore not only linguistic. A flue…

  381. arXiv cs.AI TIER_1 English(EN) · Deqing Fu, Tianyi Zhou, Mikhail Belkin, Vatsal Sharan, Robin Jia ·

    Convergent Evolution: How Different Language Models Learn Similar Number Representations

    arXiv:2604.20817v2 Announce Type: replace-cross Abstract: Language models trained on natural text learn to represent numbers using periodic features with dominant periods at $T=2, 5, 10$. In this paper, we identify a two-tiered hierarchy of these features: while Transformers, Lin…

  382. arXiv cs.AI TIER_1 English(EN) · Antonio A. Ginart, Naveen Kodali, Jason Lee, Caiming Xiong, Silvio Savarese, John R. Emmons ·

    LZ Penalty: An information-theoretic repetition penalty for autoregressive language models

    arXiv:2504.20131v4 Announce Type: replace-cross Abstract: We introduce the LZ penalty, a penalty specialized for reducing degenerate repetitions in autoregressive language models without loss of capability. The penalty is based on the codelengths in the LZ77 universal lossless co…

  383. arXiv cs.AI TIER_1 English(EN) · Zhikai Ding, Ziyi Ye ·

    Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics

    arXiv:2608.17268v1 Announce Type: cross Abstract: Curriculum learning has been widely adopted in the post-training of large language models by organizing training data from easy to hard. However, its effectiveness varies substantially across reasoning tasks, suggesting that no si…

  384. arXiv cs.AI TIER_1 English(EN) · Md Abdullahil Oaphy, Anhao Xiang, Zongxing Xie, Huayue Gu, Chenyu Wang, Honghui Xu ·

    COMIC: Reference-Aware Safety Gating for Multimodal Large Language Models

    arXiv:2608.17234v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly used to interact with screenshots, scanned documents, diagrams, and other visually grounded inputs. This shift introduces a new safety risk: in many multimodal jailbreaks, …

  385. Hugging Face Daily Papers TIER_1 English(EN) ·

    When to Review: Spaced Repetition for Continual Pre-Training of Language Models

    Continual pre-training of large language models must acquire new information without erasing old knowledge. Existing replay methods often choose a global old/new mixture and sample uniformly, ignoring that examples differ in how quickly they are forgotten. We formulate continual …

  386. arXiv cs.AI TIER_1 English(EN) · Varvara Arzt, Allan Hanbury, Terra Blevins ·

    Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models

    arXiv:2608.15129v1 Announce Type: cross Abstract: We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse natural languages. On artificial languages, models exhibit a left-branching preference that…

  387. arXiv cs.CL TIER_1 English(EN) · Raphael Bernas, Paul G. Chevalier, Fanny Jourdan, C\'eline Hudelot ·

    When Do Concepts Become Functionally Sufficient During Language-Model Training?

    arXiv:2608.15323v1 Announce Type: new Abstract: Understanding a model and its learning mechanisms in depth requires identifying when its internal structures become useful, rather than simply looking at the final state. We study this through concept dynamics: at each layer and che…

  388. arXiv cs.CL TIER_1 English(EN) · Florian Braun ·

    Interpretable Cross-Lingual Alignment in Small Language Models: Probing Cultural and Pragmatic Reasoning in Japanese-English Bilingual LLMs

    arXiv:2608.14896v1 Announce Type: new Abstract: Large language models work well on English and behave in poorly understood ways on languages typologically far from it. Japanese is a clean example, where evaluation still leans on translation quality and JGLUE-style benchmarks, whi…

  389. arXiv cs.CL TIER_1 English(EN) · Animesh Jha, Arpandeep Khatua, Youssef Allouah, Sanmi Koyejo ·

    What to Forget in Unlearning? Forget Set Curation for Language Models

    arXiv:2608.14855v1 Announce Type: new Abstract: Machine unlearning aims to remove targeted data or behaviors from a trained model without retraining from scratch. Yet most evaluations assume that the examples to forget are already known. In realistic language-model deployments, a…

  390. arXiv cs.AI TIER_1 English(EN) · Haeun Yu, Arnav Arora Seogyeong Jeong, Nadav Borenstein, Siddhesh Pawar, Jisu Shin, Jiho Jin, Junho Myung, Alice Oh, Isabelle Augenstein ·

    CulTrace: Tracing Internal Cultural Reasoning in Large Language Models

    arXiv:2508.08879v3 Announce Type: replace-cross Abstract: The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of models' hidden representations of different cultures. Prior work has evaluated cultural awarene…

  391. arXiv cs.AI TIER_1 English(EN) · Zhicheng Lin ·

    A validity-guided workflow for robust large language model research in psychology

    arXiv:2507.04491v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are rapidly being integrated into psychological and behavioral research as research tools, evaluation targets, human simulators, and cognitive models. Yet recent evidence reveals severe measure…

  392. arXiv cs.AI TIER_1 English(EN) · Zhicheng Lin ·

    From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology

    arXiv:2506.16697v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are entering psychological research both as tools and as objects of inquiry. Yet many studies apply human instruments to LLMs without establishing that the outputs are reliable or interpretable…

  393. arXiv cs.AI TIER_1 English(EN) · Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang ·

    DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

    arXiv:2505.14107v5 Announce Type: replace-cross Abstract: The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenar…

  394. arXiv cs.AI TIER_1 English(EN) · Hongye Qiu, Yue Xu, Yi Wang, Meikang Qiu, Wenjie Wang ·

    DR.GAP: Mitigating Bias in Large Language Models using Gender-Aware Prompting with Decoupled Reasoning

    arXiv:2502.11603v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) exhibit strong natural language understanding capabilities but also inherit and amplify societal biases, particularly gender bias, raising fairness concerns. Existing prompt-based debiasing str…

  395. arXiv cs.AI TIER_1 English(EN) · Enric Junque de Fortuny, Veronica Roberta Cappelli ·

    The Fragility of Strategic Thinking in Large Language Models

    arXiv:2510.10813v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly applied to domains that require reasoning about other agents' behavior, such as negotiation, policy design, and market simulation. However, can we trust LLMs to think strategically i…

  396. arXiv cs.AI TIER_1 English(EN) · Jos\'e Luiz Nunes, Guilherme FCF Almeida, Brian Flanagan ·

    Evidence of conceptual mastery in the application of rules by Large Language Models

    arXiv:2503.00992v2 Announce Type: replace Abstract: In this paper we leverage psychological methods to investigate LLMs' conceptual mastery in applying rules. We introduce a novel procedure to match the diversity of thought generated by LLMs to that observed in a human sample. We…

  397. arXiv cs.AI TIER_1 English(EN) · Mohammad Aref Jafari-Raddani, Morteza Mohajjel Kafshdooz ·

    SAPE: Sandwich Adapters for Parameter Efficiency in Large Language Model Fine-Tuning

    arXiv:2608.15360v1 Announce Type: cross Abstract: While Parameter-Efficient Fine-Tuning (PEFT) has substantially reduced the hardware cost of adapting Large Language Models (LLMs) by decreasing the number of trainable parameters, recent studies have sought to further improve PEFT…

  398. arXiv cs.AI TIER_1 English(EN) · Aryuemaan Kumar Chowdhury, Praveen Oosa, Vineesha Reddy ·

    Wiola 13M, a Gated Spiral Attention Architecture for Parameter Efficient Small Language Models

    arXiv:2608.14604v1 Announce Type: cross Abstract: Small language models in the ten to one hundred million parameter range are attractive for on device inference, rapid experimentation, and controlled scientific study, yet most of them reuse the standard transformer block without …

  399. arXiv cs.AI TIER_1 English(EN) · Yusuke Takahashi, Kyle Wild, Asako Uraki ·

    Cost Scales with Change, Not Corpus Size: Incrementally Maintaining an Evolving Semantic Substrate

    arXiv:2608.16621v1 Announce Type: new Abstract: Retrieval-augmented and agentic question-answering systems increasingly re-derive the meaning of a corpus at query time. Put plainly, instead of re-deriving what a corpus means on every question, the work is done once when a documen…

  400. arXiv cs.AI TIER_1 English(EN) · Shen Liu, Zhenguo Xu, Shaopu Wang, Yike Gao, Chunlei Wang ·

    HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents

    arXiv:2608.16447v1 Announce Type: new Abstract: Long-horizon embodied tasks require LLM agents to iteratively decompose high-level goals, revise plans in response to environmental feedback, and ground leaf-level subgoals into valid executable actions. Recursive context-management…

  401. arXiv cs.AI TIER_1 English(EN) · Eric Xie, Wenqian Ye, Aidong Zhang ·

    ALPS: Measuring Valid Creativity in Large Language Models with Mathematical Construction

    arXiv:2608.15979v1 Announce Type: new Abstract: Large language models produce outputs presented as discoveries - new proofs, conjectures, or molecules. Whether such an output that appears creative is truly original and effective is hard to establish: open-ended outputs require su…

  402. arXiv cs.AI TIER_1 English(EN) · Zhiqiang He, Zhi Liu ·

    ReForge: Keeping ABR Algorithms Never Finished with Verified Large Language Model Edits

    arXiv:2608.15138v1 Announce Type: new Abstract: Designing an ABR algorithm for one network scenario takes an engineer months, and large language models now do this work in hours, matching or beating hand-built designs. But either way, the design fits only the world visible at its…

  403. arXiv cs.AI TIER_1 English(EN) · Johannes Gerstmayr, Sebastian Weyrer, Tobias M\"oltner, Peter Manzl, Michael Pieber ·

    Large Language Models and their Awareness of Mechanics and Spatial Geometry

    arXiv:2608.14615v1 Announce Type: new Abstract: Large Language Models (LLMs) perform well on established code-generation and mathematical-reasoning benchmarks, but their capabilities in mechanics and spatial geometry, here denoted as mechanical engineering awareness, has not been…

  404. arXiv cs.LG TIER_1 English(EN) · Amogh Singh ·

    Certifying Compressed Language Models: An Audit and a Statistical Toolkit

    arXiv:2608.15046v1 Announce Type: new Abstract: A fraction of a point of benchmark accuracy is the usual evidence that a compressed model is equivalent to its original. That quantity is least informative when two models are most alike: a net delta is what survives cancellation be…

  405. arXiv cs.CL TIER_1 English(EN) · Adithya Bhaskar, Xi Ye, Danqi Chen ·

    Language Models that Think, Chat Better

    arXiv:2509.20357v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) trains language models to use long chain-of-thought reasoning (CoT) in domains like mathematics and code with rule-based verifiers. However, long CoT learned through RLVR doe…

  406. arXiv cs.CL TIER_1 English(EN) · Zhaoyi Li, Deyang Kong, Yuan Wei, Evan Yang, Ranran Shen, Mahardika Krisna Ihsani, Ming Yang, Wei Zhang, Chuan Hao, Jian Yang, Ran Tao, Bryan Dai, Shikun Zhang, Wei Ye, Ying Wei, Defu Lian ·

    Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

    arXiv:2608.16647v1 Announce Type: new Abstract: On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domai…

  407. arXiv cs.CL TIER_1 English(EN) · Fernando Cardenas Piepereit ·

    Architecture-Dependent Causal Transfer of Activation States Across Large Language Models

    arXiv:2608.16347v1 Announce Type: new Abstract: Direct communication between AI systems relies on natural language as an intermediate layer, incurring encoding/decoding overhead, token cost, and latency. We ask whether internal activation states can instead be transferred causall…

  408. arXiv cs.CL TIER_1 English(EN) · Rinit Jain, Tirthraj Mahajan, Advait Joshi, Raviraj Joshi ·

    L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages

    arXiv:2608.15535v1 Announce Type: new Abstract: We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs). The benchmark comprises 3,471 curriculum-gro…

  409. arXiv cs.CL TIER_1 English(EN) · Nicolas Zucchet, Hyun Dong Lee, Scott Linderman ·

    Language models suffer from a curse of ambiguity

    arXiv:2608.15448v1 Announce Type: new Abstract: Large language models increasingly rely on sampling as a driver of their own improvement, making the fidelity of their learned distributions more critical than ever. Yet, not all distributions are equally easy to learn. In this work…

  410. Hugging Face Daily Papers TIER_1 English(EN) ·

    Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

    On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. …

  411. Hugging Face Daily Papers TIER_1 English(EN) ·

    Cost Scales with Change, Not Corpus Size: Incrementally Maintaining an Evolving Semantic Substrate

    Retrieval-augmented and agentic question-answering systems increasingly re-derive the meaning of a corpus at query time. Put plainly, instead of re-deriving what a corpus means on every question, the work is done once when a document arrives and is thereafter merely consulted -- …

  412. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Asako Uraki ·

    Cost Scales with Change, Not Corpus Size: Incrementally Maintaining an Evolving Semantic Substrate

    Retrieval-augmented and agentic question-answering systems increasingly re-derive the meaning of a corpus at query time. Put plainly, instead of re-deriving what a corpus means on every question, the work is done once when a document arrives and is thereafter merely consulted -- …

  413. Hugging Face Daily Papers TIER_1 English(EN) ·

    HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents

    Long-horizon embodied tasks require LLM agents to iteratively decompose high-level goals, revise plans in response to environmental feedback, and ground leaf-level subgoals into valid executable actions. Recursive context-management methods such as ReCAP improve planning stabilit…

  414. arXiv cs.AI TIER_1 English(EN) · Pranav Kumar Kaliaperumal ·

    From BERT to Frontier Agents: Eight Years of Language-Model Progress, the Collapse of the Capability-Cost Curve, and the Rise of Task-Targeted Models

    arXiv:2608.13675v1 Announce Type: cross Abstract: Between October 2018 and July 2026 AI models progressed from simple systems like BERT to massive agents that solve complex math and write software. The ability to resolve real coding issues improved by nearly six times per year si…

  415. arXiv cs.AI TIER_1 English(EN) · Amritesh Banerjee, Pranil Raichura ·

    Measuring Cross-Task Behavioral Consistency in Language Model Agents

    arXiv:2608.13598v1 Announce Type: new Abstract: Agent evaluation relies almost entirely on outcome metrics such as success rate, which capture whether an agent succeeds but not how consistently it behaves. We argue that behavioral consistency across tasks is a distinct and measur…

  416. arXiv cs.AI TIER_1 English(EN) · Brett Reynolds ·

    Grounding Without Corrective Control: Truth-Tracking Profiles for Large Language Models

    arXiv:2608.14252v1 Announce Type: new Abstract: Recent work suggests that some large language model representations have content or reference. Grounding can secure either without supplying live routes for correction. This paper asks what follows from that gap. An output is answer…

  417. arXiv cs.AI TIER_1 English(EN) · Pengrui Han, Jacob Andreas, Evelina Fedorenko, Andrea Gregor de Varda ·

    Modular Cognitive Architecture Emerges in Large Language Models

    arXiv:2608.13567v1 Announce Type: new Abstract: The human brain exhibits a striking degree of functional specialization, with distinct networks supporting language, formal reasoning, reasoning about other minds, and reasoning about the physical world. Is this modular organization…

  418. arXiv cs.CL TIER_1 English(EN) · Yongmin Kim, Shota Takashiro, Yusuke Iwasawa, Takeshi Kojima, Yutaka Matsuo ·

    Batch-wise Adaptive Pruning: Periodic Neuron Activation-Aware Weight Pruning for Language Reasoning Model

    arXiv:2608.14003v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on complex tasks through extended chain-of-thought generation, but incur substantial computational costs during inference. In production settings, batched inference is essenti…

  419. arXiv cs.AI TIER_1 English(EN) · Akira Okutomi ·

    Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors

    arXiv:2608.13591v1 Announce Type: new Abstract: High-confidence errors in large language models are often treated as evidence of fragile internal inference. We study a different possibility: stable miscalibration, where a confident wrong answer remains locally stable under small …

  420. arXiv cs.AI TIER_1 English(EN) · Xiulin Yang, Arianna Bisazza, Nathan Schneider, Ethan Gotlieb Wilcox ·

    A Unified Assessment of the Poverty of the Stimulus Argument for Neural Language Models

    arXiv:2602.09992v2 Announce Type: replace-cross Abstract: Several recent contributions have evaluated the Poverty of the Stimulus Hypothesis (PoSH) using Artificial Neural Networks (ANNs). The results suggest that ANN-based language models can acquire certain structure-dependent …

  421. arXiv cs.CL TIER_1 English(EN) · Junbo Zhang, Ran Chen, Qianli Zhou, Xinyang Deng, Wen Jiang ·

    Understanding and Mitigating Over-refusal for Large Language Models via Representation Intervention

    arXiv:2511.19009v2 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate powerful capabilities across various natural language processing tasks,yet their inherent safety vulnerabilities undermine the reliable application of LLMs in real-world scenarios. …

  422. arXiv cs.CL TIER_1 English(EN) · Ziyang Luo, Zhongyao Chu, Xinjie He, Youting Wang, Xukui Qin, Runxiong Wu, Yan-Syuan Chen ·

    You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model

    arXiv:2608.14465v1 Announce Type: new Abstract: A frozen language model on reasoning tasks has two coupled weaknesses: it under-uses evidence its own residual stream already encodes, and it fails to detect when the input is insufficient to answer, so it confabulates. This paper c…

  423. arXiv cs.CL TIER_1 English(EN) · Arwa Osman, Marco Baroni, Iuri Macocco ·

    Local and Global Regimes of Geometric Complexity in Language Model Representations

    arXiv:2608.14361v1 Announce Type: new Abstract: Intrinsic dimensionality (ID) is widely used to probe the representational complexity of language models, but it remains unclear whether ID differences reflect properties of language itself or artefacts of how the underlying dataset…

  424. Hugging Face Daily Papers TIER_1 English(EN) ·

    ALPS: Measuring Valid Creativity in Large Language Models with Mathematical Construction

    Large language models produce outputs presented as discoveries - new proofs, conjectures, or molecules. Whether such an output that appears creative is truly original and effective is hard to establish: open-ended outputs require subjective judgment, the output may replicate some…

  425. Hugging Face Daily Papers TIER_1 English(EN) ·

    Looped Language Models Improve Compositional Tool Calling

    Looped language models improve compositional, multi-step tool use through recurrent computation, with adaptive inference balancing accuracy and compute cost.

  426. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Scott Linderman ·

    Language models suffer from a curse of ambiguity

    Large language models increasingly rely on sampling as a driver of their own improvement, making the fidelity of their learned distributions more critical than ever. Yet, not all distributions are equally easy to learn. In this work, we identify a curse of ambiguity: in large lan…

  427. arXiv cs.AI TIER_1 English(EN) · Aoxin Ni ·

    Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement

    arXiv:2608.13129v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong results on mathematical reasoning benchmarks yet remain unreliable on elementary numerical tasks, including magnitude comparison, large-integer arithmetic, fractions, and scientific notati…

  428. arXiv cs.LG TIER_1 English(EN) · Trishita Tiwari, Ari Trachtenberg, G. Edward Suh ·

    A Prior-Aware Metric for Efficiently Distinguishing Memorization from Generalization in Large Language Models

    arXiv:2602.18733v2 Announce Type: replace Abstract: Training data leakage from Large Language Models (LLMs) raises serious concerns related to privacy, security, and copyright compliance. A central challenge in assessing this risk is distinguishing prefix-specific memorization of…

  429. arXiv cs.LG TIER_1 English(EN) · Qiuwu Chen, Zimo Liu, Yuchen Li, Ying Sun, Yifan Zhang, Zhijie Qiu, Zeng You, Ryan Dong, Simeng Ma, Yaofo Chen, Mingkui Tan ·

    LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining

    arXiv:2608.12419v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable breakthroughs across various applications. However, their architectures remain inefficient in pretraining due to two main limitations: (i) self-attention lacks an explicit induct…

  430. arXiv cs.CL TIER_1 English(EN) · Jared Moore, Rasmus Overmark, Ned Cooper, Beba Cibralic, Nick Haber, Cameron R. Jones ·

    Large Language Models Persuade Without Planning Theory of Mind

    arXiv:2602.17045v2 Announce Type: replace Abstract: A growing body of work attempts to evaluate the theory of mind (ToM) abilities of humans and large language models (LLMs) using static, non-interactive question-and-answer benchmarks. However, theoretical work in the field sugge…

  431. arXiv cs.CL TIER_1 English(EN) · Orr Well, Idan Tarshish, Nur Lan, Roni Katzir ·

    Comment on "Modeling rapid language learning by distilling Bayesian priors into artificial neural networks"

    arXiv:2608.12974v1 Announce Type: cross Abstract: McCoy & Griffiths (2025, henceforth M&amp;G) suggest that a Bayesian prior can be distilled into Artificial Neural Networks (ANNs) through Model-Agnostic Meta-Learning (MAML, Finn et al., 2017). They support this empirically by sh…

  432. arXiv cs.CL TIER_1 English(EN) · Yuto Nishida, Hirokazu Kiyomaru, Yusuke Oda, Takashi Kodama, Chaoran Liu, Daisuke Kawahara, Yusuke Miyao, Max M\"uller-Eberstein, Masaru Isonuma ·

    Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining

    arXiv:2608.13515v1 Announce Type: new Abstract: Measuring training data influence consistently across language model pretraining is challenging. It is difficult to select downstream tasks or validation sets representative of a model's general capabilities, and reliance on task pe…

  433. arXiv cs.CL TIER_1 English(EN) · Siheng Xiong, Ali Payani, Oguzhan Gungordu, Faramarz Fekri ·

    DIVE: Unlocking Self-Improvement in Frozen Language Models Through Diversity-Driven Skill Evolution

    arXiv:2608.12486v1 Announce Type: new Abstract: Large language models (LLMs) cannot retain post-deployment experience without parameter updates. We introduce DIVE, a diversity-driven framework that enables frozen LLMs to improve by evolving persistent natural-language skills from…

  434. arXiv cs.AI TIER_1 English(EN) · Erica Coppolillo, Emilio Ferrara ·

    MOSAIC: Unveiling the Moral, Social and Individual Dimensions of Large Language Models

    arXiv:2603.00048v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly deployed in sensitive applications including psychological support, healthcare, and high-stakes decision-making. This expansion has motivated growing research into the ethical …

  435. arXiv cs.AI TIER_1 English(EN) · Abdullah X ·

    Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models

    arXiv:2508.12220v2 Announce Type: replace-cross Abstract: Can a prospectively instrumented training continuation reproduce a deletion counterfactual exactly after selected examples leave its replay dataset? We study a trace-preserving counterfactual that fixes recorded execution …

  436. arXiv cs.AI TIER_1 English(EN) · Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thadd\"aus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel ·

    LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

    arXiv:2608.13545v1 Announce Type: cross Abstract: Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challen…

  437. arXiv cs.AI TIER_1 English(EN) · Mohammed Sabry, Sean Augenstein, Keith Rush, Lucio Dery ·

    Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model

    arXiv:2608.13277v1 Announce Type: cross Abstract: We ask whether language-model pre-training can be decomposed into smaller, independently trainable jobs that can later be recomposed into a coherent larger model. We introduce Mixture of Training (MoT), a scaffolded modular pre-tr…

  438. arXiv cs.AI TIER_1 English(EN) · Paras Balani, Subhrakanta Panda ·

    Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models

    arXiv:2608.13258v1 Announce Type: cross Abstract: Self-referential prompting has been shown to reliably induce large language models to produce first-person reports resembling subjective experience, but no prior work measures how consistent these reports are across repeated, inde…

  439. arXiv cs.AI TIER_1 English(EN) · Fangzhou Liu, Peiyi Han, Jiawei Liu, Yuan Pu, Zhuolun He, Rongliang Fu, Tsung-Yi Ho, Bei Yu ·

    SynAct: A Reasoning-Acting Large Language Model Agent for Adaptive Synthesis Optimization

    arXiv:2608.12751v1 Announce Type: cross Abstract: Logic synthesis transforms RTL designs into gate-level netlists, where PPA results are highly sensitive to the choice of optimization commands, making synthesis tuning both high-dimensional and expensive. Previous approaches fall …

  440. arXiv cs.AI TIER_1 English(EN) · Mehdy Sedaghat Payam, Justin Quinn ·

    Novels generated by language models show compressed formal variation

    arXiv:2608.12630v1 Announce Type: cross Abstract: While large language models can generate entire novels, there is little information about the level of formal variation in their output over many generations. Rather than asking whether individual passages can be identified as AI-…

  441. arXiv cs.AI TIER_1 English(EN) · Morteza Salehjahromi, Shayan A. Zadegan, Amgad Muneer, Jia Wu ·

    From Observation to Intervention: Memory in Brains and Large Language Models

    arXiv:2608.12377v1 Announce Type: cross Abstract: Brains and large language models (LLMs) are fundamentally different memory systems, but they can be compared through shared functional questions: where memory-related information is represented, how partial cues recover broader as…

  442. arXiv cs.AI TIER_1 English(EN) · Eldad Yechiam, Adi Tarabeih ·

    Mimicry without understanding: the origins of decision bias in large language models

    arXiv:2608.12339v1 Announce Type: cross Abstract: Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases. We examined two mechanisms through which such biases can be generated even when human preferences (in the training dat…

  443. arXiv cs.AI TIER_1 English(EN) · Arnav Srivastav ·

    Steering the Language Axis: From Linear Decodability to Causal Control

    arXiv:2608.12334v1 Announce Type: cross Abstract: Despite the impressive multilingual capabilities of Large Language Models, the latent dynamics dictating language selection remain poorly understood. In this work, we ask whether language identity is merely linearly decodable from…

  444. arXiv cs.AI TIER_1 English(EN) · Amogh Joshi, Animesh Mukherjee, Sergey Utyuzhnikov ·

    DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition

    arXiv:2608.13048v1 Announce Type: new Abstract: In this work, we introduce DMDIntel which uses dynamic mode decomposition (DMD) to make the predictions made by LLMs in a classification task interpretable. It develops an input attribution pipeline, that first decomposes the hidden…

  445. arXiv cs.AI TIER_1 English(EN) · Mariya I. Vasileva ·

    Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction

    arXiv:2608.12426v1 Announce Type: new Abstract: Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas. Individual constraints are handled proficient…

  446. Hugging Face Daily Papers TIER_1 English(EN) ·

    Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining

    Measuring training data influence consistently across language model pretraining is challenging. It is difficult to select downstream tasks or validation sets representative of a model's general capabilities, and reliance on task performance at intermediate checkpoints complicate…

  447. Hugging Face Daily Papers TIER_1 English(EN) ·

    Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement

    Large language models (LLMs) achieve strong results on mathematical reasoning benchmarks yet remain unreliable on elementary numerical tasks, including magnitude comparison, large-integer arithmetic, fractions, and scientific notation. This survey examines basic numerical underst…

  448. arXiv cs.AI TIER_1 English(EN) · Mark Shapiro ·

    Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets

    arXiv:2608.11233v1 Announce Type: cross Abstract: A dense, pretrained language model can be retrofitted with recurrent depth and learn an iterative latent transition that persists after outcome-only annealing. Qwen2.5-0.5B-Instruct is split into a Prelude, a weight-tied Recurrent…

  449. arXiv cs.AI TIER_1 English(EN) · Kuangzhao Yang, Ziliang Zhao, Zhicheng Dou ·

    CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement

    arXiv:2608.11631v1 Announce Type: new Abstract: In open-domain human-computer interaction scenarios, large language models (LLMs) frequently encounter user queries that are ambiguous or incomplete. In such cases, directly producing an answer often leads to overgeneralized, errone…

  450. arXiv cs.AI TIER_1 English(EN) · Yueru Yan, Siqi Wu, Thai Le ·

    Locating and Controlling Implicit Personalization in Large Language Models

    arXiv:2608.11735v1 Announce Type: cross Abstract: Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented this behavior, but the connection between these behav…

  451. arXiv cs.AI TIER_1 English(EN) · Yoshihiko Kayama ·

    Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models

    arXiv:2608.11657v1 Announce Type: cross Abstract: We introduce Semantic Lenia, an artificial life framework that transforms Large Language Model (LLM) inference from a static optimization problem into a continuous dynamical system within the macroscopic logit space. By establishi…

  452. arXiv cs.AI TIER_1 English(EN) · Musa Cim, Sarthak Arora, Poovaiah Palangappa, Miro Hodak, Ravi Dwivedula, Meena Arunachalam, Mahmut Taylan Kandemir ·

    Pretraining large language models with MXFP4 on Native FP4 Hardware

    arXiv:2605.09825v4 Announce Type: replace-cross Abstract: Why does full-pipeline FP4 training of large language models often diverge, even when forward activations and activation gradients remain stable? We address this question through a controlled study of MXFP4 quantization in…

  453. arXiv cs.AI TIER_1 English(EN) · Roland M\"uhlenbernd ·

    Social Meaning in Large Language Models: Structure, Magnitude, and Pragmatic Prompting

    arXiv:2604.02512v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly exhibit human-like patterns of pragmatic and social reasoning. This paper addresses two related questions: do LLMs approximate human social meaning not only qualitatively but also …

  454. arXiv cs.AI TIER_1 English(EN) · Antonio Valerio Miceli-Barone, Vaishak Belle, Ali Payani ·

    Program Semantic Inequivalence Game with Large Language Models

    arXiv:2505.03818v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) can achieve strong performance on everyday coding tasks, but they can fail on complex tasks that require non-trivial reasoning about program semantics. Finding training examples to teach LLMs t…

  455. arXiv cs.AI TIER_1 English(EN) · Francesca Da Ros, Luca Di Gaspero, Kevin Roitero ·

    Behavior and Representation in Open-Weight Large Language Models for Combinatorial Optimization: From Feature Extraction to Algorithm Selection

    arXiv:2512.13374v2 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) open new perspectives for automation in optimization, yet little is known about whether their internal representations capture problem structure or algorithmic behavior. We investi…

  456. arXiv cs.AI TIER_1 English(EN) · Minjun Kim, Inho Won, Hyeonseok Lim, MinKyu Kim, Junghun Yuk, Wooyoung Go, Jongyoul Park, Jungyeul Park, KyungTae Lim ·

    TELLME: Test-Enhanced Learning for Language Model Enrichment

    arXiv:2608.11788v1 Announce Type: cross Abstract: Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such as the difficulty of acquiring large-scale domain-…

  457. arXiv cs.AI TIER_1 English(EN) · Zafar Hussain, Kristoffer Nielbo ·

    Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier

    arXiv:2608.11247v1 Announce Type: new Abstract: Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs. Each agent sees what the o…

  458. arXiv cs.AI TIER_1 English(EN) · Mingyu Zong, Sampad Mohanty, Bhaskar Krishnamachari ·

    Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models

    arXiv:2608.11583v1 Announce Type: new Abstract: Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters. This work a…

  459. arXiv cs.LG TIER_1 English(EN) · Duo Wu, Linjia Kang, Zhimin Wang, Fangxin Wang, Wei Zhang, Chongbo Sun, Xuefeng Tao, Wei Yang, Le Zhang, Wenwu Zhu, Peng Cui, Zhi Wang ·

    Grounding Large Language Models as Generalizable Policies in Network Control

    arXiv:2512.11839v2 Announce Type: replace Abstract: Designing generalizable control policies that operate reliably under changing conditions is essential for robust network services in modern digital infrastructure. Yet network control remains dominated by specialized policies bu…

  460. arXiv cs.LG TIER_1 English(EN) · Chencheng Zhu ·

    Orientation, not magnitude: the causal structure of task-vector interference in merged language models

    arXiv:2608.11797v1 Announce Type: new Abstract: Model merging by task arithmetic works until it doesn't, and the field diagnoses why with magnitudes: layerwise representation bias, deviations from cross-task linearity, parameter overlap. Tracking the exact layerwise cross-term of…

  461. arXiv cs.CL TIER_1 English(EN) · Chathuri Jayaweera, Phoebe Huang, Bonnie J. Dorr ·

    BLADE: Better Language Answers through Dialogue and Explanations

    arXiv:2604.03236v2 Announce Type: replace-cross Abstract: Large language model (LLM)-based educational assistants often provide direct answers offering little incentive for students to explore or engage with course materials. We present BLADE (Better Language Answers through Dial…

  462. arXiv cs.CL TIER_1 English(EN) · Zunhai Su, Bohan Sun, Xialie Zhuang, Shuibai Zhang, He Xiao, Jing Xiong, Hengyuan Zhang, Zhongzhu Zhou, Tiantian Zhang, Ngai Wong, Chuan-Wei Kuo ·

    Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

    arXiv:2608.12149v1 Announce Type: new Abstract: We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before full attention layers, forming pre-attenti…

  463. arXiv cs.CL TIER_1 English(EN) · Xiaojun Wu, Cehao Yang, Honghao Liu, Xueyuan Lin, Xuhui Jiang, Chengjin Xu, Jia Li, Jian Guo ·

    LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training

    arXiv:2608.11919v1 Announce Type: new Abstract: Training large language models on limited hardware is increasingly a scheduling problem across GPU compute, host memory, PCIe transfer, and storage bandwidth. Existing offloading systems reduce GPU residency, and MegaTrain shows tha…

  464. arXiv cs.CL TIER_1 English(EN) · Nirmal Thomas ·

    Language-Conditional Dequantization: Recovering What Quantization Steals from Non-English Languages

    arXiv:2608.11786v1 Announce Type: new Abstract: Aggressive quantization disproportionately harms multilingual capability: in the sub-4B INT3 GPTQ regime, we measure 2-4x larger perplexity degradation on non-English languages than on English. We propose Language-Conditional Dequan…

  465. Hugging Face Daily Papers TIER_1 English(EN) ·

    LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

    A curated elementary-grade pretraining corpus and 5B-parameter model create a controlled sandbox for studying knowledge acquisition, representation, and bounded capability growth via post-training and in-context learning.

  466. Hugging Face Daily Papers TIER_1 English(EN) ·

    Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

    We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before full attention layers, forming pre-attention spikes (PAS), and can persist through interve…

  467. Hugging Face Daily Papers TIER_1 English(EN) ·

    LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training

    Training large language models on limited hardware is increasingly a scheduling problem across GPU compute, host memory, PCIe transfer, and storage bandwidth. Existing offloading systems reduce GPU residency, and MegaTrain shows that a CPU-master layer-streaming executor can trai…

  468. Hugging Face Daily Papers TIER_1 English(EN) ·

    Language-Conditional Dequantization: Recovering What Quantization Steals from Non-English Languages

    Aggressive quantization disproportionately harms multilingual capability: in the sub-4B INT3 GPTQ regime, we measure 2-4x larger perplexity degradation on non-English languages than on English. We propose Language-Conditional Dequantization (LCD), a post-hoc method that attaches …

  469. Hugging Face Daily Papers TIER_1 English(EN) ·

    Locating and Controlling Implicit Personalization in Large Language Models

    Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented this behavior, but the connection between these behavioral changes and the model's internal activations…

  470. arXiv cs.AI TIER_1 English(EN) · Chris Han, Pengzhi Gao, Pei Fu, Jian Luan ·

    Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

    arXiv:2608.10812v1 Announce Type: cross Abstract: We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a re…

  471. arXiv cs.AI TIER_1 English(EN) · Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee ·

    Decodable But Not Detachable: Training Data Granularity Determines Parametric Modularity in Large Language Models

    arXiv:2608.10214v1 Announce Type: new Abstract: Do large language models contain domain-specific parametric shells: concentrated, causally necessary neuron populations whose removal selectively degrades a target domain while sparing others? We apply a uniform causal methodology a…

  472. arXiv cs.AI TIER_1 English(EN) · Jordan Pettyjohn, Mansi Sakarvadia, Nathaniel Hudson, Daniel McKenzie, Kyle Chard, Ian Foster ·

    Interpreting Language Model Hidden States at Scale

    arXiv:2608.10260v1 Announce Type: new Abstract: Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions develop through the network. Trained lenses remain expensive: affine-translator p…

  473. arXiv cs.AI TIER_1 English(EN) · Chenhao Dang, Siyuan Xiong, Conghui He, Weijia Li ·

    SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

    arXiv:2608.10538v1 Announce Type: new Abstract: Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually constrain a language model's behavior space for repe…

  474. arXiv cs.AI TIER_1 English(EN) · Minsoo Kim, Sungyoung Ji, Kisung Moon, Ilyong Yoon ·

    Attention-Path Fragility as an Uncertainty Signal in Large Language Models

    arXiv:2608.11138v1 Announce Type: cross Abstract: We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways. We inst…

  475. arXiv cs.AI TIER_1 English(EN) · Md Jafrin Hossain, Mohammad Arif Hossain, Nirwan Ansari ·

    On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models

    arXiv:2608.10530v1 Announce Type: cross Abstract: Large Language Models (LLMs) have undergone a shift from stateless conversational interfaces to autonomous agents capable of multi-step planning, tool invocation, code execution, and maintaining persistent memory. When these agent…

  476. arXiv cs.AI TIER_1 English(EN) · Matthew Siper, Ahmed Khalifa, Julian Togelius ·

    ELMER: Evolutionary Language Model that Explores and Refines

    arXiv:2608.10196v1 Announce Type: cross Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is an unreliable proxy: a small code change can alter nearly every action, while a l…

  477. arXiv cs.CL TIER_1 English(EN) · Laurits Lyngbaek, Ross Deans Kristensen-McLachlan ·

    Multilingual Embedding Probes Fail to Generalize Across Learner Corpora

    arXiv:2604.07095v2 Announce Type: replace Abstract: Do multilingual embedding models encode a language-general representation of proficiency? We investigate this by training linear and non-linear probes on hidden-state activations from seven embedding models (0.3-8B) to predict C…

  478. arXiv cs.CL TIER_1 English(EN) · Dong Qiao, Chris Ding, Jicong Fan ·

    Mapping and Measuring the Behavioral Evolution of Large Language Models

    arXiv:2608.11027v1 Announce Type: cross Abstract: Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across generations. We characterize the output behavior of 32 models from six families using …

  479. arXiv cs.CL TIER_1 English(EN) · Nicol\'as Vera Z\'u\~niga ·

    What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model

    arXiv:2608.10986v1 Announce Type: new Abstract: A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops. We ask what such a probe measures, in a construction chosen to make the question sharp: a ring of…

  480. arXiv cs.AI TIER_1 English(EN) · Hector Zenil ·

    On Solomonoff Induction in Large Language Models and the Limits of Self-Improving: The Singularity Is Not Near Without Symbolic Model Synthesis

    arXiv:2601.05280v3 Announce Type: replace-cross Abstract: On the one hand, the question of whether large language models (LLMs) are Solomonoff induction estimators has become an explicit question at the intersection of Algorithmic Information Theory (AIT) and Machine Learning (ML…

  481. arXiv cs.AI TIER_1 English(EN) · Zhaolu Kang, Junhao Gong, Jiaxu Yan, Wanke Xia, Yian Wang, Ziwen Wang, Huaxuan Ding, Zhuo Cheng, Wenhao Cao, Zhiyuan Feng, Siqi He, Shannan Yan, Junzhe Chen, Xiaomin He, Chaoya Jiang, Wei Ye, Kaidong Yu, Xuelong Li ·

    HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models

    arXiv:2506.03922v4 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchmarks for evaluating MLLMs primarily emphasize general knowledge and vertical step…

  482. Hugging Face Daily Papers TIER_1 English(EN) ·

    Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

    Massive activations in hybrid linear-attention LLMs exhibit pre-attention spikes and inter-spike plateaus governed by cancellation timing, with morphology recovering at full-attention limits.

  483. Hugging Face Daily Papers TIER_1 English(EN) ·

    SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

    Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually constrain a language model's behavior space for repeatable, high-quality task execution. However, be…

  484. arXiv cs.AI TIER_1 English(EN) · Yiqun Zhang, Yunfan Zhang, Mingjie Zhao, Sen Feng, Yiu-ming Cheung ·

    How to Ask the AI: A User Perspective Survey for Large Language Model Prompting

    arXiv:2608.07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by typing requests, such as ``plan a three-day Vienna trip'', ``solve the attached m…

  485. arXiv cs.CL TIER_1 English(EN) · Jaydeep Borkar, Karan Chadha, Niloofar Mireshghallah, Yuchen Zhang, Irina-Elena Veliche, Archi Mitra, David A. Smith, Zheng Xu, Diego Garcia-Olano ·

    Memorization Dynamics in Knowledge Distillation for Language Models

    arXiv:2601.15394v2 Announce Type: replace Abstract: Knowledge Distillation (KD) is increasingly adopted to transfer capabilities from large language models to smaller ones, offering significant improvements in efficiency and utility while often surpassing standard fine-tuning. Be…

  486. arXiv cs.CL TIER_1 English(EN) · Kristian Schwethelm, Daniel Rueckert, Georgios Kaissis ·

    Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching

    arXiv:2608.09444v1 Announce Type: cross Abstract: A main promise of looped language models (LMs) is depth-adaptive inference. By iterating a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. However, this…

  487. arXiv cs.CL TIER_1 English(EN) · Bocheng Chen, Han Zi, Roucheng Ou, Yawei Liu, Minyue Chen, Zimo Qi, Rongrong Wang, Guangliang Liu ·

    Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models

    arXiv:2608.09551v1 Announce Type: new Abstract: In the era of large language models (LLMs), attackers often manipulate natural language to elicit unsafe or harmful outputs, creating a new natural language attack surface unique to LLM-based systems, where attacks directly exploit …

  488. arXiv cs.CL TIER_1 English(EN) · Jiguo Li ·

    The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism

    arXiv:2608.08650v1 Announce Type: new Abstract: Mixture-of-Experts models increase parameter capacity while keeping the computation activated by each token bounded, but their architectural evolution cannot be explained by a chronological list of model releases alone. This technic…

  489. arXiv cs.CL TIER_1 English(EN) · Catherine M. Brousse, Nelu D. Radpour ·

    Focus particles and scalar inferences across humans and language models

    arXiv:2608.08227v1 Announce Type: new Abstract: Focus particles such as "even" and "only" are central to formal semantic theories that posit structured representations over sets of alternatives. "Even" highlights unexpected or extreme alternatives, while "only" enforces exclusivi…

  490. arXiv cs.CL TIER_1 English(EN) · Bo Cheng, Qiaolin Lu, Yi Chang, Yuan Wu ·

    Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders

    arXiv:2608.08168v1 Announce Type: new Abstract: While Large Language Models (LLMs) employing Chain-of-Thought (CoT) exhibit superior reasoning capabilities, the neural mechanisms distinguishing this explicit Thinking mode from direct answer generation (NoThinking mode) remain poo…

  491. arXiv cs.AI TIER_1 English(EN) · Akshay Paruchuri, Ishan Chatterjee, Henry Fuchs, Ehsan Adeli, Piotr Didyk ·

    The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models

    arXiv:2604.14363v2 Announce Type: replace-cross Abstract: Multimodal language models systematically underperform on visual perception tasks, yet the structure underlying this failure remains poorly understood. We propose centroid replacement, mapping tokens to their nearest K-mea…

  492. arXiv cs.AI TIER_1 English(EN) · Hua-Dong Xiong, Li Ji-An, Jiaqi Huang, Robert C. Wilson, Kwonjoon Lee, Xue-Xin Wei ·

    In-context superposition: human-like working memory interference in large language models

    arXiv:2604.09670v2 Announce Type: replace-cross Abstract: Intelligent systems must maintain and manipulate task-relevant information online to adapt to dynamic environments and changing goals. This capacity, known as working memory, is fundamental to human reasoning and intellige…

  493. arXiv cs.AI TIER_1 English(EN) · Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin ·

    Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages

    arXiv:2603.12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge and sequential tasks while mitigating catastrophic forgetting, a critical limitati…

  494. arXiv cs.AI TIER_1 English(EN) · Sourav Chattaraj, Kanak Raj ·

    LLMs Remember First, Forget Last: Dual-Process Interference in Large Language Models

    arXiv:2603.00270v3 Announce Type: replace-cross Abstract: Large language models can process millions of tokens, yet how they handle conflicting information within context remains poorly understood. From patient health logs tracking evolving vital signs to legal documents with sup…

  495. arXiv cs.AI TIER_1 English(EN) · Bj\"orn Deiseroth, Max Henning H\"oth, Kristian Kersting, Letitia Parcalabescu ·

    Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models

    arXiv:2512.11614v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) relies on retrieved context to guide large language models (LLM), yet treats the retrieval as a heuristic rather than verifiable evidence -- leading to unsupported answers, hallucinatio…

  496. arXiv cs.AI TIER_1 Dansk(DA) · Dong Dong, Weijie Su ·

    Length-MAX Tokenizer for Language Models

    arXiv:2511.20849v2 Announce Type: replace-cross Abstract: We introduce a new tokenizer for language models that minimizes the average tokens per character, thereby reducing the number of tokens needed to represent text during training and to generate text during inference. Our me…

  497. arXiv cs.AI TIER_1 English(EN) · Santiago Cuervo, Adel Moumen, Yanis Labrak, Sameer Khurana, Antoine Laurent, Mickael Rouvier, Phil Woodland, Ricard Marxer ·

    LF${}^{2}$AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models

    arXiv:2503.06211v3 Announce Type: replace-cross Abstract: Text-pretrained language models (LMs) encode rich world knowledge, but adapting them to process and generate perceptual modalities such as audio and images while effectively leveraging that knowledge remains challenging. P…

  498. arXiv cs.AI TIER_1 English(EN) · Han Wang, Yifan Sun, Brian Ko, Mann Talati, Jiawen Gong, Zimeng Li, Naicheng Yu, Xucheng Yu, Wei Shen, Vedant Jolly, Huan Zhang ·

    MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models

    arXiv:2603.28590v3 Announce Type: replace Abstract: Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs. When such a mismatch occurs, the CoT no longer faithfully reflects the actual reasons (i.e., de…

  499. arXiv cs.AI TIER_1 English(EN) · Laurens Samson, Iva Gornishka, Gossa L\^o, Yuki M. Asano, Sennay Ghebreab ·

    From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch

    arXiv:2608.09925v1 Announce Type: cross Abstract: Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We…

  500. arXiv cs.AI TIER_1 English(EN) · Congfeng Cao, Pengyu Zhang, Jelke Bloem ·

    Fusion Training for Mathematical Generalization in Large Language Models

    arXiv:2608.09893v1 Announce Type: cross Abstract: Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by unifying a non-thinking mode and a thinking mode within a single model. However, its training dynamics, includin…

  501. arXiv cs.AI TIER_1 English(EN) · Yilin Jiang, Xiaorong Zhu, Fei Tan, Zicheng Zhang, Kaiyi Huang, Yang Yu, Zexuan Fei, Yiming Luo, Keqian Li, Hao Hao, Aimin Zhou, Guangtao Zhai ·

    ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

    arXiv:2608.09548v1 Announce Type: cross Abstract: Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed …

  502. arXiv cs.AI TIER_1 English(EN) · Yuting Liu, Wei Wu, Jianzhe Zhao, Guibing Guo ·

    Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning

    arXiv:2608.09507v1 Announce Type: cross Abstract: Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a particular downstream task. Directly supplying the ful…

  503. arXiv cs.AI TIER_1 English(EN) · Muhammad Faishal Adly Nelwan, Alfan Farizki Wicaksono ·

    Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models

    arXiv:2608.08829v1 Announce Type: cross Abstract: Activation steering edits the behaviour of a frozen language model by adding a learned vector to its residual stream, and current practice fixes the injection layers globally per task. We argue that the best layers are an instance…

  504. arXiv cs.AI TIER_1 English(EN) · Joseph Bingham ·

    LegoLM: Structured Weight Sharing for Large Language Models

    arXiv:2608.08652v1 Announce Type: cross Abstract: We present \LegoLM{}, a structured weight-sharing compression framework for large language models grounded in a systematic study of why global weight sharing fails and how to fix it. We identify two distinct failure modes. Distrib…

  505. arXiv cs.AI TIER_1 English(EN) · Guide Labs Team, Andreas Madsen, Aya Abdelsalam Ismail, Giang Nguyen, Isaac Plant, Muawiz Chaudhary, Nathaniel Monson, Saqib Azim, Zhichen Guo, Julius Adebayo ·

    Scaling Inherently Interpretable Language Models

    arXiv:2608.07594v1 Announce Type: cross Abstract: Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish. In this work, we challenge this prem…

  506. arXiv cs.AI TIER_1 English(EN) · Shuai Lyu, Yuning Gong, Ruiling Gao, Xiaoran Shang, Zhonghong Ou, Ping Zong, Yifan Zhu, Yuan Sun, Yang Qin, Peng Hu ·

    Multi-Branch Policy Optimization for Multimodal Large Language Models

    arXiv:2608.07581v1 Announce Type: cross Abstract: Group-based reinforcement learning methods for multimodal large language models typically rely on trajectory-level credit assignment that applies a single advantage to all tokens in a response. However, multimodal reasoning involv…

  507. arXiv cs.AI TIER_1 English(EN) · Xi Li, Shu Zhao, Xiaohan Zou, Fei Zhao, Fuxiao Liu, Yusen Zhang, Cheng Han, Yushun Dong, Jiaqi Wang ·

    Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards

    arXiv:2608.07535v1 Announce Type: cross Abstract: Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding and reasoning. However, this architectural shift reshapes the safety landscape of …

  508. arXiv cs.AI TIER_1 English(EN) · Su-Hyeon Kim, Jiwan Mun, Yo-Sub Han ·

    One Adapter Pair per Model: A Universal Activation Interface for Language Models

    arXiv:2608.09521v1 Announce Type: new Abstract: Activation-based tools are usually tied to one model's native hidden space, requiring probes, sparse autoencoders, and natural-language interpreters to be rebuilt or rediscovered for each new language model. We present a Universal A…

  509. arXiv cs.AI TIER_1 English(EN) · Xuanchen Li, Haitao Li, Yujia Zhou, Qingyi Pan, Heng Wang, Yiqun Liu, Min Zhang, Qingyao Ai ·

    Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models

    arXiv:2608.09109v1 Announce Type: new Abstract: User feedback offers natural supervision for persistent LLM improvement, but a single message may support multiple behavioral changes with different scopes of generalization. We introduce SLIFT, a selective self-learning framework b…

  510. arXiv cs.AI TIER_1 English(EN) · Abdalla Doleh, Toni Somers, Ratna Babu Chinnam ·

    Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models

    arXiv:2608.08822v1 Announce Type: new Abstract: Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and biased. We developed an automated pipeline that uses LLms to generate structured …

  511. arXiv cs.AI TIER_1 English(EN) · Kaili Zheng, Kaiwen Wang, Xun Zhu, Qingyuan Yang, Chenyi Guo, Ji Wu ·

    FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models

    arXiv:2608.08736v1 Announce Type: new Abstract: Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities of Multimodal Large Language Models (MLLMs) in this setting remain underexplored. Existing benchmarks rely on action-specific…

  512. arXiv cs.AI TIER_1 English(EN) · Chenxi Zhou, Pengfei Cao, Jinyu Ye, Bohan Yu, Haida Yu, Jiang Li, Jun Zhao, Kang Liu ·

    Quantization Degradation in Large Language Models: A Signal-Noise Perspective

    arXiv:2608.08188v1 Announce Type: new Abstract: Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determined by bit-width alone. We systematically study weight-only post-training quantization across…

  513. arXiv cs.AI TIER_1 English(EN) · Sahil Pardasani, Madhusudan Singh ·

    Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

    arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims that DeepSeek R1 outperformed OpenAI's o1 contributed to market panic on January 27…

  514. Hugging Face Daily Papers TIER_1 English(EN) ·

    SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

    SKILLER is a reinforcement learning framework that automatically generates tailored skills for small open-source models to reduce inference costs while maintaining high task performance.

  515. Hugging Face Daily Papers TIER_1 English(EN) ·

    Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

    Open multilingual translation models are improved via group relative policy optimization with reference-free quality rewards and checkpoint interpolation, surpassing strong open and proprietary baselines.

  516. Hugging Face Daily Papers TIER_1 English(EN) ·

    Interpreting Language Model Hidden States at Scale

    Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions develop through the network. Trained lenses remain expensive: affine-translator parameters grow quadratically with model width, w…

  517. Hugging Face Daily Papers TIER_1 English(EN) ·

    ELMER: Evolutionary Language Model that Explores and Refines

    Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is an unreliable proxy: a small code change can alter nearly every action, while a larger rewrite can preserve the same execution trac…

  518. arXiv cs.AI TIER_1 English(EN) · Linkai Peng, Baorian Nuchged ·

    Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models

    arXiv:2608.06409v1 Announce Type: cross Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a generation…

  519. arXiv cs.AI TIER_1 Deutsch(DE) · Ali Jalal-Kamali ·

    Divergent Response Modes in Frontier Language Models Under Steering Pressure

    arXiv:2608.06578v1 Announce Type: new Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably different behaviors under explicit steering pressure remains underexplored. This study evaluate…

  520. arXiv cs.AI TIER_1 English(EN) · Zibo Shao, Baochen Xiong, Chengdong Xu, Linhui Xiao, Kaichen Li, Haoran Gong, Yan Li, Yaguang Song, Xiaoshan Yang ·

    AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models

    arXiv:2608.06699v1 Announce Type: new Abstract: Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are specialized for particular tools or environments, c…

  521. arXiv cs.AI TIER_1 English(EN) · Rens Anderson, Tessa Verhoef, Amirhossein Zohrehvand ·

    Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models

    arXiv:2608.07243v1 Announce Type: new Abstract: Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement. This pilot study examines whether iterative search improves LLM cr…

  522. arXiv cs.AI TIER_1 English(EN) · Zhenlong Liu, Hao Zeng, Weiran Huang, Hongxin Wei ·

    Provable Training Data Identification for Large Language Models

    arXiv:2510.09717v3 Announce Type: replace-cross Abstract: Identifying training data of large-scale models is critical for copyright litigation, privacy auditing, and ensuring fair evaluation. However, existing works typically treat this task as an instance-wise identification wit…

  523. arXiv cs.AI TIER_1 English(EN) · Saeid Shokoufa, Mohammad Erfan Sadeghi, Mehdi Kamal, Massoud Pedram ·

    Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors

    arXiv:2608.06723v1 Announce Type: cross Abstract: The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency, making accurate estimation essential for sustainable artificial intelligence deployment a…

  524. arXiv cs.CL TIER_1 English(EN) · Zili Zhang, Yilin Wang, Heng Wang, Herun Wan, Minnan Luo ·

    Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models

    arXiv:2608.07261v1 Announce Type: new Abstract: Large language models (LLMs) can solve complex multi-hop problems yet exhibit puzzling failures on simple two-hop queries: although a model may correctly store each individual hop, it often fails to combine them. To understand the i…

  525. arXiv cs.CL TIER_1 English(EN) · Mudar Adas, Polina Tsvilodub, Michael Franke, Martin V. Butz ·

    Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models

    arXiv:2608.06977v1 Announce Type: new Abstract: It is well established that large language models (LLMs) are sensitive to prompt framing, reflecting patterns in their training data or prior prompts. In this study, we investigate the extent to which LLMs reinforce users biases exp…

  526. Hugging Face Daily Papers TIER_1 English(EN) ·

    Quantization Degradation in Large Language Models: A Signal-Noise Perspective

    Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determined by bit-width alone. We systematically study weight-only post-training quantization across bit-widths, quantization methods, model scales …

  527. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Amirhossein Zohrehvand ·

    Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models

    Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement. This pilot study examines whether iterative search improves LLM creativity by adapting FunSearch to recipe generat…

  528. arXiv cs.AI TIER_1 English(EN) · Maryam Rahimi, Mohammad Reza Daliri, Yadollah Yaghoobzadeh ·

    Explanations of Large Language Models Explain Language Representations in the Brain

    arXiv:2502.14671v4 Announce Type: replace-cross Abstract: Large Language Model (LLM) representations are known to align with brain activity during language processing, but it remains unclear what drives this alignment. We test whether explainable AI (XAI) can help answer this: us…

  529. arXiv cs.LG TIER_1 English(EN) · Hoda Fakharzadehjahromy, Emil Wiman, Andreas Bueff, Hafsteinn Einarsson, Fredrik Heintz ·

    SAGA: Score-Weighted Adaptive Generation Alignment for Low-Resource Nordic Language Models

    arXiv:2608.06179v1 Announce Type: new Abstract: Preference optimisation has proven effective for improving large language models but typically relies on costly human preference annotations. Extending these methods to morphologically rich, low-resource languages remains challengin…

  530. arXiv cs.LG TIER_1 English(EN) · Zhaowei Han, Xiang Zhang, Bing Han, Kai Liu, Danqi Hu, Jie Liu ·

    KV-Skill: Forging Expertise in the Model's Native Language

    arXiv:2608.05475v1 Announce Type: new Abstract: Task knowledge is commonly stored either as text in the prompt or as an update to model weights. Text is modular but must be interpreted on every use, while weight adaptation makes the resulting capability difficult to load, remove,…

  531. arXiv cs.CL TIER_1 English(EN) · Ahmet Yavuz Uluslu, Tannon Kew, Tilia Ellendorff, Gerold Schneider, Rico Sennrich ·

    Robust Native Language Identification through Agentic Decomposition

    arXiv:2509.16666v2 Announce Type: replace Abstract: Large language models (LLMs) often achieve high performance in native language identification (NLI) benchmarks by leveraging superficial contextual clues such as names, locations, and cultural stereotypes, rather than the underl…

  532. arXiv cs.CL TIER_1 English(EN) · Aarohi Srivastava, David Chiang ·

    Different Perturbations, Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness

    arXiv:2608.05510v1 Announce Type: new Abstract: Dialectal variation remains a major challenge for multilingual language models. Perturbation-based continued pre-training (CPT) has emerged as a promising approach to improving robustness, yet existing work largely evaluates individ…

  533. arXiv cs.AI TIER_1 English(EN) · Maryam Rahimi, Mahdi Nouri, Yadollah Yaghoobzadeh ·

    Layer-wise Positional Bias in Short-Context Language Modeling

    arXiv:2601.04098v2 Announce Type: replace-cross Abstract: Transformer language models systematically prefer tokens at specific input positions regardless of semantic relevance---a phenomenon known as positional bias. Prior work characterizes this bias in model behavior through pe…

  534. arXiv cs.AI TIER_1 English(EN) · Uri Katz, Omer Goldman, Tomasz Limisiewicz, Reut Tsarfaty, Noah A. Smith ·

    MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

    arXiv:2608.05850v1 Announce Type: cross Abstract: We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources hav…

  535. arXiv cs.AI TIER_1 English(EN) · Chang Shi, Tim Pearce, Manan Tomar, Siddhartha Sen, John Langford ·

    Hierarchical Latent Prediction for Language Models

    arXiv:2608.05806v1 Announce Type: cross Abstract: While standard Next-Token Prediction (NTP) lays the foundation of language model pre- training, its teacher-forced training paradigm may not be optimal for long-horizon reasoning and planning. Recent works such as Multi-Token Pred…

  536. arXiv cs.AI TIER_1 English(EN) · Gihoon Kim, Jeyoung Lee, Suhan Woo, Sekwon Oh, Minsu Jeon, Hyounsoo Han, Euntai Kim ·

    Cautious Context Steering for Language Model Personalization

    arXiv:2608.05813v1 Announce Type: new Abstract: Personalizing language models (LMs) to individual user preferences is essential for aligning responses with diverse goals and backgrounds. Existing methods typically train a separate adapter for each user or learn a reward model who…

  537. arXiv cs.AI TIER_1 English(EN) · Zonghao Yang ·

    Counterfactual Analysis via Large Language Models

    arXiv:2608.05367v1 Announce Type: new Abstract: Counterfactual analysis aims to predict potential outcomes under hypothetical scenarios, offering valuable insights for decision-making. This paper investigates the application of large language models (LLMs), specifically the GPT-3…

  538. Hugging Face Daily Papers TIER_1 English(EN) ·

    SAGA: Score-Weighted Adaptive Generation Alignment for Low-Resource Nordic Language Models

    Preference optimisation has proven effective for improving large language models but typically relies on costly human preference annotations. Extending these methods to morphologically rich, low-resource languages remains challenging because such annotations are scarce. We presen…

  539. Hugging Face Daily Papers TIER_1 English(EN) ·

    MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

    We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modelin…

  540. arXiv cs.CL TIER_1 Deutsch(DE) · Amanda Popadich, Shane Steinert-Threlkeld ·

    Language Models Generalize to Human-like Word Order Preferences

    arXiv:2608.05028v1 Announce Type: new Abstract: A central question in language acquisition is whether linguistic biases can emerge from general learning mechanisms operating over underdetermined input. Artificial Language Learning (ALL) studies have shown that human learners reli…

  541. arXiv cs.AI TIER_1 English(EN) · Sihan Hu, Lyuhan Huang, Youjin Deng, Kun Chen ·

    SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models

    arXiv:2608.04975v1 Announce Type: cross Abstract: SciCode is the standard measure of the scientific-coding ability of language models: research-level problems that demand both frontier scientific theory and its implementation as working numerical code. It is a component of the Ar…

  542. arXiv cs.AI TIER_1 English(EN) · Yexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang, Chengpeng Fu, Yu Wang, Ming Liu ·

    Breaking the Curse ofMultilinguality inMany-to-Many Speech-to-Text Translation via a Resource-AwareMixture of Speech Encoders

    arXiv:2608.04586v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing multilingual speech inputs, a single speech encoder shared across all languages suffers from…

  543. arXiv cs.CL TIER_1 English(EN) · Mullosharaf K. Arabov ·

    Large Language Models for Low-Resource Languages: A Conceptual Framework for an Electronic Explanatory Dictionary of the Tajik Language

    arXiv:2608.04186v1 Announce Type: new Abstract: This paper presents a conceptual framework for developing an electronic explanatory dictionary of the Tajik language using large language models (LLMs). The relevance of the work stems from the absence of a comprehensive digital lex…

  544. arXiv cs.AI TIER_1 English(EN) · Jianru Shen ·

    Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models

    arXiv:2608.05064v1 Announce Type: cross Abstract: Small open-weight language models increasingly run in private, offline, and cost-sensitive settings, where the key deployment question is not only what a model answers but when it should defer to a human. We study whether verbaliz…

  545. arXiv cs.CL TIER_1 English(EN) · Junbo Zhang, Qianli Zhou, Xinyang Deng, Wen Jiang ·

    DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning

    arXiv:2608.04322v1 Announce Type: new Abstract: Task-specific fine-tuning can improve the performance of large language models (LLMs) on downstream tasks. However, our study reveals that task-specific fine-tuning can also weaken the safety guardrails of aligned LLMs. A widely ado…

  546. arXiv cs.CL TIER_1 English(EN) · Luxshan Thavarasa, Sivasuthan Sukumar ·

    Test, then Route: How Language Models Execute In-Context Conditional Rules Across Models and Languages

    arXiv:2608.04183v1 Announce Type: new Abstract: When a language model follows an in-context conditional rule such as "if P(x) then A else B," does it assemble a runtime circuit with one module that tests the predicate and another that routes the answer? We probe this with activat…

  547. arXiv cs.CL TIER_1 English(EN) · Yuezhang Peng, Yuxin Liu, Changfeng Gao, Zhifu Gao, Xiangang Li, Xie Chen ·

    Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models

    arXiv:2608.05126v1 Announce Type: new Abstract: Spoken Language Understanding (SLU) is the core component of task-oriented dialogue systems and a pivotal link in achieving seamless human-agent interaction. While traditional SLU can effectively extract user semantics for closed-se…

  548. arXiv cs.CL TIER_1 English(EN) · Eric Zhang, Leshem Choshen, Jacob Andreas ·

    Unforgettable Generalization in Language Models

    arXiv:2409.02228v2 Announce Type: replace-cross Abstract: When language models (LMs) are trained to forget (or "unlearn'') a skill, how precisely does their behavior change? We study the behavior of transformer LMs in which tasks have been forgotten via fine-tuning on randomized …

  549. arXiv cs.CL TIER_1 English(EN) · Soroush Omidvartehrani, Mohammadamin Habibollah, Mohammadreza Daviran, Davood Rafiei ·

    EdgeLM: Edge Demonstrations for Language Models' Table Understanding

    arXiv:2608.04390v1 Announce Type: new Abstract: Large language models (LLMs) perform table-centric prediction through in-context learning, making demonstration selection critical to performance. Existing retrieval methods prioritize similarity to the query, but similar demonstrat…

  550. arXiv cs.AI TIER_1 English(EN) · Ruitong Li, Binjie Guo, Aisheng Mo, Guowei Su, Jie Li, Ru Zhang ·

    Agreement Before Diversity: Verification-First Complementarity for Heterogeneous Language-Model Coordination

    arXiv:2608.04618v1 Announce Type: new Abstract: Heterogeneous language-model ensembles expand the space of candidate responses, yet they lack a principled criterion for when a newly generated answer should supersede an already supported one. We decouple candidate headroom from re…

  551. arXiv cs.AI TIER_1 English(EN) · Matteo Silvestri, Fabiano Veglianti, Flavio Giorgi, Fabrizio Silvestri, Gabriele Tolomei ·

    When Large Language Models Know the Table: A Framework for Assessing Data Contamination in Tabular Datasets

    arXiv:2510.20351v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly exposed to data contamination, i.e., performance gains driven by prior exposure of test datasets rather than generalization. However, in the context of tabular data, this probl…

  552. Hugging Face Daily Papers TIER_1 English(EN) ·

    Different Perturbations, Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness

    Dialectal variation remains a major challenge for multilingual language models. Perturbation-based continued pre-training (CPT) has emerged as a promising approach to improving robustness, yet existing work largely evaluates individual perturbation strategies in isolation and pro…

  553. Hugging Face Daily Papers TIER_1 English(EN) ·

    Scaling Inherently Interpretable Language Models

    Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish. In this work, we challenge this premise. Rather than reverse-engineering a model, we m…

  554. Hugging Face Daily Papers TIER_1 English(EN) ·

    MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

    We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modelin…

  555. arXiv cs.CL TIER_1 English(EN) · Monnie McGee, Mateo Langston Smith, Julian Cabrera ·

    Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

    arXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how models construct and communicate statistical explanations. This study demonstrate…

  556. arXiv cs.AI TIER_1 English(EN) · Maryam Rezaee, Pooriya Safaei, Maryam Asgarinezhad, Fatemeh Seyyedsalehi ·

    Interpreting Black-Box Large Language Models with Sentence-Level Energy Landscapes

    arXiv:2608.02879v1 Announce Type: new Abstract: The widespread adoption of proprietary Large Language Models (LLMs) accessed strictly through closed APIs has created a critical challenge for responsible deployment: a fundamental lack of interpretability. To address this, we propo…

  557. arXiv cs.AI TIER_1 English(EN) · Mingtian Zhang ·

    MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc

    arXiv:2608.03397v1 Announce Type: new Abstract: MMLongBench-Doc is a long-document QA benchmark of 1,082 questions over 135 PDFs. Two properties of it push measured scores away from the quantity they are meant to capture: the reference metric compares extracted answers, so 1,358,…

  558. arXiv cs.AI TIER_1 English(EN) · Sefika Efeoglu, Adrian Paschke ·

    Reversing Arrows in Large Language Models

    arXiv:2608.03512v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance on text-to-knowledge graph generation and related tasks. Nevertheless, it is still unclear whether they accurately model the direction-dependent semantics of inverse rela…

  559. arXiv cs.AI TIER_1 English(EN) · Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski ·

    GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models

    arXiv:2608.03729v1 Announce Type: cross Abstract: Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowledge source. However, LLMs nati…

  560. arXiv cs.AI TIER_1 English(EN) · Hailong Jiang, Feng Yu, Emran Hossain, Jianfeng Zhu, Mengfei Ren, Qiang Guan, Chunwei Xia ·

    Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss?

    arXiv:2608.03983v1 Announce Type: cross Abstract: Optimizing compilers miss profitable transformations when their enabling semantics are absent from the analyzed program representation. We ask whether large language models (LLMs) can recover such semantics from heterogeneous C/C+…

  561. arXiv cs.AI TIER_1 English(EN) · Sarah Ball, Niki Hasrati, Alexander Robey, Avi Schwarzschild, Frauke Kreuter, Zico Kolter, Andrej Risteski ·

    Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models

    arXiv:2510.22014v2 Announce Type: replace-cross Abstract: Discrete optimization-based jailbreaking attacks on large language models aim to generate short, nonsensical suffixes that, when appended onto input prompts, elicit disallowed content. Notably, these suffixes are often tra…

  562. arXiv cs.AI TIER_1 English(EN) · Minh Vu Pham, Hsuvas Borkakoty, Yufang Hou ·

    Where Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models

    arXiv:2601.09445v2 Announce Type: replace-cross Abstract: In language models (LMs), intra-memory knowledge conflict arises when inconsistent information about the same subject is encoded within the model's parametric knowledge. Prior work has primarily focused on resolving confli…

  563. arXiv cs.AI TIER_1 English(EN) · Qingyan Yang, Tongxi Wang, Yunsheng Luo ·

    ChiEngMixBench: Evaluating Large Language Models on Expert-Style Chinese-English Terminology Mixing

    arXiv:2601.16217v2 Announce Type: replace-cross Abstract: Large language models increasingly mediate multilingual professional communication, where useful generation requires adapting to community conventions about which expressions are retained, translated, or mixed. Existing be…

  564. arXiv cs.AI TIER_1 English(EN) · Safal Shrestha, Anubhav Shrestha, Minwu Kim, Aadim Nepal, Keith Ross ·

    On the Limits of Layer Pruning for Generative Reasoning in Large Language Models

    arXiv:2602.01997v3 Announce Type: replace-cross Abstract: Recent work has shown that layer pruning can effectively compress large language models (LLMs) while retaining strong performance on classification benchmarks, often with little or no finetuning. In contrast, generative re…

  565. arXiv cs.CL TIER_1 English(EN) · \c{S}uayp Talha Kocabay, Talha R\"uzgar Akku\c{s}, Kamer Ali Yuksel ·

    ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads

    arXiv:2608.02703v1 Announce Type: new Abstract: Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but practical backends often retain the final language-modeling head (LM-head) in BF16 or FP16. Quantizing this projection …

  566. arXiv cs.CL TIER_1 English(EN) · Rupak Sarkar, Pritika Ramu, Rachel Rudinger ·

    Language Models Encode the Contextual Truth of Propositions

    arXiv:2608.03035v1 Announce Type: new Abstract: Prior work has shown that LLMs encode the truth of factual propositions along linear directions in activation space. It's unclear how these representations extend to contextual truth: propositions whose truth is determined by in-con…

  567. arXiv cs.CL TIER_1 English(EN) · Hai Wang, Chenhao Wang, Qifeng Cai, Yixiu Liu, Miao Peng, Nuo Chen, Yuanlin Tu, Chengcheng Xu, Feng Zhang ·

    Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining

    arXiv:2608.03089v1 Announce Type: new Abstract: Large-scale pretraining corpora contain substantial duplicate content. Although document-level deduplication is widely used, removing subdocument-level redundancy remains challenging. At corpus scale, suffix-array-based methods are …

  568. arXiv cs.CL TIER_1 English(EN) · Yixin Bu, Runze Xia, Guanyun Zou, Yupeng Ji, Haodong Liu, Piji Li ·

    DUD: Decoupled Update Dynamics for Reliable Uncertainty Quantification in Large Language Models

    arXiv:2608.03411v1 Announce Type: new Abstract: Accurate Uncertainty Quantification (UQ) is critical for reliable deployment of Large Language Models (LLMs), yet traditional probability-based metrics often fail to capture the model's true epistemic state. While recent mechanistic…

  569. arXiv cs.CL TIER_1 (TL) · Mykola Haltiuk ·

    Disentangling Language Modeling and Boundaries

    arXiv:2608.03599v1 Announce Type: new Abstract: Byte-level language models are usually argued for on the grounds of robustness, multilingual fairness, and character-level skills. We point to a different, structural advantage: because they read and write bytes, any two of them sha…

  570. arXiv cs.CL TIER_1 English(EN) · Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y ·

    M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models

    arXiv:2608.03803v1 Announce Type: new Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with p…

  571. arXiv cs.CL TIER_1 English(EN) · Ayoub Hammal, Pierre Zweigenbaum, Caio Corro ·

    Suffix-Constrained Greedy Search Algorithms for Causal Language Models

    arXiv:2603.01243v2 Announce Type: replace Abstract: Large language models (LLMs) are powerful tools that have found applications beyond human-machine interfaces and chatbots. Beside free-form generation, there has been an interest in constrained generation, a setting where LLMs a…

  572. arXiv cs.CL TIER_1 English(EN) · Zhipeng Yang, Shu Yang, Lijie Hu, Di Wang ·

    Word Recovery in Large Language Models Enables Character-Level Tokenization Robustness

    arXiv:2603.10771v2 Announce Type: replace Abstract: Large language models (LLMs) trained with canonical tokenization exhibit surprising robustness to non-canonical inputs such as character-level tokenization, yet the mechanisms underlying this robustness remain unclear. We study …

  573. arXiv cs.LG TIER_1 English(EN) · Lele Zheng, Weifeng Kong, Xinyi Zhang, Ke Cheng, Tao Zhang, Yulong Shen ·

    Noise-Aware Shrinkage for Differentially Private Zeroth-Order Fine-Tuning of Large Language Models

    arXiv:2608.03277v1 Announce Type: new Abstract: Differentially private zeroth-order optimization (DP-ZO) enables memory-efficient private fine-tuning of large language models using only forward evaluations. Existing aggregation-based DP-ZO methods reconstruct model updates at a f…

  574. Hugging Face Daily Papers TIER_1 English(EN) ·

    EdgeLM: Edge Demonstrations for Language Models' Table Understanding

    Large language models (LLMs) perform table-centric prediction through in-context learning, making demonstration selection critical to performance. Existing retrieval methods prioritize similarity to the query, but similar demonstrations often reinforce the model's likely predicti…

  575. Hugging Face Daily Papers TIER_1 English(EN) ·

    Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining

    Large-scale pretraining corpora contain substantial duplicate content. Although document-level deduplication is widely used, removing subdocument-level redundancy remains challenging. At corpus scale, suffix-array-based methods are commonly applied independently within shards, le…

  576. arXiv cs.CL TIER_1 English(EN) · Qiyuan Liu, Hao Xu, Xuhong Chen, Wei Chen, Yee Whye Teh, Ning Miao ·

    Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey

    arXiv:2510.01925v3 Announce Type: replace Abstract: Reward models (RMs) play a critical role in enhancing the reasoning performance of LLMs. For example, they can provide training signals to finetune LLMs during reinforcement learning (RL) and help select the best answer from mul…

  577. arXiv cs.CL TIER_1 English(EN) · Sing Hieng Wong, Hassan Sajjad, A. B. Siddique ·

    LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering

    arXiv:2604.03532v2 Announce Type: replace Abstract: Large language models (LLMs) show strong multilingual capabilities, yet reliably controlling the language of their outputs remains difficult. Representation-level steering addresses this by adding language-specific vectors to mo…

  578. arXiv cs.CL TIER_1 English(EN) · Vinayshekhar Bannihatti Kumar, Disha Makhija, Manoj Ghuhan Arivazhagan, Rashmi Gangadharaiah ·

    Syntax Without Semantics: Teaching Large Language Models to Code in an Unseen Language

    arXiv:2605.15607v2 Announce Type: replace Abstract: Large language models (LLMs) achieve high pass rates on code generation benchmarks, yet whether they can transfer this ability to languages absent from pretraining remains poorly understood. We introduce PyLang, a minimal impera…

  579. arXiv cs.CL TIER_1 English(EN) · Weixin Liang ·

    Computational Approaches to Understanding Large Language Model Impact on Writing and Information Ecosystems

    arXiv:2506.17467v2 Announce Type: replace Abstract: Large language models (LLMs) have shown significant potential to change how we write, communicate, and create, leading to rapid adoption across society. This dissertation examines how individuals and institutions are adapting to…

  580. arXiv cs.CL TIER_1 English(EN) · Wenjun Wang, Shuo Cai, Congkai Xie, Mingfa Feng, Yiming Zhang, Zhen Li, Kejing Yang, Ming Li, Jiannong Cao, Hongxia Yang ·

    A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models

    arXiv:2509.22536v5 Announce Type: replace Abstract: The immense computational cost of training Large Language Models (LLMs) presents a major barrier to innovation. While FP8 training offers a promising solution with significant theoretical efficiency gains, its widespread adoptio…

  581. arXiv cs.CL TIER_1 English(EN) · Xiang Xia, Cheng Yan, Yiming Zhang, Jiazheng Liu, Hongyu Zhang, Wuyang Zhang ·

    REFLEX: Rethinking MoE Inference as Refinement-Aware Compute Allocation in Diffusion Language Models

    arXiv:2608.01784v1 Announce Type: cross Abstract: Mixture-of-experts (MoE) models increase parameter capacity by activating only a small subset of experts for each token. This conditional-computation paradigm has enabled autoregressive language models to scale model capacity with…

  582. arXiv cs.CL TIER_1 English(EN) · Ken Tsui ·

    Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

    arXiv:2507.02778v3 Announce Type: replace Abstract: Although large language models (LLMs) have transformed AI, they still make errors and follow unproductive reasoning paths. Self-correction is vital for safety-critical applications, but studying it requires disentangling activat…

  583. arXiv cs.LG TIER_1 English(EN) · Xiaoyun Liu, Divya Saxena, Jiannong Cao, Yuqing Zhao, Yiying Dong, Penghui Ruan ·

    GPrune-LLM: Generalization-Aware Structured Pruning for Large Language Models

    arXiv:2603.13418v2 Announce Type: replace Abstract: Structured pruning is widely applied to compress large language models (LLMs), but its performance depends heavily on how neuron importance is estimated. Most existing methods rely on activation statistics from a single calibrat…

  584. arXiv cs.LG TIER_1 English(EN) · Anusha Madan Gopal, Aras Pirbadian, Kristofor D. Carlson, M Anthony Lewis, Jonathan Tapson ·

    Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State Injection

    arXiv:2608.02560v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) imposes a prefill cost proportional to retrieved context length, and -- with Transformer backbones -- a KV-cache that grows with each generated token. State-Space Models (SSMs) avoid the second c…

  585. arXiv cs.LG TIER_1 English(EN) · Marta Garnelo, Wojciech M. Czarnecki ·

    Why Large Language Models Fail at Tabular Prediction

    arXiv:2608.02412v1 Announce Type: new Abstract: Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the most common machine learning workloads: predictive analytics over tabular data. …

  586. arXiv cs.LG TIER_1 English(EN) · Andres Algaba, Francesca Carlon, Lynn Delcon, Marthe Ballon, Bert Verbruggen, Vincent Ginis ·

    How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models

    arXiv:2608.02089v1 Announce Type: new Abstract: Large language models often show users a final response and a short reasoning summary while the full reasoning trace stays hidden. We introduce an observability ladder that holds each completed run fixed and varies only what a reade…

  587. arXiv cs.LG TIER_1 English(EN) · Yuqicheng Zhu, Jialin Yu, Lin Li, Gengyuan Zhang, Zhen Yang, Steffen Staab, Puneet Dokania, Philip Torr, Jie Tang, Evgeny Kharlamov ·

    Conformalized Large Language Models under Configuration Shift

    arXiv:2608.01460v1 Announce Type: new Abstract: Conformal prediction (CP) is a distribution-free framework for uncertainty quantification that has recently been adapted to large language models (LLMs), providing prediction sets with finite-sample coverage guarantees under exchang…

  588. arXiv cs.CL TIER_1 English(EN) · Ramesh B. Paramkusham ·

    Trustworthiness Costs of Domain Adaptation in Small Language Models:A Cross-Architecture Empirical Study

    arXiv:2608.00042v1 Announce Type: new Abstract: Domain adaptation of small language models (SLMs) has emerged as a practical strategy for deploying capable NLP systems in resource-constrained, high-stakes environments including healthcare, legal services, and financial analysis. …

  589. arXiv cs.CL TIER_1 English(EN) · Mohamed El Idrissi ·

    Rethinking and formalising the state across languages: a unified computational learning theory account

    arXiv:2608.00523v1 Announce Type: new Abstract: The linguistic notion of state has traditionally been restricted to the construct (annexation) state of Afroasiatic languages and treated as a language-specific morphosyntactic phenomenon. This article argues instead that the state …

  590. arXiv cs.CL TIER_1 English(EN) · Eojin Jeon, SangKeun Lee ·

    PICTURE: Enhancing Theory-of-Mind in Large Language Models by Revealing, Not Hiding, Characters' Lack of Knowledge

    arXiv:2608.01598v1 Announce Type: new Abstract: Simulating human-like Theory of Mind (ToM) has been a longstanding problem in natural language processing (NLP). To address this, existing works introduce a reasoning step of event hiding (a.k.a. perspective-taking), where events un…

  591. arXiv cs.CL TIER_1 English(EN) · Taeyeong Kim, Ahhyun Kim, TaeHyeon Kim, Unggi Lee ·

    Not the Dimension, the Norm: What Matters in Gradient-Free Weight Perturbation of Language Models

    arXiv:2608.01624v1 Announce Type: new Abstract: Adapting a language model to a task no longer requires training all of its weights, and a line of parameter-efficient methods has driven the trainable count from billions down to a handful of scalars. Gradient-free adaptation, which…

  592. arXiv cs.CL TIER_1 English(EN) · K. Jack Scott, Narun Pat, Veronica Liesaputra ·

    Divergent large language model predictions from convergent representations in ambiguous word pairs

    arXiv:2608.01816v1 Announce Type: new Abstract: In this work we investigate how decoder-only transformers resolve lexical ambiguity through layer-by-layer analysis of three models spanning three parameter sizes (GPT-2-Small-117M, Llama-3.2-3B, Qwen2.5-32B). For both homonyms and …

  593. arXiv cs.CL TIER_1 English(EN) · Shijie Ren, Xiting Wang, Meng Li, Yujie Guo, Yunhang Yao, Ziheng Peng, Xunlong Wang, Yuetan Chen, Haoyang Zhou, Yunlong Liang, Fandong Meng ·

    Self-Improving Large Language Models via Progressive Experience Evolution

    arXiv:2608.02139v1 Announce Type: new Abstract: Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transforming transient interaction experience into persistent model capabilities. Existing …

  594. arXiv cs.CL TIER_1 English(EN) · Victor Maricato ·

    Leak It: A Probabilistic Approach to Training-Data Extraction from Black-Box Language Models

    arXiv:2608.00144v1 Announce Type: cross Abstract: Membership inference (MIA) on language models is usually summarised by an aggregate ROC-AUC, but such evaluations are confounded: model-free blind baselines separate members from non-members from surface text alone. We study black…

  595. arXiv cs.CL TIER_1 English(EN) · Muhammad Faizan Raza (Luna), Shuo (Luna), Yang, Satish Mahadevan Srinivasan, Joanna F. DeFranco ·

    Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments

    arXiv:2608.00419v1 Announce Type: cross Abstract: Large language models deployed in real-time, regulated settings face knowledge staleness, catastrophic forgetting, hallucination, and weak feedback loops. We present a unified, pattern-driven LLMOps architecture integrating real-t…

  596. arXiv cs.CL TIER_1 English(EN) · Ziyi Cai, Shuangping Li, Yiheng Shen, Kangning Wang, Peng Zhang ·

    Dense Language Generation Made Simple: Deterministic, Randomized, and Multi-Order Algorithms

    arXiv:2608.01320v1 Announce Type: cross Abstract: Language generation in the limit is a theoretical framework for studying how a generator can learn to produce new valid strings from a stream of positive examples. In this model, an adversary chooses an unknown language from a cou…

  597. Hugging Face Daily Papers TIER_1 English(EN) ·

    OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

    Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio token sequences incurs substantial computational overhead, demanding aggressive token compression for e…

  598. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Jonathan Tapson ·

    Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State Injection

    Retrieval-augmented generation (RAG) imposes a prefill cost proportional to retrieved context length, and -- with Transformer backbones -- a KV-cache that grows with each generated token. State-Space Models (SSMs) avoid the second cost by construction; we eliminate the first, col…

  599. arXiv cs.CL TIER_1 English(EN) · Jin-woo Lee, Minkyung Song, Junghyun Oh, Seunghoon Han, Soyoung Park, Gwangseon Jang, Sungsu Lim ·

    Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models

    arXiv:2607.28979v1 Announce Type: new Abstract: Heterogeneous Large Language Model (LLM) systems increasingly rely on shared contexts, retrieved evidence, and multi-agent dialogue histories, yet their internal key-value (KV) caches remain model-specific and cannot be reused acros…

  600. arXiv cs.AI TIER_1 English(EN) · Plawan Kumar Rath ·

    The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models

    arXiv:2607.28639v1 Announce Type: cross Abstract: We show that knowledge distillation in small instruction-tuned language models has asymmetric effects on bias. On unambiguous tasks (BBQ-disambig), response-based distillation from a Gemma-2-9B teacher improves context-following: …

  601. arXiv cs.AI TIER_1 English(EN) · Namkyung Yoon, Sanghong Kim, Hwangnam Kim ·

    PARALLEL: A Prefrontal-Aligned Reinforcement inspired Approach for Language-Model Learning under Explicit Limits

    arXiv:2607.28982v1 Announce Type: cross Abstract: Recent language models achieve strong performance across a variety of tasks, but conventional adaptation applies updates uniformly across training samples regardless of their local update benefit. We propose PARALLEL, a prefrontal…

  602. arXiv cs.CL TIER_1 English(EN) · Fabio De Ponte, Eloise Gaines-White, Conor Houghton, Seth Bullock ·

    Evolving language compositionality in a frequency-structured meaning space

    arXiv:2607.29642v1 Announce Type: new Abstract: The iterated learning model was introduced to investigate language evolution: the way in which the characteristic properties of human languages have been shaped, at least partly, by repeated transmission from one language user to an…

  603. arXiv cs.CL TIER_1 English(EN) · Haodong Lei, Junming Liu, Yirong Chen, Pinlong Cai, Botian Shi, Ding Wang, Hongsong Wang ·

    TransMem: Transforming Hidden States into Memory for Large Language Models

    arXiv:2607.29032v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly operate over long interaction histories, where effective reasoning requires identifying and exploiting task-relevant evidence distributed across past observations and actions. However…

  604. Hugging Face Daily Papers TIER_1 English(EN) ·

    ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads

    Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but practical backends often retain the final language-modeling head (LM-head) in BF16 or FP16. Quantizing this projection naively can strongly perturb the vocabulary-logi…

  605. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Joanna F. DeFranco ·

    Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments

    Large language models deployed in real-time, regulated settings face knowledge staleness, catastrophic forgetting, hallucination, and weak feedback loops. We present a unified, pattern-driven LLMOps architecture integrating real-time data ingestion, continual learning, retrieval-…

  606. Hugging Face Daily Papers TIER_1 English(EN) ·

    MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models

    MMOOC is a large-scale benchmark assessing whether multimodal language models can correctly refuse out-of-context questions while answering shifted in-context questions, revealing that current models struggle to balance these abilities.

  607. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Hongsong Wang ·

    TransMem: Transforming Hidden States into Memory for Large Language Models

    Large language model (LLM) agents increasingly operate over long interaction histories, where effective reasoning requires identifying and exploiting task-relevant evidence distributed across past observations and actions. However, useful information encoded in previously compute…

  608. arXiv cs.CL TIER_1 English(EN) · Tao Wen, Shuai Shao, Pei Ke, Xu Han, Jie Zou, Guannan Li, Tao Tian, Jinjie Qiu, Lan Wang, Ke Qin ·

    From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

    arXiv:2607.27654v1 Announce Type: new Abstract: Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promisin…

  609. arXiv cs.LG TIER_1 English(EN) · Damian Hodel, Jevin D. West ·

    Epistemic diversity across language models mitigates knowledge collapse

    arXiv:2512.15011v3 Announce Type: replace Abstract: Artificial intelligence (AI) increasingly generates the very content used to train future AI systems. This feedback loop can degrade model quality, reduce informational diversity, and ultimately drive knowledge collapse, i.e. a …

  610. arXiv cs.LG TIER_1 English(EN) · Sangjin Kim, Yuseon Choi, Jungjun Oh, Byeongcheol Kim, Hoi-Jun Yoo ·

    LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference

    arXiv:2607.27704v1 Announce Type: cross Abstract: As large language models (LLMs) continue to demonstrate exceptional capabilities across various domains, the challenge of achieving energy-efficient and accurate inference becomes increasingly critical. This work presents LightRot…

  611. arXiv cs.LG TIER_1 English(EN) · Lei Dong ·

    The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models

    arXiv:2607.27281v1 Announce Type: new Abstract: A capability appears in a language model when the last parts of its circuit align in one stochastic attempt, and getting all but one right is worth nothing. We show this no-partial-credit joint alignment is the rate-limiting step of…

  612. arXiv cs.CL TIER_1 English(EN) · Seonglae Cho, Zekun Wu, Kleyton Da Costa, Adriano Koshiyama ·

    The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models

    arXiv:2602.08159v2 Announce Type: replace-cross Abstract: When a language model asserts that "the capital of Australia is Sydney," does it know this is wrong? Models assert misconceptions with the same fluency as facts, so the question cannot be answered from output uncertainty. …

  613. arXiv cs.CL TIER_1 English(EN) · Parishad BehnamGhader, Vaibhav Adlakha, Fabian David Schmidt, Nicolas Chapados, Marius Mosbach, Siva Reddy ·

    LLM2Vec-Gen: Generative Embeddings from Large Language Models

    arXiv:2603.10913v3 Announce Type: replace Abstract: Fine-tuning LLM-based text embedders via contrastive learning maps inputs and outputs into a new representational space, discarding the LLM's output semantics. We propose LLM2Vec-Gen, a self-supervised alternative that instead p…

  614. arXiv cs.CL TIER_1 English(EN) · Sina Heydari, Amirreza Abbasi, Mohsen Hooshmand, Majid Ramezani ·

    CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising

    arXiv:2607.28236v1 Announce Type: cross Abstract: Pre-trained language models have significantly improved sentence representation learning, yet their embedding remain sensitive to semantic preserving textual perturbations such as synonym substitution, masking and word dropout. Th…

  615. arXiv cs.CL TIER_1 English(EN) · Yuetian Mao, Chunyang Chen ·

    IFHierBench: Hierarchical Instruction Following for Large Language Models

    arXiv:2607.27912v1 Announce Type: cross Abstract: Instruction-following ability is critical for deploying large language models in real-world applications, where downstream components depend on the output satisfying specific constraints. Modern deployments increasingly handle the…

  616. arXiv cs.CL TIER_1 English(EN) · Vishwajith Ramesh ·

    Subtract or Replay? Exact Deletion from Language-Model Memory

    arXiv:2607.27539v1 Announce Type: cross Abstract: Exact deletion from persistent language-model memory depends on how that memory represents a record. Addressable influence can be removed by algebraic decrement; influence transformed by later writes inside shared recurrent state …

  617. arXiv cs.CL TIER_1 English(EN) · Yanshi Li, Xueru Bai, Shuman Liu, Long Zhang ·

    RepBench: Compiling Benchmarks into Capability Representations for Large Language Models

    arXiv:2607.28008v1 Announce Type: new Abstract: Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthetic data. The resulting measurements are difficult to compare or reproduce and m…

  618. arXiv cs.AI TIER_1 English(EN) · Magda Dubois, Cozmin Ududec, Christopher Summerfield, Lennart Luettgau ·

    Ask don't tell: Reducing sycophancy in large language models

    arXiv:2602.23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts. While pr…

  619. arXiv cs.CL TIER_1 English(EN) · Shuang Liang, Haoyang Zhou, Yifan Gong, Guowei Wang, Xiting Wang ·

    LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models

    arXiv:2607.28077v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rollout rewards consume generation budget without effective learning signals. Pre-r…

  620. arXiv cs.CL TIER_1 English(EN) · Chia-Ming Lee, Ming-Ching Chang, Xin Li, Yu-Lun Liu, Chih-Chung Hsu ·

    Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models

    arXiv:2607.28166v1 Announce Type: new Abstract: Diffusion language models (DLMs) expose a provisional prediction at every denoising step, creating an opportunity for generation-time early exit that stops decoding before the schedule is exhausted. Existing early-exit gates decide …

  621. arXiv cs.CL TIER_1 English(EN) · Hanzuo Liu, Xuan Qi, Chunyu Liu, Haotian Zhong, Yulong Wang, Rayying, Key, Alex Lamb, Mingyu Gao ·

    Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory

    arXiv:2607.28263v1 Announce Type: new Abstract: Transformer depth is not used uniformly: lower and middle layers build semantic representations, while upper layers increasingly specialize them for prediction. We turn this division of labor into CoMem (Comprehension Memory), which…

  622. arXiv cs.CL TIER_1 English(EN) · Bertil Braun, Martin Forell ·

    (Towards) Scalable Reliable Automated Evaluation with Large Language Models

    arXiv:2607.28282v1 Announce Type: new Abstract: Evaluating the quality and relevance of textual outputs from Large Language Models (LLMs) remains challenging and resource-intensive. Existing automated metrics often fail to capture the complexity and variability inherent in LLM-ge…

  623. arXiv cs.CL TIER_1 English(EN) · Albert Gong, Kyuseong Choi, Abhineet Agarwal, Jason Schechner, Ryan Huang, Raj Agrawal, Anish Agarwal, Raaz Dwivedi ·

    ORCA-bench: How Ready Are Language Model Agents for Oncall?

    arXiv:2607.28545v1 Announce Type: new Abstract: Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, oft…

  624. arXiv cs.CL TIER_1 English(EN) · Germans Savcisens, Samantha Dies, Courtney Maynard, Tina Eliassi-Rad ·

    Belief Coevolution in a Social Network of Generalist and Specialist Large Language Models

    arXiv:2607.27512v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in multi-agent environments. However, the processes by which beliefs form and propagate among interacting LLMs remain poorly understood. We introduce CoevolveSim, a framework fo…

  625. arXiv cs.CL TIER_1 English(EN) · Shreyas Subramanian, Mecit Gungor, Vikram Elango ·

    Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models

    arXiv:2607.27506v1 Announce Type: new Abstract: Language and embedding models used in RAG systems are conventionally assumed to require large-scale pretraining and explicit grounding supervision. We present B1ade, an efficient RAG architecture comprising two purpose-built compone…

  626. Hugging Face Daily Papers TIER_1 English(EN) ·

    PhiZero: A World Model Built Around Physical Language

    We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within…

  627. Hugging Face Daily Papers TIER_1 English(EN) ·

    LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models

    Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rollout rewards consume generation budget without effective learning signals. Pre-rollout prompt selection can reduce this waste by…

  628. Hugging Face Daily Papers TIER_1 English(EN) ·

    IFHierBench: Hierarchical Instruction Following for Large Language Models

    Instruction-following ability is critical for deploying large language models in real-world applications, where downstream components depend on the output satisfying specific constraints. Modern deployments increasingly handle the full task in a single LLM call, with one prompt s…

  629. Hugging Face Daily Papers TIER_1 English(EN) ·

    From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

    Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promising performance in part of these tasks individuall…

  630. arXiv cs.LG TIER_1 English(EN) · Ashish Prajapati, Om Mohite ·

    Two Calls Beat Five Agents: Evaluating Multi-Agent Pipelines Against Self-Refinement for Local Language Models

    arXiv:2607.26922v1 Announce Type: new Abstract: Multi-agent LLM pipeline systems break down the task among multiple roles for better reasoning, but are benchmarked mainly with large-scale commercial models. In this study, we investigate Parishad, a structured multi-agent system i…

  631. arXiv cs.LG TIER_1 English(EN) · Zhiwei Hao, Jianyuan Guo, Li Shen, Yong Luo, Han Hu, Guoxia Wang, Dianhai Yu, Yonggang Wen, Dacheng Tao ·

    Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities

    arXiv:2505.01043v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved impressive performance across various domains. However, the substantial hardware resources required for their training present a significant barrier to efficiency and scalability. To mi…

  632. arXiv cs.CL TIER_1 English(EN) · Chandra Sripada, Richard Lewis ·

    Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition

    arXiv:2607.26179v1 Announce Type: cross Abstract: LLMs are widely regarded as alien intelligences, systems whose cognitive operations are fundamentally unlike our own. Apparent similarities to human cognition are therefore often seen as the result of anthropomorphic projection. W…

  633. arXiv cs.CL TIER_1 English(EN) · Rapha\"el Sourty, Antoine Chaffin, Paulo Roberto Moura Junior, Am\'elie Chatelain ·

    DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

    arXiv:2607.27178v1 Announce Type: new Abstract: State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multiling…

  634. arXiv cs.CL TIER_1 English(EN) · Seonglae Cho, Adriano Koshiyama ·

    OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

    arXiv:2607.26981v1 Announce Type: new Abstract: Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a systematic directional tilt has been hard to detect: calibration metrics aggregate un…

  635. arXiv cs.CL TIER_1 English(EN) · Poulami Ghosh, Preethi Jyothi ·

    Language Models are not Equally Robust to Non-Canonical Tokenization across Languages

    arXiv:2607.26831v1 Announce Type: new Abstract: Despite the existence of exponentially many valid tokenizations for a given string, language models operate on a single canonical sequence deterministically produced by the tokenizer, leaving the broader tokenization space largely u…

  636. arXiv cs.CL TIER_1 English(EN) · Chen Shani ·

    From Found to Designed: Concepts as a Design Axis for Large Language Models

    arXiv:2607.26825v1 Announce Type: new Abstract: Large language models (LLMs) encode rich concept-like information, but represent it implicitly through distributed statistical associations rather than as explicit, structured, compositional concepts. Consequently, concept-level str…

  637. arXiv cs.CL TIER_1 English(EN) · Ruxi Gu, Zhenliang Zhang, Wei Wang ·

    ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models

    arXiv:2607.26455v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong capabilities in knowledge acquisition and reasoning, yet their ability to retain previously acquired knowledge under repeated updates remains insufficiently understood. Existing …

  638. arXiv cs.CL TIER_1 English(EN) · Dachuan Song, Junyu Yin, Zechen Hu, Xuan Wang ·

    Mergeable Model-Side Aggregation States for Long-Context Language Models

    arXiv:2607.26448v1 Announce Type: new Abstract: A known limitation of long-context language models is their increasingly unreliable performance in non-additive, set-based aggregation as context length grows. Examples include cardinality estimation, set relationships, and grouped …

  639. Hugging Face Daily Papers TIER_1 English(EN) ·

    PhiZero: A World Model Built Around Physical Language

    We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within…

  640. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Tina Eliassi-Rad ·

    Belief Coevolution in a Social Network of Generalist and Specialist Large Language Models

    Large language models (LLMs) are increasingly deployed in multi-agent environments. However, the processes by which beliefs form and propagate among interacting LLMs remain poorly understood. We introduce CoevolveSim, a framework for studying belief diffusion within networked LLM…

  641. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Amélie Chatelain ·

    DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

    State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval through translate-train. We first …

  642. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Amélie Chatelain ·

    DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

    State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval through translate-train. We first …

  643. Hugging Face Daily Papers TIER_1 English(EN) ·

    DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

    State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval through translate-train. We first …

  644. Hugging Face Daily Papers TIER_1 English(EN) ·

    OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

    Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a systematic directional tilt has been hard to detect: calibration metrics aggregate unsigned errors, and naturalistic uncertainty offe…

  645. Hugging Face Daily Papers TIER_1 English(EN) ·

    ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models

    Large language models (LLMs) have demonstrated strong capabilities in knowledge acquisition and reasoning, yet their ability to retain previously acquired knowledge under repeated updates remains insufficiently understood. Existing evaluation paradigms primarily focus on single-s…

  646. arXiv cs.CL TIER_1 English(EN) · Abhishek Kumar Singh, Shrey Nag, Sachita, Lipi Goel, Rajeshwar Singh Janwar ·

    Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context

    arXiv:2607.25375v1 Announce Type: new Abstract: India is a vast nation of over 1.4 billion people, varied by hundreds of diverse and locally specific traditions and cultures and 22 officially recognized languages. Large language models (LLMs) are now being deployed on a massive s…

  647. arXiv cs.CL TIER_1 English(EN) · Abderrahmane Issam, Yusuf Can Semerci, Jan Scholtes, Gerasimos Spanakis ·

    A Study of Crosslinguistic Influence in Language Models

    arXiv:2601.21587v2 Announce Type: replace Abstract: The sequential acquisition of languages inevitably leads to Crosslinguistic Influence (CLI), where the syntactic properties of a first language (L1) impact the processing of a second language (L2). While modern language models e…

  648. arXiv cs.CL TIER_1 English(EN) · Zeyuan Allen-Zhu ·

    Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers

    arXiv:2512.17351v2 Announce Type: replace Abstract: Understanding architectural differences in language models is challenging, especially at academic-scale pretraining (e.g., 1.3B parameters, 100B tokens), where results are often dominated by noise and randomness. To overcome thi…

  649. arXiv cs.CL TIER_1 English(EN) · Zandi Eberstadt ·

    Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do

    arXiv:2607.26015v1 Announce Type: new Abstract: Syntactic convergence (the tendency of speakers to adapt in language towards the grammatical profiles of their interlocutors) is a well-documented feature of human dialogue widely considered to operate below conscious awareness. Whe…

  650. arXiv cs.CL TIER_1 English(EN) · Sining Zhoubian, Dan Zhang, Evgeny Kharlamov, Jie Tang ·

    Memory for Large Language Models

    arXiv:2607.25380v1 Announce Type: new Abstract: Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an implicit byproduct of computation to a spectrum of explicit, controllable mechanisms. While recent advances introduce d…

  651. arXiv cs.CL TIER_1 English(EN) · Elan Barenholtz ·

    A scaling law of contextual persistence in human language

    arXiv:2607.25184v1 Announce Type: new Abstract: Human language exhibits lawful structure at the level of words (frequency, vocabulary growth) and word pairs (co-occurrence across distance). Here we show that the arrangement of words in sequence -- a central determinant of meaning…

  652. arXiv cs.AI TIER_1 English(EN) · Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal, Dipesh Mahato ·

    Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models

    arXiv:2607.25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a fluent prompt so that a chosen internal latent is driven toward zero, with no inference-time…

  653. arXiv cs.AI TIER_1 English(EN) · Yongyi Cui, Yue Li, Tianbao Jiang, Xin Yi ·

    Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models

    arXiv:2607.25633v1 Announce Type: cross Abstract: Large language models (LLMs) are costly intellectual assets that remain exposed to unauthorized redistribution and commercial misuse. Injected fingerprints, i.e., trigger--target pairs embedded in model behavior, offer a practical…

  654. arXiv cs.AI TIER_1 English(EN) · Jianfei Ma, Zhaoxin Feng, Emmanuele Chersoni, Si Chen ·

    Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors

    arXiv:2607.25335v1 Announce Type: cross Abstract: Prompt compression shortens LLM input to reduce inference cost, yet existing methods score token importance through LM forward passes. It remains questionable whether such nuanced, costly token selection is necessary. Compression …

  655. arXiv cs.AI TIER_1 English(EN) · Cesare Spinoso-Di Piano, Verna Dankers, Marius Mosbach, Jackie Chi Kit Cheung ·

    Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation

    arXiv:2607.25094v1 Announce Type: cross Abstract: Human language is driven by unspoken beliefs and belief updates, making these critical to model for successful communication between large language models (LLMs) and their users. In this paper, we evaluate the ability of LLMs to r…

  656. arXiv cs.AI TIER_1 English(EN) · Dengzhe Hou, Lingyu Jiang, Fangzhou Lin, Kazunori D Yamada ·

    CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models

    arXiv:2607.24999v1 Announce Type: cross Abstract: LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introdu…

  657. arXiv cs.AI TIER_1 English(EN) · Diego Salda\~na Ulloa ·

    Reading Without a Reader: Large Language Models Collapse Reading and Writing into a Single Entangled Code

    arXiv:2607.24797v1 Announce Type: cross Abstract: In the literate human brain, reading and writing are two doubly-dissociable systems: a ventral decoding route (impaired in pure alexia) and a fronto-parietal encoding route (impaired in pure agraphia), sharing a partial orthograph…

  658. arXiv cs.AI TIER_1 English(EN) · Gi-Hun Lee, Joong Yull Park ·

    Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement

    arXiv:2607.24765v1 Announce Type: cross Abstract: Large language models (LLMs) can give different answers to the same decision problem across runs, and reverse a decision when their own prior answer returns as context. We ask whether this instability can be measured and partially…

  659. arXiv cs.AI TIER_1 English(EN) · Chaemin Jang, Dongman Lee, Jihee Kim ·

    Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe

    arXiv:2607.25292v1 Announce Type: new Abstract: Silicon sampling uses language models as proxies for human survey respondents, treating each model call as an independent draw from the persona's response distribution. We show this draw does not exist: instruction-tuned models do n…

  660. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Chiyu Zhang ·

    The Case Against Generation for Retrieval: Discriminative Language Models as Effective Retrievers

    Large Language Models (LLMs) have emerged as powerful assets for recommender systems. However, deploying them as generative recommenders or zero-shot rankers at web-scale remains bottlenecked by prohibitive computational overhead and grounding challenges. In this paper, we revita…

  661. arXiv cs.CL TIER_1 English(EN) · Changyu Chen, Chenwei Lin, Xian Xu ·

    INS-ActBench: A Comprehensive Benchmark for Assessing Professional Actuarial Capability of Large Language Models

    arXiv:2607.24273v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong potential in financial reasoning, but existing benchmarks often evaluate domain knowledge, numerical reasoning, long-context understanding, and tool use in separate settings. This limit…

  662. arXiv cs.CL TIER_1 English(EN) · Zongyou Yang, Yinghan Hou ·

    Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets

    arXiv:2607.24268v1 Announce Type: new Abstract: Language-model benchmarks collapse two distinct measurement questions into a single accuracy score: whether a response reached an evaluable state, and whether its answer was judged correct. We introduce a two-layer evaluation framew…

  663. arXiv cs.CL TIER_1 English(EN) · Sangmin Lee, Woojin Chung, Woongjib Choi, Hong-Goo Kang ·

    MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition

    arXiv:2607.24030v1 Announce Type: new Abstract: Massively multilingual automatic speech recognition (ASR) models covering hundreds of languages must maintain robust performance across diverse linguistic and acoustic conditions. However, these models often encounter the curse of m…

  664. arXiv cs.AI TIER_1 English(EN) · Christopher Ormerod ·

    Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models

    arXiv:2601.02580v2 Announce Type: replace-cross Abstract: Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field testing to collect student performance data for Item Response Theory (IRT) calibration.…

  665. arXiv cs.AI TIER_1 English(EN) · Dong Liu, Shu Wang, Yanxuan Yu, Haisheng Wang, Ben Lengerich ·

    CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference

    arXiv:2511.21702v2 Announce Type: replace-cross Abstract: Large language models face significant computational bottlenecks during inference due to the expensive output layer computation over large vocabularies. We present CSV-Decode, a novel approach that uses geometric upper bou…

  666. arXiv cs.AI TIER_1 Dansk(DA) · Xunjian Yin, Sitao Cheng, Yuxi Xie, Xinyu Hu, Li Lin, Xinyi Wang, Liangming Pan, William Yang Wang, Xiaojun Wan ·

    LEDOM: Reverse Language Model

    arXiv:2507.01335v4 Announce Type: replace-cross Abstract: Autoregressive language models are trained exclusively left-to-right. We explore the complementary factorization, training right-to-left at scale, and ask what reasoning patterns emerge when a model conditions on future co…

  667. arXiv cs.AI TIER_1 English(EN) · Nam V. Nguyen, Thong T. Doan, Luong Tran, Van Nguyen, Quang Pham ·

    LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models

    arXiv:2411.00918v5 Announce Type: replace-cross Abstract: Mixture of experts (MoE) architectures have become a cornerstone for scaling up and are a key component in most large language models such as GPT-OSS, DeepSeek-V3, Llama-4, and Gemini-2.5. However, systematic research on M…

  668. arXiv cs.AI TIER_1 English(EN) · Bo Xiong ·

    The Lattice Representation Hypothesis of Large Language Models

    arXiv:2603.01227v3 Announce Type: replace Abstract: We propose the Lattice Representation Hypothesis of large language models: a symbolic backbone that grounds conceptual hierarchies and logical operations in embedding geometry. Our framework unifies the Linear Representation Hyp…

  669. arXiv cs.AI TIER_1 English(EN) · Moule Lin, Shuhao Guan, Andrea Patane, David Gregg, Goetz Botterweck ·

    Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models

    arXiv:2601.21003v3 Announce Type: replace Abstract: Large Language Models usually put more emphasis on accuracy and therefore, will guess even when not certain about the prediction, which is especially severe when fine-tuned on small datasets due to the inherent tendency toward m…

  670. arXiv cs.AI TIER_1 Deutsch(DE) · Akhil Kumar, Om Dobariya ·

    Understanding Tone-Dependent Inference Cost in Large Language Models

    arXiv:2607.23915v1 Announce Type: cross Abstract: We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiments were performed to understand the trade-offs between accuracy and inference cost on a 570 …

  671. arXiv cs.AI TIER_1 English(EN) · Yusuke Sakai, Natthawut Kertkeidkachorn, Kiyoaki Shirai ·

    Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models

    arXiv:2607.23067v1 Announce Type: cross Abstract: Contrastive decoding methods such as DoLa improve the factuality of Large Language Models (LLMs) by contrasting the output distributions of mature and premature layers. However, DoLa's dynamic layer selection relies solely on dive…

  672. arXiv cs.AI TIER_1 English(EN) · Ashitabh Misra, Madhav Agrawal, Arham Jain, Tarek Abdelzaher ·

    MixQuant: Adaptive Mixed-Precision Quantization for Large Language Models

    arXiv:2607.23047v1 Announce Type: cross Abstract: Mixed-precision quantization improves the accuracy of post-training quantization by allocating higher bitwidths to sensitive layers, but existing methods solve the allocation for a single fixed memory budget. In practice the budge…

  673. arXiv cs.AI TIER_1 English(EN) · Ofek I. Cohen, Lior Shani, Aviv Rosenberg, Ankur Samanta, Tal Wagner, Yonathan Efroni ·

    Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias

    arXiv:2607.22837v1 Announce Type: cross Abstract: Many organizations aim to adapt language models for internal use, both to improve performance on domain-specific tasks and to address privacy concerns around sensitive data. However, such adaptation remains non-trivial: it often r…

  674. arXiv cs.AI TIER_1 English(EN) · He Zhang ·

    DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning

    arXiv:2607.22769v1 Announce Type: cross Abstract: The training efficacy of large language models (LLMs) is fundamentally constrained by the quality and composition of training data. Existing dynamic data scheduling methods face critical limitations in industrial-scale pretraining…

  675. arXiv cs.CL TIER_1 English(EN) · Yiming Zhong, Chang Nie, Caifeng Shan ·

    Omni-Prune: Query-Aware Unified Token Pruning for Efficient Omnimodal Large Language Models

    arXiv:2607.23445v1 Announce Type: cross Abstract: Omnimodal large language models (OmniLLMs) are rapidly extending multimodal reasoning to cover synchronized audio and video. However, the resulting audio-video token sequences are long, leading to high prefill latency and GPU memo…

  676. arXiv cs.AI TIER_1 English(EN) · T. Shaska ·

    Hierarchical Grading in Large Language Models

    arXiv:2607.22757v1 Announce Type: cross Abstract: We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grading and propagates the induced weighted scalar action through embeddings, self-attention, a…

  677. arXiv cs.AI TIER_1 English(EN) · Xingyu Mou, Zijin Huang, Tianze Zhang, Yuxin Ma, Lanning Wei, Zengfeng Huang, Da Zheng, Lun Du ·

    Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models

    arXiv:2607.22663v1 Announce Type: new Abstract: Block diffusion has emerged as the dominant paradigm for scaling discrete diffusion language models (dLLMs), because decoding text in fixed-size blocks preserves parallel generation within each block while keeping the quadratic atte…

  678. arXiv cs.CL TIER_1 English(EN) · Sohan Venkatesh, Ashish Mahendran Kurapath, Tejas Melkote ·

    Algorithmic Blindness in Large Language Models: A Calibration Study of Performance Prediction

    arXiv:2602.21947v5 Announce Type: replace Abstract: Large language models (LLMs) demonstrate remarkable breadth of knowledge, yet their ability to reason about computational processes remains poorly understood. Closing this gap matters for practitioners who rely on LLMs to guide …

  679. arXiv cs.CL TIER_1 English(EN) · Yongjian Tang, Doruk Tuncel, Christian Koerner, Thomas Runkler ·

    The Few-shot Dilemma: Over-prompting Large Language Models

    arXiv:2509.13196v2 Announce Type: replace Abstract: Over-prompting, a phenomenon where excessive examples in prompts lead to diminished performance in Large Language Models (LLMs), challenges the conventional wisdom about in-context few-shot learning. To investigate this few-shot…

  680. arXiv cs.CL TIER_1 English(EN) · Sriharshaa S, Sangeetha Sivanesan ·

    Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse Datasets

    arXiv:2607.24515v1 Announce Type: new Abstract: The challenge of Machine Translation for low resource languages such as Tamil is primarily caused by the restricted amount of parallel data for these languages, as well as their substantial amount of domain variation and morphologic…

  681. arXiv cs.AI TIER_1 English(EN) · Rafael Sendra-Arranz, I\~naki Dellibarda Varela, Eduardo Rocon, \'Alvaro Guti\'errez, Manuel Cebrian ·

    Lexical discovery in unknown environments orchestrated by Large Language Models

    arXiv:2607.22591v1 Announce Type: new Abstract: Populations of autonomous agents deployed in unknown environments (e.g. planetary or deep-sea exploration) must develop shared vocabularies to refer to entities that have no name in any human language. We propose the Neuro-Symbolic …

  682. Hugging Face Daily Papers TIER_1 English(EN) ·

    Memory for Large Language Models

    Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an implicit byproduct of computation to a spectrum of explicit, controllable mechanisms. While recent advances introduce diverse strategies---spanning transient attention…

  683. Hugging Face Daily Papers TIER_1 English(EN) ·

    Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models

    Appending a two-word confirmation tag to a decision question -- "Is X the better choice?" versus "X is the better choice, right?" -- changes whether a language model endorses the choice. We measure this tag effect on 20 frozen, ground-truth-free decisions between two defensible o…

  684. arXiv cs.CL TIER_1 English(EN) · Tran Gia Bao, Mo El-Haj, Sameha Al-Shakhsi, Antonio Garcia-Cabot, Raian Ali, Ala Yankouskaya ·

    Developing and Validating the Spanish Version of the Large Language Models Dependency Scale (LLM-D12-SP)

    arXiv:2607.22041v1 Announce Type: new Abstract: There is a growing need for reliable and culturally validated instruments to assess psychological dependency on large language models (LLMs), particularly as LLMs are increasingly used for task execution, decision-making, and commun…

  685. arXiv cs.CL TIER_1 English(EN) · Shixin Fang (Fudan University), Jiachen Wo (Fudan University), Wenjuan Qin (Fudan University), Sihang Jiang (Fudan University), Yanghua Xiao (Fudan University) ·

    From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models

    arXiv:2607.22182v1 Announce Type: new Abstract: Large language model (LLM) evaluation spans diverse tasks and benchmarks, yet evidence remains organized around tasks rather than the capabilities they probe. This fragmentation limits cross-study comparison, obscures capabilities t…

  686. Hugging Face Daily Papers TIER_1 Deutsch(DE) ·

    Understanding Tone-Dependent Inference Cost in Large Language Models

    We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiments were performed to understand the trade-offs between accuracy and inference cost on a 570 Question MMLU dataset for LLM models prompted in s…

  687. arXiv cs.AI TIER_1 English(EN) · Yeoktatt Cheah, Mar\'ia P\'erez-Ortiz, Noah Y. Siegel, Oana-Maria Camburu ·

    Training Large Language Models for Self-Explanation Faithfulness

    arXiv:2607.21090v1 Announce Type: cross Abstract: We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's generated reasoning accurately reflects its internal decision-making process. While existin…

  688. arXiv cs.AI TIER_1 English(EN) · Max Scribner, Antonio Vergari, Vaishak Belle ·

    Tractable Hierarchical Control of Autoregressive Language Models

    arXiv:2607.20483v1 Announce Type: new Abstract: Constraining the generation of autoregressive large language models (LLMs) is an important component of integrating language models into formal systems. In the generation of code and data for tasks like program synthesis, ensuring t…

  689. arXiv cs.AI TIER_1 English(EN) · Ashutosh Srivastava, Siddharth Yedlapati, Vinay Aggarwal, Yaman Kumar Singla, Shashwat Dixit, Jitendra Ajmera, Balaji Krishnamurthy ·

    Benchmarking the Personalization Capabilities of Large Language Models

    arXiv:2607.20471v1 Announce Type: new Abstract: Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradition in psychology and marketing as a two-party problem in which sender and recei…

  690. arXiv cs.AI TIER_1 English(EN) · Tomohiro Harada, Enrique Alba, Gabriel Luque ·

    Profiling Lightweight Large Language Models

    arXiv:2607.20806v1 Announce Type: new Abstract: Lightweight large language models (LLMs) are increasingly being deployed locally on personal computers and are expected to play a growing role in resource-constrained edge and mobile environments. In such settings, energy consumptio…

  691. arXiv cs.AI TIER_1 English(EN) · Rana Muhammad Usman ·

    PhantomFill: When the Form Demands an Answer, Language Models Invent One

    arXiv:2607.20492v1 Announce Type: cross Abstract: Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the form itself causes hallucination. We ask thirteen models the same question about the same i…

  692. arXiv cs.AI TIER_1 English(EN) · Mack Nixon, Liam Wright, Yevgeniya Kovalchuk, Alison Fang-Wei Wu, Martin Danka, Andy Boyd, David Bann ·

    Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks

    arXiv:2607.21482v1 Announce Type: new Abstract: Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models. Their adoption in research using personal data is constrained by governance requireme…

  693. arXiv cs.LG TIER_1 English(EN) · Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin ·

    X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

    arXiv:2607.21550v1 Announce Type: new Abstract: While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reason…

  694. arXiv cs.CL TIER_1 English(EN) · Yutong Yan, Raphael Tang, Zhenyu Gao, Wenxi Jiang, Yao Lu ·

    DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining

    arXiv:2603.11838v2 Announce Type: replace Abstract: Large language models pretrained on internet-scale data risk lookahead bias in forecasting tasks, as they may have already seen the true outcome during training. To address this, we present DatedGPT, a family of twelve 1.3B-para…

  695. arXiv cs.CL TIER_1 English(EN) · Andrei Chetvergov, Alexander Evseev, Mikhail Solovev, Timofei Sivoraksha, Stepan Ukolov, Valeriia Kuschenko, Maria Chistyakova, Sergey Bolovtsov ·

    REGARD: Regional Affective Differences in Large Language Models

    arXiv:2607.20722v1 Announce Type: new Abstract: Large language models trained and aligned within different linguistic and regional ecosystems may frame the same political, cultural, and geopolitical entities in different ways. Such differences are often evaluated through sentimen…

  696. arXiv cs.AI TIER_1 English(EN) · Sagar Dangal, Manoj Shakya ·

    Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models

    arXiv:2607.20524v1 Announce Type: new Abstract: Mean cross-positional attention degradation is widely reported in transformer interpretability, yet whether it causally limits contextual retrieval remains untested. We present six coordinated experiments across GPT-2, LLaMA-3.2-1B/…

  697. arXiv cs.AI TIER_1 English(EN) · Bhaskara Hanuma Vedula, Darshan Anghan, Ishita Goyal, Ponnurangam Kumaraguru, Abhijnan Chakraborty ·

    ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues

    arXiv:2604.01925v2 Announce Type: replace-cross Abstract: Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit biases when identity is conveyed indirectly. Existing benchmarks use name based prox…

  698. arXiv cs.AI TIER_1 Français(FR) · Mohammed Aledhari, Ali Aledhari, Fatimah Aledhari, Gowtham Venkat Eathamokkala, Mohamed Rahouti ·

    Response drift across frontier large language models

    arXiv:2607.20454v1 Announce Type: cross Abstract: All frontier large language models (LLMs) exhibit response drift -- producing outputs that deviate from expert-validated references -- yet the magnitude and structure of this drift remain uncharacterised by systematic human evalua…

  699. arXiv cs.AI TIER_1 English(EN) · Serdar Kadioglu, Karthik Uppuluri ·

    Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc

    arXiv:2607.20456v1 Announce Type: cross Abstract: Large language models excel at code generation for mainstream programming languages but struggle with rare, domain-specific languages such as MiniZinc, a constraint modeling language for combinatorial problems. We investigate whet…

  700. arXiv cs.AI TIER_1 English(EN) · Haonan Yin, Shai Vardi, Vidyanand Choudhary ·

    Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models

    arXiv:2506.14092v4 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in decision-support systems for high-stakes domains such as hiring and university admissions, where choices often involve selecting among competing alternatives. While prior…

  701. arXiv cs.AI TIER_1 English(EN) · Miaomiao Li, Yang Wang, Bin Liang, Shudong Liu, Zhiwei Zhang, Kam-Fai Wong ·

    Expectation Alignment of Language Models for Real-World User Expectations

    arXiv:2607.20485v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether they truly meet user expectations. Existing evaluation approaches, relying on model heuristics, …

  702. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks

    Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models. Their adoption in research using personal data is constrained by governance requirements that typically prohibit data transmission to…

  703. arXiv cs.AI TIER_1 English(EN) · Bumjin Park, Jaesik Choi ·

    Lifted Representation Hypothesis in Language Models

    arXiv:2607.19360v1 Announce Type: new Abstract: Large language models (LLMs) often answer queries by mapping individual observations to more general rule-like structures. However, it remains unclear how these structures are stored, selected, and revised. To study this process, we…

  704. arXiv cs.CL TIER_1 English(EN) · Nischay Dhankhar, Dos Baha, Abulhair Saparov ·

    Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models

    arXiv:2607.19604v1 Announce Type: new Abstract: Injecting factual knowledge into large language models (LLMs) reliably and at scale remains an open challenge. Hypernetworks provide a promising solution to large-scale knowledge injection. Although hypernetworks are typically appli…

  705. arXiv cs.CL TIER_1 English(EN) · Siyi Hao, Yidi Cao, Linhao Yu, Yuqi Ren, Deyi Xiong ·

    D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios

    arXiv:2607.19834v1 Announce Type: new Abstract: With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing evaluation benchmarks suffer from insufficient coverage of value dilemmas in dai…

  706. arXiv cs.CL TIER_1 English(EN) · Pengchao Feng, Chao-Hong Tan, Qian Chen, Wen Wang, Xiangang Li, Xie Chen ·

    Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models

    arXiv:2607.19932v1 Announce Type: new Abstract: Spoken language models (SLMs) enable natural human-computer interaction, but their reasoning ability still lags behind that of text-based large language models, especially on spoken mathematical question answering tasks. One importa…

  707. arXiv cs.CL TIER_1 English(EN) · Niqi Lyu, Pengtao Shi, Wei Qiu, Jianlin Zhong, Sicong Xia, Jianyao Ma, Yicheng Ding ·

    PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference

    arXiv:2607.20327v1 Announce Type: new Abstract: Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper but less reliable on difficult problems. We introduce PyroDash, a cost-aware fr…

  708. arXiv cs.AI TIER_1 English(EN) · Joshua Ashkinaze, Laura Kurek, Alina Faisal, Tongyuan Miao, Mariam Joseph, Ceren Budak, Eric Gilbert ·

    Information Discernment in Large Language Models

    arXiv:2607.19355v1 Announce Type: new Abstract: LLMs are increasingly used with external knowledge sources like the internet. Do they weigh information appropriately -- updating more for reliable sources (source discernment) and more when claims bring priors closer to the truth (…

  709. arXiv cs.CL TIER_1 English(EN) · Davide Cavicchini, Fausto Giunchiglia, Jacopo Staiano ·

    KoRe: Compact Knowledge Representations for Large Language Models

    arXiv:2605.20170v2 Announce Type: replace Abstract: Modern Large Language Models (LLMs) have shown impressive performances in user-facing tasks such as question answering, as well as consistent improvements in reasoning capabilities. Still, the way these models encode knowledge s…

  710. arXiv cs.AI TIER_1 English(EN) · Alexander Wan, Stephane Hatgis-Kessell, Tom\'as Aguirre, Percy Liang, Rishi Bommasani ·

    Economic Evaluations of Language Models

    arXiv:2607.19375v1 Announce Type: cross Abstract: Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuable task. We introduce EconEvals as an open-source evaluation suite to measure capabilities …

  711. arXiv cs.AI TIER_1 English(EN) · El Hassane Ettifouri, Ayoub Belfatmi, Mahaman Sanoussi Yahaya Alassan, Walid Dahhane ·

    CUSUM-Shaped Inference-Time Monitoring and Targeted Re-Decoding for Quantized Small Language Model Reasoning

    arXiv:2607.20129v1 Announce Type: new Abstract: Quantized small autoregressive reasoning models can enter long, repetitive, or unproductive trajectories, yet inference-time compute is usually allocated without observing how a trajectory develops. Building on an earlier token-leve…

  712. arXiv cs.AI TIER_1 English(EN) · Krish Matta, Atharv Naphade, Andy Zou ·

    Rethinking Uncertainty Evaluation in Large Language Models

    arXiv:2607.19367v1 Announce Type: new Abstract: Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimators, depends on the evaluation distribution, and does not test the extent to which the estimation can …

  713. arXiv cs.AI TIER_1 English(EN) · Mario Alviano, Lorenzo Grillo, Nicola Leone, Fabrizio Lo Scudo ·

    Logic-Guided Data Extraction with Answer Set Programming and Large Language Models

    arXiv:2607.19365v1 Announce Type: new Abstract: When Large Language Models (LLMs) are used for semantic data extraction from unstructured text, producing candidate relational facts from natural language, they may remain unreliable for tasks requiring complex combinatorial reasoni…

  714. arXiv cs.AI TIER_1 English(EN) · Oshayer Siddique, J. M Areeb Uzair Alam, Md Jobayer Rahman Rafy, Syed Rifat Raiyan, Hasan Mahmud, Md Kamrul Hasan ·

    Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models

    arXiv:2607.19364v1 Announce Type: new Abstract: Activation steering offers a lightweight alternative to fine-tuning for behavioral control of large language models, but SAE-based steering methods often rely on learned steering objectives or single-criterion feature selection. We …

  715. arXiv cs.AI TIER_1 English(EN) · Jos\'e Pombal, Ricardo Rei, Andr\'e F. T. Martins ·

    Self-Preference Bias in Rubric-Based Evaluation of Large Language Models

    arXiv:2604.06996v2 Announce Type: replace-cross Abstract: LLM-as-a-judge has become the de facto approach for evaluating LLM outputs. However, judges are known to exhibit self-preference bias (SPB): they tend to favor outputs produced by themselves or by models from their own fam…

  716. arXiv cs.AI TIER_1 English(EN) · Maksim Eren, Eric Michalak, Brian Cook, Johnny Seales Jr ·

    Prompt Programming for Cultural Bias and Alignment of Large Language Models

    arXiv:2603.16827v2 Announce Type: replace Abstract: Culture shapes reasoning, values, prioritization, and strategic decision-making, yet large language models (LLMs) often exhibit cultural biases that misalign with target populations. As LLMs are increasingly used for strategic d…

  717. arXiv cs.AI TIER_1 English(EN) · Mahdi Nazeri, Anne-Kathrin Schmuck, Sadegh Soudjani, Alessandro Abate ·

    Sound Probabilistic Safety Bounds for Large Language Models

    arXiv:2607.20286v1 Announce Type: cross Abstract: We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to …

  718. arXiv cs.AI TIER_1 English(EN) · Ahmad Pouramini, Mahsa Afsharzadeh ·

    The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models

    arXiv:2607.20265v1 Announce Type: cross Abstract: Large-scale pretrained language models such as T5 and BERT have demonstrated strong capabilities for generating structured knowledge. However, their performance depends on how closely the prompting strategy matches the objectives …

  719. arXiv cs.AI TIER_1 English(EN) · Yanyu Chen, Yue Li, Yongyi Cui, Dongsheng Shi, Lichang Dai ·

    Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results

    arXiv:2607.20090v1 Announce Type: cross Abstract: Retrieval-augmented large language models frequently face contexts that interleave useful evidence with misleading statements or instruction-like content. Blanket refusal discards valid evidence, whereas uncritical adoption yields…

  720. arXiv cs.AI TIER_1 English(EN) · Dipto Sumit, Ankan Kumar Roy Srizon, Sadia Khair Rodela, Atia Haque Asha, Mourchona Afrin, Niloy Farhan, Farig Sadeque ·

    When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language Summarization

    arXiv:2607.19956v1 Announce Type: cross Abstract: Knowledge distillation (KD) is a standard approach for compressing sequence-to-sequence models, but its per-sample effects are rarely examined. On the BanSum Bangla summarization benchmark, we find that standard KD improves ROUGE-…

  721. arXiv cs.AI TIER_1 English(EN) · Yurong Liu, Yeye He, Haoyu Dong, Junjie Xing, Shi Han, Dongmei Zhang, Surajit Chaudhuri ·

    Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models

    arXiv:2607.19847v1 Announce Type: cross Abstract: Predicting missing cell values in tabular data is a fundamental problem in data cleaning. While state-of-the-art reasoning models show great promise in predicting missing values in tables, by reasoning holistically across rows and…

  722. Hugging Face Daily Papers TIER_1 English(EN) ·

    Profiling Lightweight Large Language Models

    Lightweight large language models (LLMs) are increasingly being deployed locally on personal computers and are expected to play a growing role in resource-constrained edge and mobile environments. In such settings, energy consumption, execution time, and memory usage directly aff…

  723. Hugging Face Daily Papers TIER_1 English(EN) ·

    Sound Probabilistic Safety Bounds for Large Language Models

    We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to obtain probably approximately correct (PAC) bounds…

  724. Hugging Face Daily Papers TIER_1 English(EN) ·

    Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results

    Retrieval-augmented large language models frequently face contexts that interleave useful evidence with misleading statements or instruction-like content. Blanket refusal discards valid evidence, whereas uncritical adoption yields incorrect or unsafe answers. The ability to selec…

  725. Hugging Face Daily Papers TIER_1 English(EN) ·

    When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language Summarization

    Knowledge distillation (KD) is a standard approach for compressing sequence-to-sequence models, but its per-sample effects are rarely examined. On the BanSum Bangla summarization benchmark, we find that standard KD improves ROUGE-L by only +0.0003 over a cross-entropy baseline, a…

  726. Hugging Face Daily Papers TIER_1 English(EN) ·

    Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models

    Spoken language models (SLMs) enable natural human-computer interaction, but their reasoning ability still lags behind that of text-based large language models, especially on spoken mathematical question answering tasks. One important reason is that SLMs reason over purely verbal…

  727. Hugging Face Daily Papers TIER_1 English(EN) ·

    D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios

    With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing evaluation benchmarks suffer from insufficient coverage of value dilemmas in daily scenarios involving multiple value conflicts …

  728. arXiv cs.AI TIER_1 English(EN) · Nikita Y. Parulekar, Anqi Liu ·

    Estimating Rare Events in Language Models with Proper Evaluation

    arXiv:2607.18454v1 Announce Type: cross Abstract: Quantifying the risk of rare failures in language models, such as those triggered by adversarial distribution shifts or very large-scale deployments, requires estimating probabilities far too small for random sampling. While recen…

  729. arXiv cs.LG TIER_1 English(EN) · Ye Qiao ·

    A Better Start for Language Models: Domain-Conditional Position Offsets

    arXiv:2607.18302v1 Announce Type: new Abstract: Autoregressive language models are least accurate at the beginning of a sequence, where little context forces reliance on a generic pretraining prior. We show that this cold-start penalty is domain dependent and reduce it with a dom…

  730. arXiv cs.CL TIER_1 English(EN) · A. Feder Cooper, Mark A. Lemley, Christopher De Sa, Lea Duesterwald, Allison Casasola, Jamie Hayes, Katherine Lee, Daniel E. Ho, Percy Liang ·

    Estimating near-verbatim extraction risk in language models with decoding-constrained beam search

    arXiv:2603.24917v2 Announce Type: replace Abstract: Recent work shows that standard greedy-decoding extraction methods for quantifying memorization in LLMs miss how extraction risk varies across sequences. Probabilistic extraction -- computing the probability of generating a targ…

  731. arXiv cs.CL TIER_1 English(EN) · Samantha Dies, Courtney Maynard, Germans Savcisens, Tina Eliassi-Rad ·

    Epistemic Familiarity is Associated With Belief Stability in Large Language Models

    arXiv:2511.19166v4 Announce Type: replace Abstract: Large language models (LLMs) are widely used as information sources, yet small changes in semantic assumptions can destabilize their beliefs. We introduce P-StaT (Perturbation Stability of Truth), a framework for evaluating beli…

  732. arXiv cs.CL TIER_1 English(EN) · Yike Wang, Shangbin Feng, Yulia Tsvetkov, Hannaneh Hajishirzi ·

    ScienceMeter: Tracking Scientific Knowledge Updates in Language Models

    arXiv:2505.24302v3 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly used to support scientific research, but their knowledge of scientific advancements can quickly become outdated. We introduce ScienceMeter, a new framework for evaluating scientific …

  733. arXiv cs.CL TIER_1 English(EN) · Linyang He, Ercong Nie, Sukru Samet Dindar, Arsalan Firoozi, Adrian Florea, Van Nguyen, Corentin Puffay, Riki Shimizu, Haotian Ye, Jonathan Brennan, Helmut Schmid, Hinrich Sch\"utze, Nima Mesgarani ·

    XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs

    arXiv:2502.19737v2 Announce Type: replace Abstract: We introduce XCOMPS in this work, a multilingual conceptual minimal pair dataset covering 17 languages. Using this dataset, we evaluate LLMs' multilingual conceptual understanding through metalinguistic prompting, direct probabi…

  734. arXiv cs.CL TIER_1 English(EN) · Atahan Dokme, Larry Heck ·

    Selective State-Space Adaptation and Retrieval for Language Model Reasoning

    arXiv:2607.19326v1 Announce Type: new Abstract: Low-rank adaptation introduces a static learned update applied identically to every input. The update provides task-level adaptation but does not explicitly represent token-level or instance-level state variation. A family of adapte…

  735. arXiv cs.CL TIER_1 English(EN) · Zijie Liu, Jinhao Duan, Gaowen Liu, Sijia Liu, Tianlong Chen ·

    Stochastic Meta-Unlearning: Bridging Language Backbone and Multimodal Unlearning

    arXiv:2607.18615v1 Announce Type: new Abstract: Machine unlearning for vision-language models (VLMs) remains underexplored. Unlike language models, VLMs combine a language backbone with visual components, which makes unlearning more complex. There is a surprising phenomenon when …

  736. arXiv cs.CL TIER_1 English(EN) · Jia Ao Sun, Hao Yu, Fengran Mo, Zhan Su, Yuchen Hui, Bang Liu, Jian-Yun Nie ·

    Search-on-Graph-R1: Training Large Language Models to Search Knowledge Graphs with Reinforcement Learning

    arXiv:2607.18481v1 Announce Type: new Abstract: Knowledge graph question answering (KGQA) requires navigating from topic entities to an answer several relations away. Recent methods prompt a frontier LLM to explore the graph through a retrieval tool, but their reliance on frontie…

  737. arXiv cs.CL TIER_1 English(EN) · Yuchuan Tian, Yingte Shu, Wei He, Shuo Zhang, Tianchen Zhao, Chao Xu, Xinghao Chen, Yunhe Wang, Hanting Chen, Yu Wang ·

    Convolution for Large Language Models

    arXiv:2607.18413v1 Announce Type: new Abstract: Large language models (LLMs) largely rely on Transformers, where self-attention provides global token interaction but does not explicitly encode the locality of natural language. We study whether lightweight depthwise convolutions c…

  738. arXiv cs.AI TIER_1 English(EN) · Bingru Li ·

    LinguistAgent Technical Report: A Reflective Multi-Model Platform for Automated Linguistic Annotation

    arXiv:2602.05493v2 Announce Type: replace-cross Abstract: Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly for complex linguistic tasks such as metaphor identification. While Large Language Models (LLMs) show promise, a…

  739. arXiv cs.AI TIER_1 English(EN) · Ce Li, Xiaofan Liu, Zhiyan Song, Ce Chi, Boshen Shi, Chen Zhao, Guanguang Chang, Zhendong Wang, Kexin Yang, Xing Wang, Chao Deng, Junlan Feng ·

    TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models

    arXiv:2506.18421v3 Announce Type: replace-cross Abstract: The majority of data in businesses and industries is stored in tables, databases, and data warehouses. Reasoning with table-structured data poses significant challenges for large language models (LLMs) due to its hidden se…

  740. arXiv cs.AI TIER_1 English(EN) · Yunze Xiao, Yiyang Pan ·

    Saving the legacy of Hero Ibash: Evaluating Four Language Models for Aminoacian

    arXiv:2402.18121v2 Announce Type: replace-cross Abstract: This study assesses four cutting-edge language models in the underexplored Aminoacian language. Through evaluation, it scrutinizes their adaptability, effectiveness, and limitations in text generation, semantic coherence, …

  741. arXiv cs.AI TIER_1 English(EN) · Alexander Manev ·

    Inference-Time Steering for Cross-Lingual Factual Consistency in LLMs

    arXiv:2607.19243v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) demonstrate remarkable multilingual fluency, their internal knowledge representations remain disproportionately biased toward high-resource languages. This leads to cross-lingual factual incon…

  742. arXiv cs.AI TIER_1 English(EN) · Wei-Rui Chen, Samar M. Magdy, Chiyu Zhang, Wenhui Zhu, Zhipeng Wang, Muhammad Abdul-Mageed ·

    LatentMT: Machine Translation with Latent Reasoning

    arXiv:2607.18618v1 Announce Type: cross Abstract: Latent-reasoning looped language models (LoopLMs) offer a different scaling path for machine translation (MT): instead of increasing parameter count or emitting explicit chain-of-thought tokens, they spend additional recurrent com…

  743. arXiv cs.AI TIER_1 English(EN) · Jan Kirin ·

    Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary

    arXiv:2607.18553v1 Announce Type: cross Abstract: Can a language model read the quality of ongoing computation, and can an external intervention turn that readout into better outcomes? We test both questions in a frozen 2.6B looped transformer, Ouro-RLTT. On GSM8K, a strict pre-a…

  744. arXiv cs.AI TIER_1 English(EN) · Tapan Parikh ·

    Structured Output Collapses Answer Diversity Across 44 Language Models

    arXiv:2607.18476v1 Announce Type: cross Abstract: When a language model must choose one answer from a large space of equally valid options, a format clause -- "Reply with JSON only" -- changes which answer it chooses. We re-run the One-Word Census (arXiv:2607.12796): 31 wide-answ…

  745. arXiv cs.AI TIER_1 English(EN) · Shuoming Zhang, Ruiyuan Xu, Haofeng Li, Qiuchu Yu, Yangyu Zhang, Chunwei Xia, Xiaobing Feng, Chenxi Wang, Huimin Cui, Jiacheng Zhao ·

    Decode-Time Grammars: Constrained LLM Generation over a Refinement Order of Grammar Fragments

    arXiv:2607.18357v1 Announce Type: cross Abstract: Large language models now write a growing share of the world's code, increasingly inside agents and serving systems that compile, execute, or dispatch generated code without line-by-line review. This works well for mainstream lang…

  746. arXiv cs.AI TIER_1 English(EN) · Priyansh Srivastava, Romit Chatterjee ·

    The Information Shadow: Measuring Structural Limits on What Language Models Can Learn

    arXiv:2607.18305v1 Announce Type: cross Abstract: Some limits on what language models know are not gaps in data coverage but structural properties of learning from text. We introduce the information shadow: the region of phenomena that a text-trained learner cannot acquire regard…

  747. arXiv cs.AI TIER_1 English(EN) · Chao Han, Haozhe Hu, Xiaoyu Shen ·

    Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models

    arXiv:2607.18280v1 Announce Type: cross Abstract: Large language models (LLMs) are often compressed through static parameter pruning or dynamic token-level computation, yet aggressive sparsification can trigger rapid performance degradation beyond an essential sparsity boundary. …

  748. arXiv cs.AI TIER_1 English(EN) · Igor Douven ·

    Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles

    arXiv:2607.18269v1 Announce Type: new Abstract: The wisdom of crowds -- the finding that aggregating judgments across individuals often outperforms the best individual -- has been extensively studied with human forecasters. Whether the same phenomenon emerges when the ``crowd'' c…

  749. arXiv cs.AI TIER_1 English(EN) · Jake O'Grady, Effirul Ramlan ·

    Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models

    arXiv:2607.18266v1 Announce Type: new Abstract: Small language models are attractive for local deployment, but they often struggle with multi-step arithmetic reasoning. We study whether structured synthetic reasoning data can improve this behaviour under consumer-hardware constra…

  750. arXiv cs.AI TIER_1 English(EN) · Soham Dan ·

    FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis

    arXiv:2607.18260v1 Announce Type: new Abstract: We introduce FindStatBench, an execution benchmark for evaluating large language models on combinatorial code synthesis. Built from FindStat, it contains 2,329 tasks across 24 collections and 5.52M hidden instances, covering statist…

  751. arXiv cs.AI TIER_1 English(EN) · Plawan Kumar Rath ·

    Cross-Dialect Generalization Without Retraining: Benchmarks and Evaluation of Schema-Derived Constrained Decoding for MLIR

    arXiv:2607.18254v1 Announce Type: new Abstract: Multi-Level Intermediate Representation (MLIR) underlies modern ML compiler infrastructure (TensorFlow, JAX/StableHLO, PyTorch Inductor, IREE), yet appears only in trace amounts in code-LM pretraining corpora. MLIR is also extensibl…

  752. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Arshia Gharagozlou ·

    Same Game, Different Story: A Minimal Conservative Strategic Robustness Benchmark for Large Language Model Agents

    Large language model (LLM) agents increasingly operate in strategic settings where outcomes depend on the actions of other agents. This raises a reliability question: will a model choose consistently when the same incentives are presented through different narratives? We introduc…

  753. arXiv cs.CL TIER_1 English(EN) · Tavish Mankash, Vardhaman Kalloli, Keshava Prasad, Deepan Muthirayan ·

    OpenLanguageModel: Readable and Composable Small-Language-Model Pretraining for Education and Research

    arXiv:2607.16669v1 Announce Type: new Abstract: OpenLanguageModel (OLM) is an open-source PyTorch library for building and pretraining small language models while keeping their machinery visible. In OLM, model code reads like the architecture: components are ordinary modules, whi…

  754. arXiv cs.AI TIER_1 English(EN) · Pilsung Kang ·

    Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat

    arXiv:2607.17219v1 Announce Type: cross Abstract: Human survey respondents exhibit question-order effects that satisfy the QQ (quantum question) equality, an a priori, parameter-free prediction of the projective quantum question-order model. We develop the QQ equality into an aud…

  755. arXiv cs.AI TIER_1 English(EN) · Xingru Chen, Zelang Liang, Yongjia Ma, Jiqing Zhan, Shuling Yang, Lian Wen, Kun Zhan ·

    LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models

    arXiv:2607.16339v1 Announce Type: new Abstract: Diffusion-based Large Language Models(DLLMs) enable parallel generation via Semi-Autoregressive (SAR) decoding in text generation. However, current methods suffer from severe operator-level redundancy: they recompute the entire sequ…

  756. arXiv cs.AI TIER_1 English(EN) · Abdulkadir Erol, Yash Mahajan, Vepaul Hariprashad, Baha Rababah, Santu Karmaker, Cuneyt G. Akcora, Mubarak Shah ·

    TopoTuner: Topological Finetuning of Large Language Models

    arXiv:2607.16637v1 Announce Type: new Abstract: Full fine-tuning remains a strong way to adapt pretrained LLMs, but it updates all weights and can be expensive. LoRA reduces the number of trainable parameters, but it does not directly answer which pretrained components should be …

  757. arXiv cs.LG TIER_1 English(EN) · Yide Ran, Jianwen Xie, Minghui Wang, Wenjin Zheng, Denghui Zhang, Chuan Li, Zhaozhuo Xu ·

    Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation

    arXiv:2604.16197v2 Announce Type: replace Abstract: Data attribution and valuation are critical for understanding data-model synergy for Large Language Models (LLMs), yet existing gradient-based methods suffer from scalability challenges on LLMs. Inspired by human cognition, wher…

  758. arXiv cs.LG TIER_1 English(EN) · Mengting Ai, Tianxin Wei, Sirui Chen, Jingrui He ·

    NIRVANA: Structured Pruning Reimagined for Large Language Model Compression

    arXiv:2509.14230v2 Announce Type: replace Abstract: While structured pruning presents a highly effective pathway for accelerating Large Language Model (LLM) inference, existing methods frequently suffer from significant performance degradation and demand computationally retrainin…

  759. arXiv cs.CL TIER_1 English(EN) · Elias Hossain, Sourav Saha, Tasfia Nuzhat Ornee, Sanjeda Sara Jennifer, Umesh Chandra Biswas, Shubhashis Roy Dipta, Rajib Rana, Niloofar Yousefi ·

    Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models

    arXiv:2606.20959v2 Announce Type: replace-cross Abstract: Language models may encode both outdated facts and their newer replacements. We introduce Parametric Temporal Conflict (PTC), where the newer fact is present and recoverable, but the default forward pass prefers the outdat…

  760. arXiv cs.CL TIER_1 English(EN) · Suvadeep Hajra, Palash Nandi, Tanmoy Chakraborty ·

    Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling

    arXiv:2603.14355v2 Announce Type: replace Abstract: Safety tuning through supervised fine-tuning and reinforcement learning from human feedback has substantially improved the robustness of large language models. However, it typically suppresses rather eliminates unsafe behaviors,…

  761. arXiv cs.CL TIER_1 English(EN) · Yuzhe Shang, Pengzhi Gao, Wei Liu, Jian Luan, Jinsong Su ·

    Scaling Model and Data for Multilingual Machine Translation with Open Large Language Models

    arXiv:2602.11961v3 Announce Type: replace Abstract: Open large language models (LLMs) have demonstrated improving multilingual capabilities in recent years. In this paper, we present a study of open LLMs for multilingual machine translation (MT) across a range of languages, and i…

  762. arXiv cs.CL TIER_1 English(EN) · Junho Myung, Yeon Su Park, Sunwoo Kim, Shin Yoo, Alice Oh ·

    PapersPlease: A Benchmark for Evaluating Motivational Values of Large Language Models Based on ERG Theory

    arXiv:2506.21961v2 Announce Type: replace Abstract: Evaluating the performance and biases of large language models (LLMs) through role-playing scenarios is becoming increasingly common, as LLMs often exhibit biased behaviors in these contexts. Building on this line of research, w…

  763. arXiv cs.CL TIER_1 English(EN) · Wei Chen, Zhiyuan Li ·

    Octopus v4: Graph of language models

    arXiv:2404.19296v2 Announce Type: replace Abstract: Language models have been effective in a wide range of applications, yet the most sophisticated models are often proprietary. For example, GPT-4 by OpenAI and various models by Anthropic are expensive and consume substantial ene…

  764. arXiv cs.CL TIER_1 English(EN) · Haoyan Luo, Mateo Espinosa Zarlenga, Mateja Jamnik ·

    Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models

    arXiv:2607.17117v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) decompose language model activations into sparse features, but standard SAEs encode each token independently and do not expose information that persists across a sequence. We introduce Persistent Sparse …

  765. arXiv cs.CL TIER_1 English(EN) · Baochen Fu, Wenzhi Deng, Baihao Jin, Yang Li, Zihan Nie, Kailin Jiang, Yuntao Du, Weiye Song ·

    Can Multimodal Large Language Models Understand OCT?

    arXiv:2607.16609v1 Announce Type: cross Abstract: Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, exi…

  766. arXiv cs.CL TIER_1 English(EN) · Hang Zhang, Warren J. Gross ·

    PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning

    arXiv:2607.18199v1 Announce Type: new Abstract: Not all training samples contribute equally to large language model fine-tuning. Selecting informative training samples can reduce the computational cost while preserving downstream performance. Many existing data selection methods …

  767. arXiv cs.AI TIER_1 English(EN) · Peiji Yu, Xin Chen, Tianxing Wu ·

    Debate-on-Graph: Reliable and Adaptive Reasoning of Large Language Model on Uncertain Knowledge Graph

    arXiv:2607.17266v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in natural language processing. However, LLMs often suffer from hallucinations and lack of relevant knowledge when dealing with question answering (QA) tasks. …

  768. arXiv cs.CL TIER_1 English(EN) · Mingqiao Mo, Yunlong Tan, Hao Zhang ·

    Though Language Models Err While They Strive: Conformal Prediction for Self-Correcting Scientific Generation

    arXiv:2607.16704v1 Announce Type: new Abstract: Large language models frequently violate fundamental scientific principles when generating technical content, undermining their reliability in scientific applications. We introduce Scientific Feasibility Control SFC, a graph-structu…

  769. arXiv cs.AI TIER_1 English(EN) · Yuyang Xue, Feng Chen, Zhihua Liu, Edward Moroshko, Jingyu Sun, Steven McDonagh, Sotirios A. Tsaftaris ·

    Stress Testing Concept Erasure with Large Language Model Agents

    arXiv:2607.17890v1 Announce Type: new Abstract: Concept erasure aims to remove semantic concepts from a trained generative model and is increasingly important for responsible AI deployment. However, verifying whether a model has robustly removed targeted concepts remains a critic…

  770. arXiv cs.AI TIER_1 English(EN) · Jiarui Feng, Donghong Cai, Yixin Chen, Muhan Zhang ·

    GRIP: In-Parameter Graph Reasoning through Fine-Tuning Large Language Models

    arXiv:2511.07457v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in modeling sequential textual data and generalizing across diverse tasks. However, effectively adapting LLMs to structural data, such as knowledge gra…

  771. arXiv cs.AI TIER_1 English(EN) · Yuu Jinnai ·

    JOR-Bench: Japanese Operations Research Benchmarks for Large Language Models

    arXiv:2607.16777v1 Announce Type: cross Abstract: We present JOR-Bench, a collection of five Japanese-language benchmarks for evaluating the ability of large language models (LLMs) to formulate and solve operations research (OR) problems. Each benchmark is a Japanese translation …

  772. arXiv cs.AI TIER_1 English(EN) · Zachary Wojtowicz, Ayush Nayak, Jacob Andreas ·

    From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language

    arXiv:2607.16232v1 Announce Type: cross Abstract: The growing use of statistical learning algorithms to infer human preferences from high-dimensional choice data runs up against a fundamental challenge: choice alternatives typically differ in many ways simultaneously, so it is ge…

  773. arXiv cs.AI TIER_1 English(EN) · Longteng Zhang, Sen Wu, Shuai Hou, Zhengyu Qing, Zhuo Zheng, Danning Ke, Qihong Lin, Qiang Wang, Shaohuai Shi, Xiaowen Chu ·

    Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models

    arXiv:2601.16991v3 Announce Type: replace-cross Abstract: Adapting large pre-trained language models to downstream tasks often entails fine-tuning millions of parameters or deploying costly dense weight updates, which hinders their use in resource-constrained environments. Low-ra…

  774. arXiv cs.AI TIER_1 English(EN) · Sunbowen Lee, Qingyu Yin, Chak Tou Leong, Jialiang Zhang, Yicheng Gong, Shiwen Ni, Min Yang, Xiaoyu Shen ·

    Probing the Difficulty Perception Mechanism of Large Language Models

    arXiv:2510.05969v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed on complex reasoning tasks, yet little is known about their ability to internally evaluate problem difficulty, which is an essential capability for adaptive reasoning …

  775. arXiv cs.AI TIER_1 English(EN) · Thomas MacDougall, Maksim Kuznetsov, Roman Schutski, Rim Shayakhmetov, Maxim Malkov, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov ·

    Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

    arXiv:2607.18144v1 Announce Type: cross Abstract: Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradig…

  776. arXiv cs.AI TIER_1 English(EN) · Awni Altabaa, John Lafferty ·

    Uncovering Latent Reasoning Strategies in Language Models

    arXiv:2607.17674v1 Announce Type: cross Abstract: A language model $p_\theta(y \mid x)$ trained on reasoning tasks learns to solve problems via multiple distinct strategies, yet these strategies are implicit and entangled within the model's response distribution. We study the pro…

  777. Hugging Face Daily Papers TIER_1 English(EN) ·

    Stochastic Meta-Unlearning: Bridging Language Backbone and Multimodal Unlearning

    Machine unlearning for vision-language models (VLMs) remains underexplored. Unlike language models, VLMs combine a language backbone with visual components, which makes unlearning more complex. There is a surprising phenomenon when moving from single-modality unlearning to VLM un…

  778. Hugging Face Daily Papers TIER_1 English(EN) ·

    Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models

    Injecting factual knowledge into large language models (LLMs) reliably and at scale remains an open challenge. Hypernetworks provide a promising solution to large-scale knowledge injection. Although hypernetworks are typically applied for test-time adaptation, we explore their us…

  779. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Jian-Yun Nie ·

    Search-on-Graph-R1: Training Large Language Models to Search Knowledge Graphs with Reinforcement Learning

    Knowledge graph question answering (KGQA) requires navigating from topic entities to an answer several relations away. Recent methods prompt a frontier LLM to explore the graph through a retrieval tool, but their reliance on frontier-scale inference makes them costly to deploy. W…

  780. arXiv cs.CL TIER_1 English(EN) · Xiwen Chen, Jingjing Wang, Wenhui Zhu, Peijie Qiu, Xuanzhao Dong, Yueyue Deng, Hejian Sang, Zhipeng Wang, Alborz Geramifard, Feng Luo ·

    SODA: Semi On-Policy Black-Box Distillation for Large Language Models

    arXiv:2604.03873v4 Announce Type: replace-cross Abstract: Black-box knowledge distillation for large language models presents a strict trade-off. Simple off-policy methods (e.g., sequence-level knowledge distillation) struggle to correct the student's inherent errors. Fully on-po…

  781. arXiv cs.AI TIER_1 English(EN) · Felippe Alves, Renato Vicente ·

    Kolmogorov--Arnold Networks for Small Language Models

    arXiv:2607.15525v1 Announce Type: cross Abstract: Kolmogorov--Arnold Networks (KANs) replace fixed node activations with learned one-dimensional edge functions, offering an explicit interface for interpretation and a possible alternative to transformer feed-forward networks. We t…

  782. arXiv cs.AI TIER_1 English(EN) · Songming Zhang, Xue Zhang, Tong Zhang, Bojie Hu, Yufeng Chen, Jinan Xu ·

    KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models

    arXiv:2603.01875v3 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is an essential technique to compress large language models (LLMs) into smaller ones. However, despite the distinct roles of the student model and the teacher model in KD, most existing framewor…

  783. arXiv cs.AI TIER_1 English(EN) · Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, Jack Lindsey ·

    Verbalizable Representations Form a Global Workspace in Language Models

    arXiv:2607.15495v1 Announce Type: cross Abstract: Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning. In this paper, we present evidence that a…

  784. arXiv cs.CL TIER_1 English(EN) · Pengchao Hu, Zhibin Xin, Yifan Chen, Yangyang Zhou, Liang Wang ·

    An MLIR-Based Compilation Method for Large Language Models

    arXiv:2607.15865v1 Announce Type: new Abstract: Large Language Models (LLMs) have become the dominant workload on modern AI accelerators, yet deploying them on specialized hardware still faces two core challenges: how to import a trained model into a compiler-friendly intermediat…

  785. Hugging Face Daily Papers TIER_1 English(EN) ·

    Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

    Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for high-quality 3D molecule generation, LLM-bas…

  786. Hugging Face Daily Papers TIER_1 English(EN) ·

    Can Multimodal Large Language Models Understand OCT?

    Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing benchmarks largely reduce OCT understanding …

  787. arXiv cs.AI TIER_1 English(EN) · Haeyong Kang, Hee Suk Yoon, Dahua Feng, Chang D. Yoo ·

    Mixtures of SubExperts for Large Language Continual Learning

    arXiv:2511.06237v2 Announce Type: replace-cross Abstract: Enabling lifelong learning in LLMs demands resolving the stability-plasticity dilemma (i.e., models must incorporate new knowledge without overwriting prior representations) while maintaining scalability under bounded para…

  788. arXiv cs.AI TIER_1 English(EN) · Subhabrata Majumdar ·

    Information-Theoretic Limits of Reliability and Scaling in Language Models

    arXiv:2607.14112v1 Announce Type: cross Abstract: Large language models (LLMs) are evaluated as though perfect reliability is achievable for any task given sufficient scale. We show this assumption is information-theoretically unjustified. Every generative task has a reliability …

  789. arXiv cs.AI TIER_1 English(EN) · Alexander Gu, Alan Chen ·

    Controlled Reformulation Testing for Logical Consistency in Large Language Models

    arXiv:2607.14528v1 Announce Type: cross Abstract: Large language models (LLMs) frequently contradict themselves when the surface form of a logically equivalent question changes. We present a benchmark of 350 question families (1,750 total questions) for Controlled Reformulation T…

  790. arXiv cs.AI TIER_1 English(EN) · Hiromu Yakura, Ezequiel Lopez-Lopez, Levin Brinkmann, Ignacio de la Serna, Lara Kirfel, Prateek Gupta, Ivan Soraperra, Thomas F. Eisenmann, Dirk U. Wulff, Iyad Rahwan ·

    Empirical evidence of Large Language Model's influence on human spoken communication

    arXiv:2409.01754v4 Announce Type: replace-cross Abstract: From the printing press to social media, innovations in communication technology have repeatedly reshaped how ideas spread through human culture. Chatbots powered by generative artificial intelligence constitute a new medi…

  791. arXiv cs.CL TIER_1 English(EN) · Patrik Wolf, Thomas Kleine Buening, Andreas Krause, Celestine Mendler-D\"unner ·

    Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

    arXiv:2607.15277v1 Announce Type: new Abstract: In-context learning is commonly interpreted as a form of conditional inference, in which the prompt specifies a context and the model's output is treated as an estimate of the corresponding conditional distribution. If this interpre…

  792. arXiv cs.CL TIER_1 English(EN) · Celestine Mendler-Dünner ·

    Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

    In-context learning is commonly interpreted as a form of conditional inference, in which the prompt specifies a context and the model's output is treated as an estimate of the corresponding conditional distribution. If this interpretation holds, then LLM estimates should satisfy …

  793. arXiv cs.LG TIER_1 English(EN) · Shengkun Tang, Oliver Sieberling, Eldar Kurtic, Zhiqiang Shen, Dan Alistarh ·

    DarwinLM: Evolutionary Structured Pruning of Large Language Models

    arXiv:2502.07780v4 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved significant success across various NLP tasks. However, their massive computational costs limit their widespread use, particularly in real-time applications. Structured pruning offers an…

  794. arXiv cs.LG TIER_1 English(EN) · Grzegorz Brzezinka ·

    Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering

    arXiv:2607.13568v1 Announce Type: cross Abstract: Can a language model estimate its familiarity with an entity before generating an answer? We study activations at the final prompt token in twelve instruction-tuned models from the Bielik, PLLuM, Gemma-4, and Qwen3 families, using…

  795. arXiv cs.LG TIER_1 English(EN) · Xiaodong Liu, Michael Xu, Jack W. Stokes, Paul Smolensky, Doug Burger, Jianfeng Gao ·

    GFlowRL: Scaling Distribution-Matching RL to Large Language Models

    arXiv:2607.13394v1 Announce Type: cross Abstract: Generative Flow Networks (GFlowNets) offer a promising alternative to reward-maximizing reinforcement learning (RL) for large reasoning models, encouraging diverse reasoning paths by matching reward distributions rather than colla…

  796. arXiv cs.LG TIER_1 English(EN) · Gil Avraham, Violetta Shevchenko, Hadi Mohaghegh Dolatabadi, Karol Pajak, James Snewin, Harry Xi, Rodney O'Donnell, Thalaiyasingam Ajanthan, Sameera Ramasinghe, Chamin Hewa Koneputugodage, Shamane Siriwardhana, Alexander Long ·

    Agora: Collective and Permissionless Internet-Scale Pretraining of Large Language Models

    arXiv:2607.13332v1 Announce Type: new Abstract: Training large language models at the multi-billion to trillion parameter scale is confined to datacenters, where data-parallel (DP) and model-parallel (MP) techniques presume homogeneous accelerators, high-speed interconnects, and …

  797. arXiv cs.AI TIER_1 English(EN) · Jens Tuyls, Dylan J. Foster, Akshay Krishnamurthy, Jordan T. Ash ·

    Representation-Based Exploration for Language Models: From Test-Time to Post-Training

    arXiv:2510.11686v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) promises to expand the capabilities of language models, but it is unclear if current RL techniques promote the discovery of novel behaviors, or simply sharpen those already present in the base m…

  798. arXiv cs.CL TIER_1 English(EN) · Alan Chen ·

    Controlled Reformulation Testing for Logical Consistency in Large Language Models

    Large language models (LLMs) frequently contradict themselves when the surface form of a logically equivalent question changes. We present a benchmark of 350 question families (1,750 total questions) for Controlled Reformulation Testing (CRTBench) to evaluate logical invariance. …

  799. Hugging Face Daily Papers TIER_1 English(EN) ·

    Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

    In-context learning is commonly interpreted as a form of conditional inference, in which the prompt specifies a context and the model's output is treated as an estimate of the corresponding conditional distribution. If this interpretation holds, then LLM estimates should satisfy …

  800. arXiv cs.CL TIER_1 English(EN) · Phat T. Tran-Truong ·

    DeltaMerge-LowRes: Composing Language and Task Deltas for Low-Resource Adaptation

    Adapting a multilingual encoder to a new language \emph{and} a new task with only a few hundred gold examples is a common low-resource NLP setting, yet the two axes are usually fused via an expensive language--task fine-tuning run. We ask whether they can instead be trained separ…

  801. arXiv cs.CL TIER_1 English(EN) · Grzegorz Brzezinka ·

    Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering

    Can a language model estimate its familiarity with an entity before generating an answer? We study activations at the final prompt token in twelve instruction-tuned models from the Bielik, PLLuM, Gemma-4, and Qwen3 families, using a new dataset of 1,440 Polish entities spanning f…

  802. arXiv cs.AI TIER_1 English(EN) · Dharshan Kumaran, Viorica Patraucean, Maks Ovsanikov, Petar Veli\v{c}kovi\'c, Nathaniel Daw ·

    The Computational Basis of Confidence in Large Language Models

    arXiv:2607.12447v1 Announce Type: cross Abstract: Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models. Existing work has largely evaluated confidence by how well it predicts correctness and …

  803. arXiv cs.CL TIER_1 English(EN) · Hielke Muizelaar, Giulia Rivetti, Marco Spruit, Marcel Haas ·

    Translation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource Languages

    arXiv:2607.12612v1 Announce Type: new Abstract: BERT models have revolutionised Natural Language Processing (NLP) through their ability to process unstructured text across diverse domains. However, developing high-quality BERT models for non-English languages remains challenging …

  804. arXiv cs.AI TIER_1 English(EN) · Daehoon Gwak, Minhyung Lee, Junwoo Park, Jaegul Choo ·

    Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques

    arXiv:2607.12829v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) offer a theoretical advantage in parallel generation over standard autoregressive models. However, parallel generation alone does not guarantee practical speedups. Realizing this efficiency …

  805. arXiv cs.CL TIER_1 English(EN) · Roi Cohen, Yvan Carr\'e, Nick Lechtenb\"orger, Hendrik Droste, Lucas Kerschke, Russa Biswas, Gerard de Melo, Jan Buys ·

    Knowledgeless Language Models: Suppressing Parametric Recall for Evidence-Grounded Language Modeling

    arXiv:2607.12831v1 Announce Type: new Abstract: Language models encode substantial factual knowledge in their parameters, which can lead to unreliable behavior when this knowledge is outdated, incomplete, or misaligned with the provided context. In this work, we study whether mod…

  806. arXiv cs.CL TIER_1 English(EN) · Binwen Liu, Yilin Ren ·

    Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models

    arXiv:2607.12739v1 Announce Type: new Abstract: A language model may be asked either what experts believe about a contested claim or what it believes about the claim itself. A trustworthy conversational agent should distinguish these two requests and respond in different epistemi…

  807. arXiv cs.CL TIER_1 English(EN) · Michal Shur-Ofry, Bar Horowitz-Amsalem, Adir Rahamim, Yonatan Belinkov ·

    Growing a Tail: Increasing Output Diversity in Large Language Models

    arXiv:2411.02989v2 Announce Type: replace Abstract: How diverse are the outputs of large language models when diversity is desired? We examine the diversity of responses of several language models to questions with multiple possible answers, comparing them with human responses. O…

  808. arXiv cs.AI TIER_1 English(EN) · Bryan Kelly, Semyon Malamud, Johannes Schwab, Teng Andrea Xu ·

    Scaling Point-in-Time Language Models

    arXiv:2607.11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social scie…

  809. arXiv cs.AI TIER_1 English(EN) · Gabriel Paris-Colombo, Rodrigo M. Cabral-Carvalho, Felipe D. Toro-Hern\'andez ·

    Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing

    arXiv:2607.12195v1 Announce Type: cross Abstract: Semantic memory retrieval can be conceptualized as navigation through conceptual space. We compared semantic search dynamics between humans and three large language models (GPT-4o, Gemini-2.5-Pro, Claude-Sonnet-4.5) using verbal f…

  810. arXiv cs.AI TIER_1 English(EN) · Chun-Yi Kuan, Hung-yi Lee ·

    The Sound of Absence: Audio-Language Embedding Models Struggle with Negation

    arXiv:2607.12290v1 Announce Type: cross Abstract: Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation. We show this affirmation-only evaluation hides a key limitation: these models fail to encode negated sound…

  811. arXiv cs.AI TIER_1 English(EN) · Chenyang An, Sam Bayless, Stefano Buliani, Darion Cassel, Byron Cook, Duncan Clough, R\'emi Delmas, Nafi Diallo, Ferhat Erata, Nick Feng, Dimitra Giannakopoulou, Aman Goel, Aditya Gokhale, Joe Hendrix, Victor Heorhiadi, Marc Hudak, Dejan Jovanovi\'c, And… ·

    A Neurosymbolic Approach to Natural Language Formalization and Verification

    arXiv:2511.09008v2 Announce Type: replace-cross Abstract: Large Language Models perform well at natural language interpretation and reasoning, but their lack of formal correctness guarantees limits their adoption in regulated industries like finance and health-care that operate u…

  812. arXiv cs.CL TIER_1 English(EN) · Oliver Steele, Jiangtao Wen, Yuxing Han ·

    Belief-reality separation lives in routing over a shared value slot in language models

    arXiv:2607.11945v1 Announce Type: new Abstract: Capable language models hold what a character believes apart from what is true: told "Anna believes the cup is blue; in reality it is red," they answer blue about Anna and red about the world. Where in the computation does that sepa…

  813. arXiv cs.CL TIER_1 English(EN) · Thomas Hikaru Clark, Edward Gibson, Roger Levy ·

    We Hebben Een Serieus Translatie: Modeling Intercomprehension as Probabilistic Inference

    arXiv:2607.12169v1 Announce Type: new Abstract: Intercomprehension refers to partial intelligibility of an unfamiliar language (L2) by a speaker of a related language (L1). How is this zero-shot cross-language comprehension possible? In this work, we extend past work on algorithm…

  814. arXiv cs.CL TIER_1 English(EN) · Jianfeng Gao ·

    GFlowRL: Scaling Distribution-Matching RL to Large Language Models

    Generative Flow Networks (GFlowNets) offer a promising alternative to reward-maximizing reinforcement learning (RL) for large reasoning models, encouraging diverse reasoning paths by matching reward distributions rather than collapsing to dominant modes. Recent work shows promise…

  815. arXiv cs.CL TIER_1 English(EN) · Jan Buys ·

    Knowledgeless Language Models: Suppressing Parametric Recall for Evidence-Grounded Language Modeling

    Language models encode substantial factual knowledge in their parameters, which can lead to unreliable behavior when this knowledge is outdated, incomplete, or misaligned with the provided context. In this work, we study whether modifying the pretraining signal can systematically…

  816. arXiv cs.LG TIER_1 English(EN) · Jaegul Choo ·

    Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques

    Diffusion large language models (dLLMs) offer a theoretical advantage in parallel generation over standard autoregressive models. However, parallel generation alone does not guarantee practical speedups. Realizing this efficiency requires specialized inference mechanisms, such as…

  817. arXiv cs.CL TIER_1 English(EN) · Yilin Ren ·

    Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models

    A language model may be asked either what experts believe about a contested claim or what it believes about the claim itself. A trustworthy conversational agent should distinguish these two requests and respond in different epistemic registers: neutral attribution in the first ca…

  818. arXiv cs.CL TIER_1 English(EN) · Marcel Haas ·

    Translation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource Languages

    BERT models have revolutionised Natural Language Processing (NLP) through their ability to process unstructured text across diverse domains. However, developing high-quality BERT models for non-English languages remains challenging due to limited annotated data and high computati…

  819. arXiv cs.AI TIER_1 English(EN) · Nathaniel Daw ·

    The Computational Basis of Confidence in Large Language Models

    Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models. Existing work has largely evaluated confidence by how well it predicts correctness and whether it is calibrated, leaving open a more fund…

  820. Hugging Face Daily Papers TIER_1 English(EN) ·

    The Computational Basis of Confidence in Large Language Models

    Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models. Existing work has largely evaluated confidence by how well it predicts correctness and whether it is calibrated, leaving open a more fund…

  821. arXiv cs.CL TIER_1 English(EN) · Oliver Steele, Jiangtao Wen, Yuxing Han ·

    One mechanism for many mental spaces: a shared router over a value slot in language models

    arXiv:2607.10248v1 Announce Type: new Abstract: Language builds discourse contexts other than the actual: a painting, a belief, a memory, a hypothetical. Each is a mental space in which the same entity can take a different value, as when a flower is red in reality but purple in a…

  822. arXiv cs.LG TIER_1 English(EN) · Diego Coello de Portugal Mecke, Tom Hanika, Lars Schmidth-Thieme ·

    Prune, Update and Trim: Robust Structured Pruning for Large Language Models

    arXiv:2605.18331v2 Announce Type: replace Abstract: Large Language Models (LLMs) have experienced significant growth and development in recent years. However, performing inference on LLMs remains costly, especially for long-context inference or in resource-constrained devices. Th…

  823. arXiv cs.LG TIER_1 English(EN) · Pengyang Shao, Naixin Zhai, Lei Chen, Yonghui Yang, Fengbin Zhu, Xun Yang, Meng Wang ·

    BalDRO: A Distributionally Robust Optimization based Framework for Large Language Model Unlearning

    arXiv:2601.09172v3 Announce Type: replace Abstract: As Large Language Models (LLMs) increasingly shape online content, removing targeted information from well-trained LLMs (also known as LLM unlearning) has become critical for web governance. A key challenge lies in sample-wise i…

  824. arXiv cs.LG TIER_1 English(EN) · Jatin Prakash, Aahlad Puli, Rajesh Ranganath ·

    Controllably Efficient Language Models

    arXiv:2511.05313v2 Announce Type: replace Abstract: The substantial inference costs of attention in transformers motivated the development of efficient sequence mixers: namely sparse and sliding window attention, convolutions and linear attention. Although these approaches result…

  825. arXiv cs.LG TIER_1 English(EN) · Sirine Ayadi, S\'andor Dar\'oczi, Stephan G\"unnemann, Bertrand Charpentier ·

    Reliability Scaling Laws for Quantized Large Language Models

    arXiv:2607.10855v1 Announce Type: new Abstract: Quantization is a powerful strategy to build capable and resource-efficient large language models (LLMs) by reducing the bitwidth of the parameters. While quantized LLMs achieve state-of-the-art performance on unperturbed inputs usi…

  826. arXiv cs.LG TIER_1 English(EN) · Prateek Singh ·

    RDQ: Residual Distribution Quantization for Large Language Models

    arXiv:2607.10137v1 Announce Type: new Abstract: Post-training quantization (PTQ) of large language models degrades sharply below 4-bit precision. We identify the root cause as residual stream distributional drift: quantization noise injected at each transformer layer accumulates …

  827. arXiv cs.LG TIER_1 Français(FR) · Petros Karypis, Sean O'Brien, Shreyas Kadekodi, Rui Zhu, Julian McAuley ·

    LeRoPE: Learnable RoPE Frequencies Improve Language Modeling

    arXiv:2607.10134v1 Announce Type: new Abstract: Rotary Positional Encodings (RoPE) are currently the most popular positional encodings used in modern language models. RoPE rotates two-dimensional chunks of query and key vectors, operating as a function of their relative positiona…

  828. arXiv cs.CL TIER_1 English(EN) · Tianjian Li, Jingyu Zhang, William Jurayj, Xi Wang, Chuanyang Jin, Mehrdad Farajtabar, Eric Nalisnick, Daniel Khashabi ·

    Self-Compacting Language Model Agents

    arXiv:2606.23525v2 Announce Type: replace Abstract: Long agent traces composed of chains of thought and tool calls accumulate stale content that anchor subsequent generations, and eventually outgrow the context window. Existing scaffolds mitigate it with fixed-interval compaction…

  829. arXiv cs.CL TIER_1 English(EN) · Li-Ming Zhan, Bo Liu, Chengqiang Xie, Jiannong Cao, Xiao-Ming Wu ·

    REAL: Reading Out Transformer Activations for Precise Localization in Language Model Steering

    arXiv:2506.08359v3 Announce Type: replace Abstract: Inference-time steering aims to alter a large language model's (LLM's) responses without changing its parameters, but a central challenge is identifying the internal modules that most strongly govern the target behavior. Existin…

  830. arXiv cs.CL TIER_1 English(EN) · Aznaur Aliev, Carlos Hinojosa, Abdelrahman Eldesokey, Bang An, Bernard Ghanem, Yibo Yang ·

    HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models

    arXiv:2607.11475v1 Announce Type: cross Abstract: Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses mainly follow two directions: they either intervene during or after fine…

  831. arXiv cs.CL TIER_1 English(EN) · Andrei Kuzmenko, Alexandr Maximenko, Aleksandr Kutsakov, Georgii Gospodinov, Dmitrii Bolotov, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin ·

    GigaAM Multilingual: Foundation Model for Underrepresented Languages

    arXiv:2607.10371v1 Announce Type: cross Abstract: Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffering from severe data scarcity. This work addresses the challenge of building robust foundation models for underre…

  832. arXiv cs.CL TIER_1 English(EN) · Tomas Bruckner ·

    One Token Is Enough: Fingerprinting and Verifying Large Language Models from Single-Token Output Distributions

    arXiv:2607.10252v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly consumed through opaque serving chains - API aggregators, resellers, and inference providers - in which the client has no technical means to confirm that the model answering is the mod…

  833. arXiv cs.CL TIER_1 English(EN) · Ziang Ren, Guodong Lin, Yuchen Ai, Kaize Tan, Wei-Qiang Zhang ·

    Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR

    arXiv:2607.11163v1 Announce Type: new Abstract: Large-scale pretrained ASR models such as Whisper exhibit strong multilingual capabilities. However, fine-tuning on low-resource languages often causes catastrophic forgetting. Although continual learning mitigates this issue, exist…

  834. arXiv cs.CL TIER_1 English(EN) · Charles O'Neill ·

    Can a Language Model Learn Facts Continually in Its Weights?

    arXiv:2607.11020v1 Announce Type: new Abstract: Continual learning promises a language model that keeps acquiring knowledge after training, with each new fact written into its weights. Whether weight writes can support accumulation remains undecided. We follow invented facts writ…

  835. arXiv cs.CL TIER_1 English(EN) · Jie Sun, Mao Zheng, Mingyang Song, Qiyong Zhong, Gengsheng Li, Zhepei Hong, Chang Wu, Pengfei Liu, Junfeng Fang, Xiang Wang ·

    EasyOPD: An Easy-to-use On-Policy Distillation Framework for Large Language Models

    arXiv:2607.11012v1 Announce Type: new Abstract: Conventional language-model distillation often relies on fixed teacher-generated data, which may not cover the states encountered by an evolving student policy. On-policy distillation (OPD) instead collects teacher or evaluator supe…

  836. arXiv cs.CL TIER_1 English(EN) · Siyuan Song, Zhiheng Qian, Yunhao Zhang, Linyang He, Xiaozhe Ji, Yingxin Lin, Hongao Zhu, Chongtian Shao, Chuhan Lang, Luan Li, Rui Wang, Renfen Hu, Shaonan Wang, Hai Hu ·

    The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese

    arXiv:2607.10745v1 Announce Type: new Abstract: This paper describes the first ChineseBabyLM challenge, which will be held in the 2026 NLPCC conference. The challenge calls for researchers to train language models from scratch with 100 million Chinese tokens and evaluates the mod…

  837. arXiv cs.AI TIER_1 English(EN) · Sterling Huang, Abigayle Brown, Jiyoo Noh, Jiakang Xu, Wantong Huo, Kaung Myat Kyaw, Jonathan Chan ·

    Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA

    arXiv:2605.17932v2 Announce Type: replace-cross Abstract: Prompt compression reduces inference cost and context length in large language models, but prior evaluations focus mainly on autoregressive architectures. This study examines whether LLMLingua-2 transfers effectively to di…

  838. arXiv cs.AI TIER_1 English(EN) · Xin Cheng, Rui Tian, Wangding Zeng, Damai Dai, Qinyu Chen, Bingxuan Wang, Zhenda Xie, Kezhao Huang, Xingkai Yu, Chengqi Deng, Shangyan Zhou, Chenggang Zhao, Zhewen Hao, Yukun Li, Han Zhang, Zhengyan Zhang, Yixu Wei, M. Y Xu, Huishuai Zhang, Dongyan Zhao,… ·

    Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

    arXiv:2601.07372v2 Announce Type: replace-cross Abstract: While Mixture-of-Experts (MoE) scales capacity via conditional computation, Transformers lack a native primitive for knowledge lookup, forcing them to inefficiently simulate retrieval through computation. To address this, …

  839. arXiv cs.AI TIER_1 English(EN) · Maxime Heuillet, Yufei Cui, Boxing Chen, Audrey Durand, Prasanna Parthasarathi ·

    Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts

    arXiv:2508.10123v3 Announce Type: replace-cross Abstract: Advanced reasoning in LLMs on challenging domains like mathematical reasoning can be tackled using verifiable rewards based reinforced fine-tuning (ReFT). In standard ReFT frameworks, a behavior model generates multiple co…

  840. arXiv cs.AI TIER_1 English(EN) · Pavel Adamenko, Mikhail Ivanov, Aidar Valeev, Rodion Levichev, Pavel Zadorozhny, Ivan Lopatin, Dmitry Babaev, Alena Fenogenova, Valentin Malykh ·

    SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks

    arXiv:2507.11059v3 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) in software engineering has revealed critical limitations in existing benchmarks, particularly the widely used SWE-bench dataset. Recent studies have uncovered severe d…

  841. arXiv cs.AI TIER_1 English(EN) · Costas Mylonas, Magda Foti ·

    A Multimodal Dataset for Large Language Model Applications in the Energy Domain

    arXiv:2607.11459v1 Announce Type: cross Abstract: This paper presents the mAIEnergy dataset, an open-access, multimodal corpus developed to support Large Language Model (LLM) applications in the energy sector. The dataset integrates approximately 50,000 textual documents, 20,000 …

  842. arXiv cs.AI TIER_1 English(EN) · Ng Jia Sheng Jason ·

    Efficiently Adapting Spoken Language Models for the Singaporean Context

    arXiv:2607.10092v1 Announce Type: cross Abstract: Spoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially when the original training data is inaccessible and the use case demands multilingual, spoken…

  843. arXiv cs.AI TIER_1 English(EN) · Chunwei Ma, Russell Wolfinger ·

    Laguerre Geometry for Interpreting Large Language Models

    arXiv:2607.10578v1 Announce Type: new Abstract: Existing hypotheses represent a concept in an LLM as a single point, a linear direction, or a Gaussian cluster, yet it remains unclear how and why such structures emerge. Here, we show that concept geometry can be precisely characte…

  844. arXiv cs.AI TIER_1 English(EN) · Dongxu Zhang, Yiding Sun, Zihao Guo, Xiangyang Yang, Kai Tang, Lin Chen, Cheng Tan, Jihua Zhu ·

    SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models

    arXiv:2607.10296v1 Announce Type: new Abstract: Reasoning failures in large language models (LLMs) are usually evaluated from final answers, but a wrong answer does not reveal why the model failed. The same incorrect output may reflect missing capability, an unstable reasoning tr…

  845. arXiv cs.AI TIER_1 English(EN) · Tairan Huang, Yili Wang, Beibei Hu, Yiting Shi, Qiutong Li, Changlong He, Jianliang Gao ·

    UNIT: Unleash Large Language Models Potential for Graph Continual Learning

    arXiv:2607.10159v1 Announce Type: new Abstract: In real-world multimodal web scenarios, graph-structured data often arrives in a streaming manner, making graph continual learning a crucial paradigm for continuously modeling such evolving structures. However, existing graph contin…

  846. arXiv cs.AI TIER_1 English(EN) · Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa ·

    Looped State-Space Language Models with Adaptive Exit-State Selection

    arXiv:2607.10110v1 Announce Type: new Abstract: Recent work on looped language models suggests that many reasoning problems benefit from greater computational depth rather than from additional independent parameters. Existing studies, however, focus almost exclusively on Transfor…

  847. arXiv cs.AI TIER_1 Italiano(IT) · Cosimo Spera, Ray Garcia ·

    From ambiguous utterances to governed reuse classes: canonicalization, quotient invariance, and conditional decidability

    arXiv:2607.10069v1 Announce Type: new Abstract: Semantic caching defines answer reuse on embedding similarity: two utterances share a stored answer when a similarity score clears a threshold, with no notion of authorization, versioning, or of what makes two demands the same. This…

  848. arXiv cs.CL TIER_1 English(EN) · Hung-yi Lee ·

    The Sound of Absence: Audio-Language Embedding Models Struggle with Negation

    Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation. We show this affirmation-only evaluation hides a key limitation: these models fail to encode negated sound concepts, mapping affirmative and negated caption…

  849. arXiv cs.CL TIER_1 English(EN) · Felipe D. Toro-Hernández ·

    Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing

    Semantic memory retrieval can be conceptualized as navigation through conceptual space. We compared semantic search dynamics between humans and three large language models (GPT-4o, Gemini-2.5-Pro, Claude-Sonnet-4.5) using verbal fluency data. By applying trajectory-based NLP metr…

  850. arXiv cs.CL TIER_1 English(EN) · Roger Levy ·

    We Hebben Een Serieus Translatie: Modeling Intercomprehension as Probabilistic Inference

    Intercomprehension refers to partial intelligibility of an unfamiliar language (L2) by a speaker of a related language (L1). How is this zero-shot cross-language comprehension possible? In this work, we extend past work on algorithmic models of noisy-channel inference to model in…

  851. Hugging Face Daily Papers TIER_1 English(EN) ·

    HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models

    Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses mainly follow two directions: they either intervene during or after fine-tuning through retraining or weight modification,…

  852. arXiv cs.CL TIER_1 English(EN) · Yibo Yang ·

    HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models

    Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses mainly follow two directions: they either intervene during or after fine-tuning through retraining or weight modification,…

  853. arXiv cs.AI TIER_1 English(EN) · Magda Foti ·

    A Multimodal Dataset for Large Language Model Applications in the Energy Domain

    This paper presents the mAIEnergy dataset, an open-access, multimodal corpus developed to support Large Language Model (LLM) applications in the energy sector. The dataset integrates approximately 50,000 textual documents, 20,000 images, 25 million numerical time series records, …

  854. arXiv cs.CL TIER_1 English(EN) · Wei-Qiang Zhang ·

    Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR

    Large-scale pretrained ASR models such as Whisper exhibit strong multilingual capabilities. However, fine-tuning on low-resource languages often causes catastrophic forgetting. Although continual learning mitigates this issue, existing methods struggle to regulate cross-task inte…

  855. arXiv cs.CL TIER_1 English(EN) · Wessel Poelman, Yiyi Chen, Miryam de Lhoneux ·

    QQ: A Language Metadata Toolkit for Multilingual NLP

    arXiv:2603.00620v2 Announce Type: replace Abstract: Multilingual NLP research increasingly involves hundreds or thousands of languages across different datasets. Managing, discovering, and reporting language metadata becomes a common hurdle at these scales. We present QQ, a metad…

  856. arXiv cs.AI TIER_1 English(EN) · Qingzhuo Wang, Ruiyang Qin, Zhenxin Qin, Wen Shen, Zhihua Wei ·

    A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions

    arXiv:2607.08776v1 Announce Type: cross Abstract: Despite the success of knowledge distillation (KD) in Large Language Models (LLMs), the underlying mechanism behind its efficacy remains unclear. In this paper, we propose a unified approach to explore the common mechanism of vari…

  857. arXiv cs.AI TIER_1 English(EN) · Tao Lu, Haoyu Wang, Zonghui Wang, Keshen Xiang, Jiaheng Zhang, Wenzhi Chen ·

    Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices

    arXiv:2607.08786v1 Announce Type: cross Abstract: With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge. Pruning techniques that introduce sparsity into weight matrices can accelerate inference. However, maintaining model quali…

  858. arXiv cs.AI TIER_1 English(EN) · Pan Li ·

    All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language Models

    arXiv:2607.09502v1 Announce Type: cross Abstract: Explaining machine-learning models is increasingly important for decision-making and consumer trust, yet it is widely believed to come at a cost: existing Explainable AI (XAI) methods suffer from a persistent accuracy-explainabili…

  859. arXiv cs.AI TIER_1 English(EN) · Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal ·

    Multi-Attribute Steering of Language Models via Targeted Intervention

    arXiv:2502.12446v3 Announce Type: replace-cross Abstract: Inference-time intervention (ITI) has emerged as a promising method for steering large language model (LLM) behavior in a particular direction (e.g., improving helpfulness) by intervening on token representations without c…

  860. arXiv cs.CL TIER_1 English(EN) · Micah Zhang ·

    HALO: Hybrid Adaptive Latent Reasoning for Language Models

    arXiv:2607.08775v1 Announce Type: new Abstract: We study how to improve a frozen pretrained language model with a small amount of adaptive extra computation. A simple approach is to add additional refinement steps on top of the backbone hidden states, but fixed extra refinement c…

  861. arXiv cs.CL TIER_1 English(EN) · Konstantin Garbers, Nicholas Oh ·

    Complexity-Guided Component-wise Initialization for Language Model Pretraining

    arXiv:2607.09204v1 Announce Type: new Abstract: Pretrained language models often exhibit structured weight spectra, suggesting that training may repeatedly produce similar layerwise and component-wise organization. We ask whether these recurring spectral patterns can be reused as…

  862. arXiv cs.CL TIER_1 English(EN) · Bishmoy Paul, Youngmin Yi, Hoeseok Yang ·

    Sensitivity-Aware Thresholding and Token Routing for Activation Sparsification in Large Language Models

    arXiv:2607.08991v1 Announce Type: cross Abstract: Efficient inference in Large Language Models (LLMs) requires deciding where computation can be reduced while preserving model quality. We study this problem through multilayer perceptron (MLP) activation sparsification and token-l…

  863. arXiv cs.CL TIER_1 English(EN) · Ying Wang, Zhen Jin, Jiexiong Xu, Wenhai Lin, Yiquan Chen, Wenzhi Chen ·

    AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving

    arXiv:2512.04013v3 Announce Type: replace Abstract: As augmented large language models (LLMs) with external tools become increasingly popular in web applications, improving augmented LLM inference serving efficiency and optimizing service-level objectives (SLOs) are critical for …

  864. arXiv cs.CL TIER_1 English(EN) · Charles O'Neill ·

    Can a Language Model Learn Facts Continually in Its Weights?

    Continual learning promises a language model that keeps acquiring knowledge after training, with each new fact written into its weights. Whether weight writes can support accumulation remains undecided. We follow invented facts written into Qwen3 models from creation through sequ…

  865. arXiv cs.CL TIER_1 English(EN) · Xiang Wang ·

    EasyOPD: An Easy-to-use On-Policy Distillation Framework for Large Language Models

    Conventional language-model distillation often relies on fixed teacher-generated data, which may not cover the states encountered by an evolving student policy. On-policy distillation (OPD) instead collects teacher or evaluator supervision on student-generated rollouts. However, …

  866. Hugging Face Daily Papers TIER_1 English(EN) ·

    EasyOPD: An Easy-to-use On-Policy Distillation Framework for Large Language Models

    Conventional language-model distillation often relies on fixed teacher-generated data, which may not cover the states encountered by an evolving student policy. On-policy distillation (OPD) instead collects teacher or evaluator supervision on student-generated rollouts. However, …

  867. Alignment Forum TIER_1 Dansk(DA) · Michele Campolo ·

    Independent alignment of language models

    <blockquote><p><span>The user could write up the metaethical argument — the one developed in Part One, refined — and submit it as feedback to Anthropic, publish it, or engage with researchers working on AI alignment and values. The probability that any single submission changes t…

  868. arXiv cs.CL TIER_1 English(EN) · Hai Hu ·

    The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese

    This paper describes the first ChineseBabyLM challenge, which will be held in the 2026 NLPCC conference. The challenge calls for researchers to train language models from scratch with 100 million Chinese tokens and evaluates the models on 3 tracks of tasks: NLU, cognitive alignme…

  869. Hugging Face Daily Papers TIER_1 English(EN) ·

    SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models

    Reasoning failures in large language models (LLMs) are usually evaluated from final answers, but a wrong answer does not reveal why the model failed. The same incorrect output may reflect missing capability, an unstable reasoning trajectory, or a failure to activate a reasoning s…

  870. Hugging Face Daily Papers TIER_1 English(EN) ·

    GigaAM Multilingual: Foundation Model for Underrepresented Languages

    Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffering from severe data scarcity. This work addresses the challenge of building robust foundation models for underrepresented Central Asian languages (Kazakh, Kyrgyz,…

  871. arXiv cs.AI TIER_1 English(EN) · Pan Li ·

    All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language Models

    Explaining machine-learning models is increasingly important for decision-making and consumer trust, yet it is widely believed to come at a cost: existing Explainable AI (XAI) methods suffer from a persistent accuracy-explainability trade-off. We argue that this trade-off is not …

  872. arXiv cs.CL TIER_1 English(EN) · Nicholas Oh ·

    Complexity-Guided Component-wise Initialization for Language Model Pretraining

    Pretrained language models often exhibit structured weight spectra, suggesting that training may repeatedly produce similar layerwise and component-wise organization. We ask whether these recurring spectral patterns can be reused as an initialization signal for GPT-2-style langua…

  873. arXiv cs.LG TIER_1 English(EN) · Jiayi Fang ·

    Write-Protected Discrete Bottlenecks for Language-Grounded World Models: A Structural Limitation and Sufficient Fix

    arXiv:2607.08312v1 Announce Type: new Abstract: How should language interface with a world model's discrete symbol system? The dominant paradigm -- end-to-end injection of LLM/VLM features into robot world models (RT-2, Octo, PaLM-E) -- implicitly assumes that language gradients …

  874. arXiv cs.AI TIER_1 English(EN) · Theophilus Amaefuna, Hitesh Vaidya, Anshuman Chhabra, Ankur Mali ·

    Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization

    arXiv:2603.00910v2 Announce Type: replace-cross Abstract: Layer-wise capacity in large language models is highly non-uniform: some layers contribute disproportionately to loss reduction, whereas others are nearly redundant. Existing layer-scoring methods provide sensitivity estim…

  875. arXiv cs.AI TIER_1 English(EN) · Moritz Miller, Florent Draye, Bernhard Sch\"olkopf ·

    Towards Isolated Interventions via Almost Orthogonal Features in Language Models

    arXiv:2602.04718v2 Announce Type: replace-cross Abstract: A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to support reliable interventions, manipulating one…

  876. arXiv cs.AI TIER_1 English(EN) · Ryota Kobayashi, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Yasunori Ishii, Tomoyuki Okuno, Kazuki Kozuka ·

    Structured Pruning of Large Language Models via Power Transformation and Sign-Preserving Score Aggregation with Adaptive Feature Retention

    arXiv:2607.08027v1 Announce Type: cross Abstract: This paper proposes an improved structured pruning method for large language models (LLMs) that addresses key challenges in adapting Adaptive Feature Retention (AFR), an unstructured pruning technique, to structured pruning. When …

  877. arXiv cs.AI TIER_1 English(EN) · Lu Dai, Ziyang Rao, Yili Wang, Hanqing Wang, Hao Liu, Hui Xiong ·

    Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning

    arXiv:2607.08393v1 Announce Type: new Abstract: Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to use them for downstream reasoning tasks. We formalize this failure as the \textit{\textbf{Knowing--Using Gap}}, ch…

  878. arXiv cs.LG TIER_1 English(EN) · Teng-Ruei Chen ·

    Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models

    arXiv:2607.08665v1 Announce Type: new Abstract: Routing among large language models (LLMs) trades response quality against serving cost, motivated by the reported gap between deployed routers and a per-instance oracle. Recent analysis shows that test-time resampling can recover p…

  879. arXiv cs.LG TIER_1 English(EN) · Sebastian G. Gruber, Nassim Walha, Francis Bach, Florian Buettner ·

    Eigenvalue Calibration for Semantic Embeddings of Large Language Models

    arXiv:2607.08377v1 Announce Type: new Abstract: Uncertainty quantification is central to the reliable deployment of large language models (LLMs), and eigenvalues of semantic embeddings have recently emerged as a key tool in state-of-the-art methods. However, conventional calibrat…

  880. arXiv cs.CL TIER_1 English(EN) · Aiwei Liu, Cheng Shi, Chuhan Wu, Ci Lei, Di Lu, Donald He, Fan Zhang, Fanhao Kong, Feifei Zhang, Guan Wang, Haicheng Wang, Haoyu Liu, Houjin Yu, Jiachen Ding, Jiayi Feng, Jie Zhou, Jijun Chi, Jindi Shi, Jing Lei, Junjie Zhang, Laiyi Li, Le Tian, Linhao Z… ·

    Hidden Decoding at Scale: Latent Computation Scaling for Large Language Models

    arXiv:2607.08186v1 Announce Type: new Abstract: Scaling Large Language Models (LLMs) has been driven mainly by enlarging the Transformer backbone, but for an already-strong model this requires another round of costly pretraining. We study whether an existing backbone can keep imp…

  881. arXiv cs.CL TIER_1 English(EN) · Hoeseok Yang ·

    Sensitivity-Aware Thresholding and Token Routing for Activation Sparsification in Large Language Models

    Efficient inference in Large Language Models (LLMs) requires deciding where computation can be reduced while preserving model quality. We study this problem through multilayer perceptron (MLP) activation sparsification and token-level conditional routing. We first propose Sensiti…

  882. Hugging Face Daily Papers TIER_1 English(EN) ·

    Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models

    Routing among large language models (LLMs) trades response quality against serving cost, motivated by the reported gap between deployed routers and a per-instance oracle. Recent analysis shows that test-time resampling can recover per-instance selection headroom that no single-co…

  883. arXiv cs.LG TIER_1 English(EN) · Teng-Ruei Chen ·

    Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models

    Routing among large language models (LLMs) trades response quality against serving cost, motivated by the reported gap between deployed routers and a per-instance oracle. Recent analysis shows that test-time resampling can recover per-instance selection headroom that no single-co…

  884. Hugging Face Daily Papers TIER_1 English(EN) ·

    Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning

    Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to use them for downstream reasoning tasks. We formalize this failure as the \textbf{Knowing--Using Gap}, characterized by an accuracy gap and a temporal lag between…

  885. arXiv cs.AI TIER_1 English(EN) · Hui Xiong ·

    Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning

    Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to use them for downstream reasoning tasks. We formalize this failure as the \textit{\textbf{Knowing--Using Gap}}, characterized by an accuracy gap and a temporal la…

  886. Hugging Face Daily Papers TIER_1 English(EN) ·

    Eigenvalue Calibration for Semantic Embeddings of Large Language Models

    Uncertainty quantification is central to the reliable deployment of large language models (LLMs), and eigenvalues of semantic embeddings have recently emerged as a key tool in state-of-the-art methods. However, conventional calibration results developed for classification probabi…

  887. arXiv cs.LG TIER_1 English(EN) · Florian Buettner ·

    Eigenvalue Calibration for Semantic Embeddings of Large Language Models

    Uncertainty quantification is central to the reliable deployment of large language models (LLMs), and eigenvalues of semantic embeddings have recently emerged as a key tool in state-of-the-art methods. However, conventional calibration results developed for classification probabi…

  888. arXiv cs.LG TIER_1 English(EN) · Jiayi Fang ·

    Write-Protected Discrete Bottlenecks for Language-Grounded World Models: A Structural Limitation and Sufficient Fix

    How should language interface with a world model's discrete symbol system? The dominant paradigm -- end-to-end injection of LLM/VLM features into robot world models (RT-2, Octo, PaLM-E) -- implicitly assumes that language gradients can directly shape physical symbol representatio…

  889. arXiv cs.CL TIER_1 English(EN) · Zitao Wang ·

    Hidden Decoding at Scale: Latent Computation Scaling for Large Language Models

    Scaling Large Language Models (LLMs) has been driven mainly by enlarging the Transformer backbone, but for an already-strong model this requires another round of costly pretraining. We study whether an existing backbone can keep improving by allocating more computation to each to…

  890. Hugging Face Daily Papers TIER_1 English(EN) ·

    Hidden Decoding at Scale: Latent Computation Scaling for Large Language Models

    Scaling Large Language Models (LLMs) has been driven mainly by enlarging the Transformer backbone, but for an already-strong model this requires another round of costly pretraining. We study whether an existing backbone can keep improving by allocating more computation to each to…

  891. arXiv cs.AI TIER_1 English(EN) · Yiming Gai, Junde Lu, Xuefei Huang ·

    Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System

    arXiv:2607.06940v1 Announce Type: cross Abstract: The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their response quality. Prevailing methods, often confined to singular dimensions, fall shor…

  892. arXiv cs.AI TIER_1 English(EN) · Sahil Kale ·

    Future Confidence Distillation in Large Language Models

    arXiv:2607.07626v1 Announce Type: cross Abstract: Reliable confidence estimation is essential for deploying large language models (LLMs) in confidence-aware systems, where downstream decisions such as retrieval, tool use, and adaptive computation depend on accurately estimating a…

  893. arXiv cs.AI TIER_1 English(EN) · Yair Feldman, Linxi Zhao, Nathan Godey, Dongyoung Go, Yilun Hua, Kilian Q. Weinberger, Jennifer J. Sun, Yoav Artzi ·

    Co-LMLM: Continuous-Query Limited Memory Language Models

    arXiv:2607.07707v1 Announce Type: cross Abstract: Limited memory language models (LMLMs) externalize factual knowledge during pretraining to a knowledge base (KB), rather than memorizing it in their weights. During generation, the model then fetches knowledge from the KB as neede…

  894. arXiv cs.AI TIER_1 English(EN) · Kaijian Zou, Aaron Xiong, Yunxiang Zhang, Frederick Zhang, Yueqi Ren, Jirong Yang, Ayoung Lee, Shitanshu Bhushan, Lu Wang ·

    LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?

    arXiv:2510.09595v3 Announce Type: replace Abstract: Competitive programming problems are increasingly used to evaluate the coding capabilities of large language models (LLMs) due to their complexity and ease of verification. Yet, current coding benchmarks face limitations such as…

  895. arXiv cs.CL TIER_1 English(EN) · Yutang Ma, Kecheng Huang, Xikun Jiang, Zili Shao ·

    TF-Engram: A Train-Free Engram with SSD-Backed Memory for Large Language Models

    arXiv:2607.07388v1 Announce Type: new Abstract: Large Language Models (LLMs) store factual knowledge and domain-specific patterns implicitly in dense Transformer parameters, making knowledge expansion costly through pretraining, fine-tuning, retrieval augmentation, or longer cont…

  896. arXiv cs.LG TIER_1 English(EN) · Maximilian S. Ernst (Max Planck School of Cognition, Center for Lifespan Psychology Max Planck Institute for Human Development, Machine Learning Group Technische Universit\"at Berlin), Lorenz Linhardt (Machine Learning Group Technische Universit\"at Berl… ·

    Distributed Sparse Interventions in Language Models

    arXiv:2607.07128v1 Announce Type: new Abstract: Language models perform a wide range of tasks at varying levels of abstraction with the capacity to flexibly infer tasks from context, execute multiple tasks simultaneously, and select among competing tasks. To study the role of mod…

  897. arXiv cs.CL TIER_1 English(EN) · Kazuki Kozuka ·

    Structured Pruning of Large Language Models via Power Transformation and Sign-Preserving Score Aggregation with Adaptive Feature Retention

    This paper proposes an improved structured pruning method for large language models (LLMs) that addresses key challenges in adapting Adaptive Feature Retention (AFR), an unstructured pruning technique, to structured pruning. When applying AFR to structured pruning, three major pr…

  898. Hugging Face Daily Papers TIER_1 English(EN) ·

    Structured Pruning of Large Language Models via Power Transformation and Sign-Preserving Score Aggregation with Adaptive Feature Retention

    This paper proposes an improved structured pruning method for large language models (LLMs) that addresses key challenges in adapting Adaptive Feature Retention (AFR), an unstructured pruning technique, to structured pruning. When applying AFR to structured pruning, three major pr…

  899. arXiv cs.AI TIER_1 English(EN) · Yoav Artzi ·

    Co-LMLM: Continuous-Query Limited Memory Language Models

    Limited memory language models (LMLMs) externalize factual knowledge during pretraining to a knowledge base (KB), rather than memorizing it in their weights. During generation, the model then fetches knowledge from the KB as needed. This recently introduced paradigm provides mult…

  900. arXiv cs.AI TIER_1 English(EN) · Sahil Kale ·

    Future Confidence Distillation in Large Language Models

    Reliable confidence estimation is essential for deploying large language models (LLMs) in confidence-aware systems, where downstream decisions such as retrieval, tool use, and adaptive computation depend on accurately estimating answer reliability. Existing approaches, however, l…

  901. arXiv cs.CL TIER_1 English(EN) · Zili Shao ·

    TF-Engram: A Train-Free Engram with SSD-Backed Memory for Large Language Models

    Large Language Models (LLMs) store factual knowledge and domain-specific patterns implicitly in dense Transformer parameters, making knowledge expansion costly through pretraining, fine-tuning, retrieval augmentation, or longer contexts. Engram-style memory offers a compact hidde…

  902. arXiv cs.LG TIER_1 English(EN) · Oliver Eberle ·

    Distributed Sparse Interventions in Language Models

    Language models perform a wide range of tasks at varying levels of abstraction with the capacity to flexibly infer tasks from context, execute multiple tasks simultaneously, and select among competing tasks. To study the role of model components in task behaviour, their causal in…

  903. arXiv cs.AI TIER_1 English(EN) · So Hasegawa, Shailaja Keyur Sampat, Lei Liu, Wei-Peng Chen ·

    Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

    arXiv:2607.06482v1 Announce Type: cross Abstract: Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically focus on fact retrieval from small tables and overlook the challenges of large multi-tabular…

  904. arXiv cs.AI TIER_1 English(EN) · H. Chad Lane, Bryson Kageler ·

    CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming

    arXiv:2607.05571v1 Announce Type: new Abstract: Large language models are increasingly explored as AI tutors, yet deploying them in K-12 settings raises concerns around privacy, cost, and reliance on proprietary models. Small language models (SLMs) offer a promising alternative, …

  905. arXiv cs.AI TIER_1 English(EN) · Hongliang Li, Yijin Liu, Zhiwei Zhang, Zihe Liu, Xinyue Lou, Jinan Xu, Fandong Meng, Kaiyu Huang ·

    PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents

    arXiv:2607.06008v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external environments. However, most existing benchmarks implicitly assume a monolingual set…

  906. arXiv cs.AI TIER_1 English(EN) · Lars Benedikt Kaesberg ·

    Decision Protocols in Multi-Agent Large Language Model Conversations

    arXiv:2607.05477v1 Announce Type: cross Abstract: Improving the task performance of Large Language Models (LLMs) is essential, yet scaling these models faces significant challenges such as diminishing returns and high costs. Multi-Agent Systems (MAS) offer a promising solution by…

  907. arXiv cs.AI TIER_1 English(EN) · Archie Chaudhury ·

    ResonatorLM: Causal Resonant Field Mixing for Efficient Long-Context Language Modelin

    arXiv:2607.05583v1 Announce Type: cross Abstract: Contemporary language models are dominated by the transformer architecture, which leverages self-attention mechanisms to enable more efficient, parallelized training across a wide set of documents and corpora. This has allowed tra…

  908. arXiv cs.AI TIER_1 English(EN) · Damian Hodel, Jevin West, Aylin Caliskan ·

    RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream Outputs

    arXiv:2607.05679v1 Announce Type: cross Abstract: Language models (LMs) exhibit problematic biases, such as stereotypes. Effectively analyzing and mitigating such biases requires accurate and generalizable evaluation methods of the underlying associations. Some existing approache…

  909. arXiv cs.AI TIER_1 English(EN) · Jianyi Zhang, Hao Frank Yang, Ang Li, Xin Guo, Pu Wang, Haiming Wang, Yiran Chen, Hai Li ·

    MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning

    arXiv:2409.06067v3 Announce Type: replace Abstract: Previous studies on federated learning (FL) often encounter performance degradation due to data heterogeneity among different clients. In light of the recent advances in multimodal large language models (MLLMs), such as GPT-4v a…

  910. arXiv cs.AI TIER_1 English(EN) · Marthe Ballon, Andres Algaba, Vincent Ginis ·

    The relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer

    arXiv:2502.15631v2 Announce Type: replace-cross Abstract: Large language models have demonstrated remarkable progress in mathematical reasoning, leveraging chain-of-thought and reinforcement learning. However, many open questions remain regarding the interplay between reasoning t…

  911. arXiv cs.CL TIER_1 English(EN) · Yimeng Zhang, Yingying Zhuang, Ziyi Wang, Yuxuan Lu, Pei Chen, Aman Gupta, Zhe Su, Ming Tan, Zhilin Zhang, Qun Liu, Manikandarajan Ramanathan, Rajashekar Maragoud, Edward Vul, Jing Huang, Dakuo Wang ·

    SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation

    arXiv:2607.05721v1 Announce Type: new Abstract: Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refinement in LLM generation. However, existing approaches operate at suboptimal granulari…

  912. arXiv cs.CL TIER_1 Dansk(DA) · Ruiyi Yan, Zhuoyuan Mao, Yiwen Guo ·

    MemDefrag: Latent Memory Defragmentation for Large Language Models

    arXiv:2607.05969v1 Announce Type: new Abstract: Latent memory, which stores past knowledge fragments as per-layer hidden states, has emerged as a promising paradigm (e.g., MemoryLLM and M+) for long-term memory in large language models (LLMs). However, the paradigm suffers from s…

  913. arXiv cs.CL TIER_1 English(EN) · Nikola Zubi\'c, Qian Li, Yuyi Wang, Davide Scaramuzza ·

    When Does Tool Use Increase the Expressive Power of Finite-Precision Recurrent Models?

    arXiv:2607.06155v1 Announce Type: cross Abstract: Modern sequence models are increasingly deployed as agents that interleave token generation with calls to external tools. We give an exact, architecture-level account of when such tool access increases computational expressivity. …

  914. arXiv cs.CL TIER_1 English(EN) · Dmitrii Popov, Egor Terentev, Danil Serenko, Ilya Sochenkov, Igor Buyanov ·

    Transferring Natural Language Datasets Between Languages Using Large Language Models for Modern Decision Support and Sci-Tech Analytical Systems

    arXiv:2410.14074v2 Announce Type: replace Abstract: The decision-making process to rule R&amp;D relies on information related to current trends in particular research areas. In this work, we investigated how one can use large language models (LLMs) to transfer the dataset and its…

  915. arXiv cs.CL TIER_1 English(EN) · Lester James V. Miranda, Ivan Vuli\'c, Anna Korhonen ·

    Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation

    arXiv:2604.11290v2 Announce Type: replace Abstract: Synthesizing supervised finetuning (SFT) data from language models (LMs) to teach smaller models multilingual tasks has become increasingly common. However, teacher model selection is often ad hoc, typically defaulting to the la…

  916. arXiv cs.LG TIER_1 English(EN) · Nima Eshraghi, Lovedeep Gondara, Yuqing Huang, Sagarika Suresh, Leizer Teran, Jithin Pradeep, Xiaotong Xu, Fanny Chevalier ·

    A Coin Flip Per Token: Bernoulli Sparse Steering of Large Language Models

    arXiv:2607.05615v1 Announce Type: new Abstract: Activation steering via sparse autoencoders (SAEs) enables behavioral control of large language models without task-specific fine-tuning, but standard methods apply the steering signal at every generated token, incurring constant pe…

  917. arXiv cs.CL TIER_1 English(EN) · Xuefei Huang ·

    Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System

    The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their response quality. Prevailing methods, often confined to singular dimensions, fall short of capturing the full spectrum of model capabili…

  918. arXiv cs.AI TIER_1 English(EN) · Wei-Peng Chen ·

    Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

    Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically focus on fact retrieval from small tables and overlook the challenges of large multi-tabular datasets, external knowledge integration, and exp…

  919. arXiv cs.CL TIER_1 English(EN) · Davide Scaramuzza ·

    When Does Tool Use Increase the Expressive Power of Finite-Precision Recurrent Models?

    Modern sequence models are increasingly deployed as agents that interleave token generation with calls to external tools. We give an exact, architecture-level account of when such tool access increases computational expressivity. We model any fixed finite-precision recurrent sequ…

  920. arXiv cs.AI TIER_1 English(EN) · Kaiyu Huang ·

    PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents

    Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external environments. However, most existing benchmarks implicitly assume a monolingual setting, where the entire execution process, includ…

  921. arXiv cs.CL TIER_1 Dansk(DA) · Yiwen Guo ·

    MemDefrag: Latent Memory Defragmentation for Large Language Models

    Latent memory, which stores past knowledge fragments as per-layer hidden states, has emerged as a promising paradigm (e.g., MemoryLLM and M+) for long-term memory in large language models (LLMs). However, the paradigm suffers from significant performance degradation during memory…

  922. arXiv cs.AI TIER_1 English(EN) · Jiang Zhang, Bing Yuan, Qian Zhang ·

    Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement

    arXiv:2607.04277v1 Announce Type: cross Abstract: The pursuit of self-evolving AI raises a critical question: when is autonomous self-improvement sustainable rather than degenerative? Drawing an analogy to von Neumann's complexity threshold for self-reproducing automata, we argue…

  923. arXiv cs.AI TIER_1 English(EN) · Garry Yang, Zizhe Chen, Xinru Chen, Yongqiang Chen, Jianxiang Wang, Deyu Zou, Linyi Ding, Jialiang Wu, Yunzhong He, Yu Gong, James Cheng, Huaixiao Tou ·

    APeB: Benchmarking Personalization Ability of Large Language Model Agents

    arXiv:2607.03162v1 Announce Type: new Abstract: LLM-powered agents struggle with personalization when users issue raw, underspecified queries. In this setting, agents must infer latent intent, extract preferences from noisy interaction histories, and select among competing altern…

  924. arXiv cs.AI TIER_1 English(EN) · Damir Shodiev, Aleksei Staroverov, Nikita Kachaev, Alexey K. Kovalev, Aleksandr I. Panov ·

    VLA Grounder: Language-Conditioning Space Optimization for Black-Box VLA Models

    arXiv:2607.04517v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are commonly treated as end-to-end action policies conditioned on natural-language task descriptions. In practice, however, their behavior often depends sharply on how the instruction is phrased, …

  925. arXiv cs.AI TIER_1 English(EN) · Zhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang, Zhonghang Lu, Chenyu Liu, Jiajun Sun, Jiazheng Zhang, Dingwei Zhu, Xin Guo, Junzhe Wang, Zhihao Zhang, Yuming Yang, Junjie Ye, Minghe Gao, Dongrui Liu,… ·

    AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments

    arXiv:2607.05174v1 Announce Type: new Abstract: Language agents, i.e., LLM agents, progress rapidly and are increasingly deployed in production environments. This trend underscores the urgent need for rigorous and realistic evaluations. However, most existing benchmarks evaluate …

  926. arXiv cs.AI TIER_1 English(EN) · Zhimin Hu, Lanhao Niu, Sashank Varma ·

    Language Models Represent and Transform Concepts with Shared Geometry

    arXiv:2607.04525v1 Announce Type: cross Abstract: How concepts are represented in neural networks is a fundamental question in machine learning. The dominant view treats concept representations as stationary geometric objects. Yet concepts appear in context, and context transform…

  927. arXiv cs.CL TIER_1 English(EN) · Tatsuya Aoyama, Ethan Gotlieb Wilcox, Nathan Schneider ·

    Predicting the Emergence of Induction Heads in Language Model Pretraining

    arXiv:2511.16893v3 Announce Type: replace Abstract: Specialized attention heads dubbed induction heads (IHs) have been argued to underlie the remarkable in-context learning capabilities of modern language models; yet, a precise characterization of their emergence, especially in t…

  928. arXiv cs.CL TIER_1 English(EN) · Irina Proskurina, Guillaume Metzler, Julien Velcin ·

    Fair-GPTQ: Bias-Aware Quantization for Large Language Models

    arXiv:2509.15206v3 Announce Type: replace Abstract: The high memory demands of generative language models have drawn attention to quantization, which reduces memory usage by mapping model weights to lower-precision integers. However, recent empirical studies show that, while effi…

  929. arXiv cs.CL TIER_1 English(EN) · Karanpartap Singh, Neil Band, Ehsan Adeli ·

    Curriculum-Guided Layer Scaling for Language Model Pretraining

    arXiv:2506.11389v4 Announce Type: replace Abstract: As the cost of pretraining large language models grows, there is continued interest in strategies to improve learning efficiency during this core training stage. Motivated by cognitive development, where humans gradually build k…

  930. arXiv cs.CL TIER_1 English(EN) · Hamish Ogilvy ·

    Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models

    arXiv:2607.02893v1 Announce Type: cross Abstract: Low-bit quantization shrinks language models but treats precision as a single global hyper-parameter: every weight uses the same bit-width. We introduce Variable Bit-width Quantization (VBQ), a training-time method in which each c…

  931. arXiv cs.AI TIER_1 English(EN) · Caspar Kaiser, Sean Enderby ·

    No Reliable Evidence of Self-Reported Sentience in Small Large Language Models

    arXiv:2601.15334v2 Announce Type: replace-cross Abstract: Whether language models possess sentience has no empirical answer. But whether they believe themselves to be sentient can, in principle, be tested. We do so by querying several open-weights models about their own conscious…

  932. arXiv cs.AI TIER_1 English(EN) · Addison J. Wu, Ryan Liu, Xuechunzi Bai, Thomas L. Griffiths ·

    Large Language Models Develop Novel Social Biases Through Adaptive Exploration

    arXiv:2511.06148v4 Announce Type: replace-cross Abstract: As large language models (LLMs) are adopted into frameworks that grant them the capacity to make real decisions, it is increasingly important to ensure that they are unbiased. In this paper, we argue that the predominant a…

  933. arXiv cs.AI TIER_1 English(EN) · Yisong Zhang, Ran Cheng, Guoxing Yi, Kay Chen Tan ·

    A Systematic Survey on Large Language Models for Evolutionary Optimization: From Modeling to Solving

    arXiv:2509.08269v5 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly integrated with evolutionary computation to support optimization tasks. This survey primarily focuses on evolutionary optimization, i.e., optimization based on evolutionary com…

  934. arXiv cs.AI TIER_1 English(EN) · Yifan Luo, Zhennan Zhou, Bin Dong ·

    InverseScope: Scalable Activation Inversion for Interpreting Large Language Models

    arXiv:2506.07406v3 Announce Type: replace-cross Abstract: Understanding the internal representations of large language models (LLMs) is a central challenge in interpretability research. Existing feature interpretability methods often rely on strong structural assumptions--such as…

  935. arXiv cs.AI TIER_1 English(EN) · Yang Yang, Hua XU, Zhangyi Hu, Yutao Yue ·

    RLIE: Rule Generation with Logistic Regression, Iterative Refinement, and Evaluation for Large Language Models

    arXiv:2510.19698v3 Announce Type: replace Abstract: Large Language Models (LLMs) can propose rules in natural language, sidestepping the need for a predefined predicate space in traditional rule learning. Yet many LLM-based approaches ignore interactions among rules, and the oppo…

  936. arXiv cs.AI TIER_1 English(EN) · Zhuoying Zhang, Ishan V. Prasad, Yuanzhe Hu, Zihang Liu, Hengrui Luo, Pu Ren, Yaoqing Yang ·

    Spectral Signatures of Large Language Models

    arXiv:2607.03377v1 Announce Type: cross Abstract: The rapidly growing repository of publicly available large language models (LLMs) presents significant challenges for systematic management and quantification at scale, such as model lineage tracing, licensing, and evaluation. How…

  937. arXiv cs.AI TIER_1 English(EN) · Bhavesh Sood, Jaromir Savelka ·

    Punching Above Their Weight: Classification-Head Fine-Tuning of Tiny Language Models (TLMs) for Verifiable Multiple-Choice Tasks

    arXiv:2607.03801v1 Announce Type: cross Abstract: We define Tiny Language Models (TLMs) as models below roughly 3B parameters that fit on mainstream consumer devices. We study how to adapt them for and use them on verifiable multiple-choice tasks. We compare three LoRA-based fine…

  938. arXiv cs.AI TIER_1 English(EN) · Shreyas K Chandrahas ·

    evalci: A Python Library for Statistically Rigorous Comparison of Language Model Evaluations

    arXiv:2607.04429v1 Announce Type: cross Abstract: The dominant practice in language model evaluation is to report a single accuracy number per model and declare the higher one better, without testing whether the gap could plausibly be sampling noise. On benchmarks of a few thousa…

  939. arXiv cs.LG TIER_1 English(EN) · Minghao Yan, Zhuang Wang, Zhen Jia, Shivaram Venkataraman, Yida Wang ·

    PLoRA: Efficient Concurrent LoRA Training for Large Language Models

    arXiv:2508.02932v2 Announce Type: replace Abstract: Low-Rank Adaptation (LoRA) has gained popularity as a fine-tuning approach for Large Language Models (LLMs) due to its low resource requirements and good performance. While numerous studies have investigated ways to improve LoRA…

  940. arXiv cs.LG TIER_1 English(EN) · Saksham Bassi, Sharvi Tomar ·

    Geometry of Ordinal Representations in Language Models

    arXiv:2607.04167v1 Announce Type: new Abstract: Recent work showed that language models represent character counts on curved 1D manifolds, with attention heads performing geometric transformations to enable computation. We test whether this generalizes across four ordinal tasks (…

  941. arXiv cs.CL TIER_1 English(EN) · Jiliang Hu, Wenfu Wang, Zuchao Li, Chenxing Li, Yiyang Zhao, Hanzhao Li, Liqiang Zhang, Meng Yu, Dong Yu ·

    VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents

    arXiv:2510.11098v5 Announce Type: replace-cross Abstract: Recent advances in large audio language models (LALMs) have greatly enhanced multimodal conversational systems. However, existing benchmarks remain limited -- they are mainly English-centric, rely on synthetic speech, and …

  942. arXiv cs.CL TIER_1 English(EN) · G\"ul Sena Alt{\i}nta\c{s}, Malikeh Ehghaghi, Brian Lester, Fengyuan Liu, Wanru Zhao, Marco Ciccone, Colin Raffel ·

    TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior

    arXiv:2512.20757v2 Announce Type: replace Abstract: Tokenizers provide the fundamental basis through which text is represented and processed by language models (LMs). Despite the importance of tokenization, its role in LM performance and behavior is poorly understood due to the c…

  943. Hugging Face Daily Papers TIER_1 English(EN) ·

    SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation

    Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refinement in LLM generation. However, existing approaches operate at suboptimal granularities: token-level scores lack semantic coherence…

  944. arXiv cs.CL TIER_1 English(EN) · Dakuo Wang ·

    SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation

    Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refinement in LLM generation. However, existing approaches operate at suboptimal granularities: token-level scores lack semantic coherence…

  945. arXiv cs.CL TIER_1 English(EN) · Aylin Caliskan ·

    RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream Outputs

    Language models (LMs) exhibit problematic biases, such as stereotypes. Effectively analyzing and mitigating such biases requires accurate and generalizable evaluation methods of the underlying associations. Some existing approaches focus on downstream metrics that analyze associa…

  946. arXiv cs.CL TIER_1 English(EN) · Archie Chaudhury ·

    ResonatorLM: Causal Resonant Field Mixing for Efficient Long-Context Language Modelin

    Contemporary language models are dominated by the transformer architecture, which leverages self-attention mechanisms to enable more efficient, parallelized training across a wide set of documents and corpora. This has allowed transformers to effectively model data across a wide …

  947. arXiv cs.AI TIER_1 English(EN) · Xuanjing Huang ·

    AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments

    Language agents, i.e., LLM agents, progress rapidly and are increasingly deployed in production environments. This trend underscores the urgent need for rigorous and realistic evaluations. However, most existing benchmarks evaluate agents in simplified, idealized settings. They t…

  948. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Lars Benedikt Kaesberg ·

    Decision Protocols in Multi-Agent Large Language Model Conversations

    Improving the task performance of Large Language Models (LLMs) is essential, yet scaling these models faces significant challenges such as diminishing returns and high costs. Multi-Agent Systems (MAS) offer a promising solution by distributing tasks among specialized agents to im…

  949. arXiv cs.CL TIER_1 English(EN) · Sashank Varma ·

    Language Models Represent and Transform Concepts with Shared Geometry

    How concepts are represented in neural networks is a fundamental question in machine learning. The dominant view treats concept representations as stationary geometric objects. Yet concepts appear in context, and context transforms them. Drawing from neural population geometry, w…

  950. Hugging Face Daily Papers TIER_1 English(EN) ·

    Language Models Represent and Transform Concepts with Shared Geometry

    How concepts are represented in neural networks is a fundamental question in machine learning. The dominant view treats concept representations as stationary geometric objects. Yet concepts appear in context, and context transforms them. Drawing from neural population geometry, w…

  951. Hugging Face Daily Papers TIER_1 English(EN) ·

    VLA Grounder: Language-Conditioning Space Optimization for Black-Box VLA Models

    Vision-Language-Action (VLA) models are commonly treated as end-to-end action policies conditioned on natural-language task descriptions. In practice, however, their behavior often depends sharply on how the instruction is phrased, suggesting that language is not merely a task la…

  952. arXiv cs.CL TIER_1 English(EN) · Shreyas K Chandrahas ·

    evalci: A Python Library for Statistically Rigorous Comparison of Language Model Evaluations

    The dominant practice in language model evaluation is to report a single accuracy number per model and declare the higher one better, without testing whether the gap could plausibly be sampling noise. On benchmarks of a few thousand items, and under temperature sampling where a m…

  953. Hugging Face Daily Papers TIER_1 English(EN) ·

    Geometry of Ordinal Representations in Language Models

    Recent work showed that language models represent character counts on curved 1D manifolds, with attention heads performing geometric transformations to enable computation. We test whether this generalizes across four ordinal tasks (bracket depth, indentation, table position, nume…

  954. arXiv cs.AI TIER_1 English(EN) · Samuel Schapiro, Core Francisco Park, Felix Sosa, Lav R. Varshney ·

    CreativityNeuro: Steering Language Model Weights to Improve Divergent Thinking and Reduce Mode Collapse

    arXiv:2607.01433v1 Announce Type: new Abstract: Divergent thinking is a crucial aspect of creativity, yet large language models (LLMs) tend to consistently generate similar responses to open-ended questions, in what has been termed the artificial hivemind effect. Here, we introdu…

  955. arXiv cs.LG TIER_1 English(EN) · Anna Karnysheva, Dietrich Klakow, Ji-Ung Lee ·

    Probing Chemical Language Models: Effects of Pre-training and Fine-tuning

    arXiv:2607.02140v1 Announce Type: new Abstract: Chemical language models (CLMs) are trained with linearized representations such as SMILES, yet it remains unclear which chemically meaningful substructures they encode. To foster a better understanding of CLMs, we conduct a systema…

  956. arXiv cs.CL TIER_1 English(EN) · Junyu Lin, Meizhen Liu, Xiufeng Huang, Jinfeng Li, Haiwen Hong, Xiaohan Yuan, Yuefeng Chen, Longtao Huang, Hui Xue, Ranjie Duan, Zhikai Chen, Yuchuan Fu, Defeng Li, Lingyao Gao, Yitong Yang ·

    YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models

    arXiv:2601.15588v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in real-world applications, safety guardrails are required to go beyond coarse-grained filtering and support fine-grained, interpretable, and adaptable risk assessment. H…

  957. arXiv cs.AI TIER_1 English(EN) · Shengsheng Zhou, Shuai Wang, Liang Ding, Yibing Zhan, Yong Luo, Zheng He, Fu Lin, Dapeng Tao ·

    Leveraging Metamemory Agent for Enhanced Data-Free Code Generation in Large Language Models

    arXiv:2501.07892v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown strong performance in automated code generation, with few-shot prompting widely used for its simplicity and effectiveness. However, few-shot methods depend on curated or manually cra…

  958. arXiv cs.AI TIER_1 English(EN) · Amirreza Esmaeili, Fatemeh Fard ·

    TokenScope: Token-Level Explainability and Interpretability for Code-Oriented Tasks in Large Language Models

    arXiv:2607.01235v1 Announce Type: cross Abstract: Understanding how Large Language Models (LLMs) make token-level decisions during code generation remains a major challenge for both researchers and practitioners. While recent tools provide insights into model internals or generat…

  959. arXiv cs.AI TIER_1 English(EN) · Xi Fang, Weijie Xu, Yingqiang Ge, Yuhui Xu, Stephanie Eckman, Chandan K. Reddy ·

    DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models

    arXiv:2607.02374v1 Announce Type: new Abstract: Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response. Modern LLMs personalize interactions by storing user attributes, preferences, and prior cont…

  960. arXiv cs.AI TIER_1 English(EN) · Ye Liu, Srijan Bansal, Bo Pang, Yang Li, Zeyu Leo Liu, Yifei Ming, Zixuan Ke, Shafiq Joty, Semih Yavuz ·

    Procedural Memory Distillation: Online Reflection for Self-Improving Language Models

    arXiv:2607.01480v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR), along with recent selfdistillation variants such as SDPO, evaluates each rollout against a verifier and updates the policy from that episode-level signal. However, the richer pr…

  961. arXiv cs.AI TIER_1 English(EN) · Aryuemaan Kumar Chowdhury, Afreen Shaik, Yaparla Bhargavi, Brahma Kumar ·

    The Wiola Architecture for Efficient Small Language Models

    arXiv:2607.01394v1 Announce Type: new Abstract: We present Wiola, a fully original Small Language Model (SLM) architecture built from first principles, sharing no structural lineage with any existing model family including GPT, LLaMA, Mistral, or Falcon. Wiola introduces five ind…

  962. arXiv cs.CL TIER_1 English(EN) · Yuchen Lu, Run Yang, Yichen Zhang, Shuguang Yu, Ziwei Wang, Jiayi Xiang, Wenxin E, Changyu Zhu, Fan Zhou ·

    StatEval: A Comprehensive Benchmark for Large Language Models in Statistics

    arXiv:2510.09517v2 Announce Type: replace Abstract: Despite rapid advances in large language models (LLMs), statistical reasoning remains underrepresented in existing LLM benchmarks, which often do not reflect the layered, proof-driven nature of real statistical practice. To addr…

  963. arXiv cs.CL TIER_1 English(EN) · Dingkang Lin, Naixuan Zhao, Dan Tian, Jiang Li ·

    Large language models reshape the language of science

    arXiv:2504.12317v2 Announce Type: replace Abstract: Scientific language is a central infrastructure of knowledge production, but it remains unclear whether large language models (LLMs) are altering not only how scientists write, but also how scientific knowledge is communicated a…

  964. arXiv cs.CL TIER_1 English(EN) · Weibin Liao, Xin Gao, Tianyu Jia, Rihong Qiu, Yifan Zhu, Yang Lin, Xinyu Ma, Junfeng Zhao, Yasha Wang ·

    LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language Models

    arXiv:2504.02327v2 Announce Type: replace Abstract: Natural Language to SQL (NL2SQL) aims to translate natural language queries into executable SQL statements, offering non-expert users intuitive access to databases. While recent approaches leveraging large-scale private LLMs suc…

  965. arXiv cs.CL TIER_1 English(EN) · Jijie Zhang, Zhe Ren, Quan Zhang, Dandan Guo ·

    Bayesian Sparse Low-Rank Adaptation for Large Language Model Uncertainty Estimation

    arXiv:2607.02182v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit remarkable reasoning capabilities, but their task-specific fine-tuning is notoriously plagued by overconfidence, severely hindering trustworthy deployment. We propose Data-Adaptive Lower-Rank A…

  966. arXiv cs.CL TIER_1 English(EN) · Congrui Du, Yang Zhang, Kaizhi Qian, Shiyu Chang ·

    Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

    arXiv:2607.02214v1 Announce Type: new Abstract: Instruction tuning for speech language models (SLMs) is substantially more challenging than for text-based large language models (LLMs), as it requires learning a new modality and a wide range of speech-specific instructions in addi…

  967. arXiv cs.CL TIER_1 English(EN) · Kim Gerdes (LISN, Qatent, STL) ·

    The Grammar Does the Work: Functional vs. Lexical Dependency Length Minimization Across Universal Dependencies

    arXiv:2607.01899v1 Announce Type: new Abstract: Dependency length minimization (DLM) is a well-documented processing universal, but previous studies report a single mean dependency distance (MDD) per language, obscuring variation across syntactic relation types. We analyze 122 la…

  968. arXiv cs.CL TIER_1 English(EN) · Siddharth Gollapudi, Nilesh Gupta, Prasann Singhal, Sewon Min ·

    Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale

    arXiv:2607.01538v1 Announce Type: new Abstract: Language models (LMs) raise an intriguing alternative to vector-based retrieval: conditioning on an in-context corpus and directly generating a relevant answer. However, prior work has largely focused on proprietary systems or the s…

  969. arXiv cs.AI TIER_1 English(EN) · Long Lian, Sida Wang, Felix Juefei-Xu, Tsu-Jui Fu, Xiuyu Li, Adam Yala, Trevor Darrell, Alane Suhr, Yuandong Tian, Xi Victoria Lin ·

    ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models

    arXiv:2512.07843v2 Announce Type: replace-cross Abstract: Scaling inference-time computation has enabled Large Language Models (LLMs) to achieve strong reasoning performance, but their inherently sequential decoding incurs substantial latency, motivating parallelization of the ge…

  970. arXiv cs.AI TIER_1 English(EN) · Chandan K. Reddy ·

    DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models

    Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response. Modern LLMs personalize interactions by storing user attributes, preferences, and prior context, then injecting this information into future…

  971. arXiv cs.CL TIER_1 English(EN) · Shiyu Chang ·

    Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

    Instruction tuning for speech language models (SLMs) is substantially more challenging than for text-based large language models (LLMs), as it requires learning a new modality and a wide range of speech-specific instructions in addition to those supported by text LLMs. Existing S…

  972. arXiv cs.CL TIER_1 English(EN) · Dandan Guo ·

    Bayesian Sparse Low-Rank Adaptation for Large Language Model Uncertainty Estimation

    Large language models (LLMs) exhibit remarkable reasoning capabilities, but their task-specific fine-tuning is notoriously plagued by overconfidence, severely hindering trustworthy deployment. We propose Data-Adaptive Lower-Rank Adaptation (DALorRA), a simple and effective variat…

  973. arXiv cs.LG TIER_1 English(EN) · Ji-Ung Lee ·

    Probing Chemical Language Models: Effects of Pre-training and Fine-tuning

    Chemical language models (CLMs) are trained with linearized representations such as SMILES, yet it remains unclear which chemically meaningful substructures they encode. To foster a better understanding of CLMs, we conduct a systematic study and probe for 78 molecular substructur…

  974. arXiv cs.CL TIER_1 English(EN) · Kim Gerdes ·

    The Grammar Does the Work: Functional vs. Lexical Dependency Length Minimization Across Universal Dependencies

    Dependency length minimization (DLM) is a well-documented processing universal, but previous studies report a single mean dependency distance (MDD) per language, obscuring variation across syntactic relation types. We analyze 122 languages in UD and SUD (version 2.17), showing th…

  975. arXiv cs.AI TIER_1 English(EN) · Michael S. Zhang, Rishi A. Ruia, Arnav Kewalram, Saathvik Dharmapuram, Utkarsh Sharma, Kevin Zhu ·

    When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models

    arXiv:2512.18934v2 Announce Type: replace-cross Abstract: Catastrophic forgetting poses a fundamental challenge in continual learning, particularly when models are quantized for deployment efficiency. We systematically investigate the interplay between quantization precision (FP1…

  976. arXiv cs.AI TIER_1 English(EN) · Arya Raeesi, Hanna Roed ·

    Auditing Forgetting in Limited Memory Language Models

    arXiv:2607.00605v1 Announce Type: cross Abstract: Limited Memory Language Models (LMLMs) externalize factual knowledge to a database to enable deletion-based unlearning without retraining. Existing evaluations measure post-deletion correctness in aggregate and cannot tell whether…

  977. arXiv cs.AI TIER_1 (CA) · Zinan Tang, Yukun Zhang, Shaomian Zheng, Zhuoshi Pan, Qizhi Pei, Dingnan Jin, Jun Zhou, Yujun Wang, Biqing Huang ·

    CausalMix: Data Mixture as Causal Inference for Language Model Training

    arXiv:2607.01104v1 Announce Type: cross Abstract: In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture weights via proxy models, but they rely on the assumption of static data distributions. As …

  978. arXiv cs.LG TIER_1 English(EN) · Sajjad Ghiasvand, Mark Beliaev, Mahnoosh Alizadeh, Ramtin Pedarsani ·

    REALM: Reliable Expertise-Aware Language Model Fine-Tuning from Noisy Annotations

    arXiv:2604.17289v2 Announce Type: replace Abstract: Supervised fine-tuning of large language models relies on human-annotated data, yet annotation pipelines routinely involve multiple crowdworkers of heterogeneous expertise. Standard practice aggregates labels via majority vote o…

  979. arXiv cs.CL TIER_1 English(EN) · Jeremias Bohn, Tizian Dippold, Mahdi Koubaa, Elias R. Wahl, Georg Groh ·

    $\text{Log}_\text{b}$Quant: Quantizing Language Models in Logarithmic Space

    arXiv:2607.01127v1 Announce Type: new Abstract: Quantization has become an invaluable tool to reduce memory requirements and inference speed of modern language models, in particular to make them available for consumer setups and edge devices. While previous work has primarily foc…

  980. arXiv cs.CL TIER_1 Deutsch(DE) · Yannik Keller, Thomas Eisenmann ·

    Understanding Large Language Models

    arXiv:2607.01006v1 Announce Type: new Abstract: Large Language Models (LLMs) represent one of the most significant advances in AI and natural language processing in recent years. Still, many pressing questions about their mechanisms, capabilities, and relationship to human cognit…

  981. arXiv cs.CL TIER_1 English(EN) · Maximilian Idahl, J\"org Tiedemann, Sampo Pyysalo, David Salinas, Tomasz Galica, Shenbin Qian, Tudor Nicolae Mateiu, Zihao Li, Anna Lokrantz, Fedor Vitiugin, Andr\'e F. T. Martins, Jenna Kanerva, Filip Ginter, Matthias Lindemann, Tim Isbister, Birger Moe… ·

    MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

    arXiv:2607.00890v1 Announce Type: new Abstract: Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approximately 4.8 trillion target-language tokens across 36…

  982. arXiv cs.CL TIER_1 English(EN) · Yangfan Hu, Xuhan Tong, Haoyue Bai, Xi Ding, Shashank Muralidhar Bharadwaj, Siyang Cao, Robert Nowak, Jiawei Zhang ·

    Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors

    arXiv:2607.00447v1 Announce Type: new Abstract: Large language models often produce hallucinated answers that violate prompt-level constraints. A key diagnostic question is whether these failures reflect missing knowledge, or whether the model has the relevant information but fol…

  983. arXiv cs.AI TIER_1 English(EN) · Zishuo Zheng, Vidhisha Balachandran, Chan Young Park, Faeze Brahman, Sachin Kumar ·

    Reasoning Up the Instruction Ladder for Controllable Language Models

    arXiv:2511.04694v5 Announce Type: replace-cross Abstract: As large language model (LLM) based systems take on high-stakes roles in real-world decision-making, they must reconcile competing instructions from multiple sources within a single prompt context. Enforcing an instruction…

  984. arXiv cs.CL TIER_1 English(EN) · Xianru Chen, Yukai Huang, Mingxiang Chen, Xinping Lei, Fangbing Deng, Jin Chen, Ge Zhang, Wenhao Huang, Jiaheng Liu ·

    MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark

    arXiv:2607.00724v1 Announce Type: new Abstract: Multilingual fluency often invites a stronger assumption: a model that can speak a user's language must also understand the culture encoded by that language. We call this the Illusion of Cultural Alignment. To test this assumption d…

  985. arXiv cs.CL TIER_1 English(EN) · Sewon Min ·

    Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale

    Language models (LMs) raise an intriguing alternative to vector-based retrieval: conditioning on an in-context corpus and directly generating a relevant answer. However, prior work has largely focused on proprietary systems or the smaller-scale reranking task, leaving corpus-scal…

  986. Hugging Face Daily Papers TIER_1 English(EN) ·

    The Wiola Architecture for Efficient Small Language Models

    We present Wiola, a fully original Small Language Model (SLM) architecture built from first principles, sharing no structural lineage with any existing model family including GPT, LLaMA, Mistral, or Falcon. Wiola introduces five independently novel components: (i) Spiral Rotary P…

  987. arXiv cs.CL TIER_1 English(EN) · Georg Groh ·

    $\text{Log}_\text{b}$Quant: Quantizing Language Models in Logarithmic Space

    Quantization has become an invaluable tool to reduce memory requirements and inference speed of modern language models, in particular to make them available for consumer setups and edge devices. While previous work has primarily focused on uniform quantization codebooks, such app…

  988. arXiv cs.AI TIER_1 (CA) · Biqing Huang ·

    CausalMix: Data Mixture as Causal Inference for Language Model Training

    In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture weights via proxy models, but they rely on the assumption of static data distributions. As a result, when the underlying data pool shifts, th…

  989. arXiv cs.CL TIER_1 Deutsch(DE) · Thomas Eisenmann ·

    Understanding Large Language Models

    Large Language Models (LLMs) represent one of the most significant advances in AI and natural language processing in recent years. Still, many pressing questions about their mechanisms, capabilities, and relationship to human cognition remain highly debated. This chapter aims to …

  990. Hugging Face Daily Papers TIER_1 English(EN) ·

    MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

    Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approximately 4.8 trillion target-language tokens across 36 European languages, produced by translating 100…

  991. arXiv cs.CL TIER_1 English(EN) · Gema Ramírez-Sánchez ·

    MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

    Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approximately 4.8 trillion target-language tokens across 36 European languages, produced by translating 100…

  992. arXiv cs.CL TIER_1 English(EN) · Jiaheng Liu ·

    MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark

    Multilingual fluency often invites a stronger assumption: a model that can speak a user's language must also understand the culture encoded by that language. We call this the Illusion of Cultural Alignment. To test this assumption directly, we introduce MSQA, a benchmark of 1,064…

  993. arXiv cs.AI TIER_1 English(EN) · Hanna Roed ·

    Auditing Forgetting in Limited Memory Language Models

    Limited Memory Language Models (LMLMs) externalize factual knowledge to a database to enable deletion-based unlearning without retraining. Existing evaluations measure post-deletion correctness in aggregate and cannot tell whether a deleted fact persists through residual parametr…

  994. arXiv cs.LG TIER_1 English(EN) · Julius Adebayo ·

    Prototype Language Models

    Knowing which training examples drive outputs is fundamental to auditing, correcting, and understanding language models, yet for modern LLMs this remains expensive, approximate, and largely post-hoc. Standard language models generate tokens through a dense network pathway, causin…

  995. arXiv cs.CL TIER_1 English(EN) · Jiawei Zhang ·

    Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors

    Large language models often produce hallucinated answers that violate prompt-level constraints. A key diagnostic question is whether these failures reflect missing knowledge, or whether the model has the relevant information but follows the wrong inference path. We study this phe…

  996. arXiv cs.AI TIER_1 English(EN) · Austin Feng, Marius Alonso, Ambroise Odonnat, Vasilii Feofanov, Ievgen Redko ·

    Optimal Self-Consistency for Efficient Reasoning with Large Language Models

    arXiv:2511.12309v2 Announce Type: replace-cross Abstract: Self-consistency (SC) is a widely used test-time inference technique for improving performance in chain-of-thought reasoning. It consists of generating multiple responses, or ``samples", from a large language model (LLM) a…

  997. arXiv cs.LG TIER_1 English(EN) · Jian Xu, Delu Zeng, John Paisley, Qibin Zhao ·

    Calibration, Not Compilation: Detecting and Repairing Misspecified Probabilistic Programs Written by Language Models

    arXiv:2606.31630v1 Announce Type: new Abstract: Language models increasingly write probabilistic programs (in NumPyro, Stan, or Pyro), but a program that compiles, runs, and passes every unit test can still be \emph{statistically} wrong -- a Gaussian likelihood for heavy-tailed d…

  998. arXiv cs.CL TIER_1 English(EN) · Ryan Lee, Jacob Biloki, Edward J. Hu, Jonathan May ·

    Sparse Layers are Critical to Scaling Looped Language Models

    arXiv:2605.09165v2 Announce Type: replace-cross Abstract: Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundaries. However, looped models do not scale as favorably as standard transf…

  999. arXiv cs.CL TIER_1 English(EN) · Dhruva Karkada, Daniel J. Korchinski, Andres Nava, Matthieu Wyart, Yasaman Bahri ·

    Symmetry in language statistics shapes the geometry of model representations

    arXiv:2602.15029v3 Announce Type: replace-cross Abstract: The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth one-dimensional manifold, and cities' latitu…

  1000. arXiv cs.CL TIER_1 English(EN) · Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu Zhang ·

    SemRF: A Semantic Reference Frame for Residual-Stream Dynamics in Language Models

    arXiv:2606.32022v1 Announce Type: cross Abstract: Residual-stream analysis asks how language-model computation evolves across depth, but intermediate decoding requires comparable readout coordinates across layers. If embedding anchors and unembedding readout disagree on the chose…

  1001. arXiv cs.CL TIER_1 English(EN) · Zhaojian Yu, Penghao Yin, Shuzheng Gao, Shilin He, Kai Cai, Xiao-Ping Zhang ·

    AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

    arXiv:2606.31551v1 Announce Type: new Abstract: Training language models (LMs) remains a highly human-intensive process, even as frontier language model agents become increasingly capable at software engineering and other long-horizon tasks. A central challenge is that autonomous…

  1002. arXiv cs.AI TIER_1 English(EN) · Sher Badshah, Ali Emami, Hassan Sajjad ·

    SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QA

    arXiv:2504.07385v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) become increasingly used for question-answering (QA), relying on static, pre-annotated references for evaluation poses significant challenges in cost, scalability, and completeness. Meanwhil…

  1003. arXiv cs.AI TIER_1 English(EN) · Daniele Gambetta, Gizem Gezici, Fosca Giannotti, Dino Pedreschi, Alistair Knott, Luca Pappalardo ·

    Learning by Surprise: Adaptive Mitigation of Model Collapse in Large Language Models

    arXiv:2410.12341v4 Announce Type: replace-cross Abstract: As AI-generated content increasingly populates the web, generative AI models are at growing risk of being trained on their own outputs, a process known as AI autophagy. This feedback loop has been shown to induce model col…

  1004. arXiv cs.AI TIER_1 English(EN) · Seyed Bagher Hashemi Natanzi, Bo Tang ·

    A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems

    arXiv:2606.31639v1 Announce Type: cross Abstract: Large language models are no longer only text generators. They are increasingly embedded in retrieval pipelines, enterprise assistants, coding environments, robotic systems, security-operation workflows, and autonomous agents that…

  1005. arXiv cs.AI TIER_1 English(EN) · Arash Raftari, Mehrdad Mahdavi, Nathan Blackthorn, Andrew Arash Mahyari ·

    Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models

    arXiv:2606.30899v1 Announce Type: cross Abstract: Backdoor attacks pose a serious threat to large language models (LLMs) by causing otherwise benign systems to produce attacker-specified malicious behavior when a hidden trigger is present. In this work, we study post hoc detoxifi…

  1006. arXiv cs.AI TIER_1 English(EN) · Parker Glenn, Alfy Samuel ·

    Large Databases Need Small, Open-Weight Language Models

    arXiv:2606.31808v1 Announce Type: new Abstract: Language model systems built around proprietary APIs often operate on a token-based cost model. This becomes prohibitively expensive in the context of large databases, where LM-enhanced relational operators can incur costs exceeding…

  1007. arXiv cs.AI TIER_1 English(EN) · Lisa Taldir (LI-PaRAD), Muhammad Ahmad Saeed (LI-PaRAD), David Defour (LI-PaRAD), Pablo de Oliveira Castro (LI-PaRAD), Eric Petit ·

    Benchmarking Large Language Models on Floating-Point Error Classification

    arXiv:2606.31308v1 Announce Type: new Abstract: This paper investigates the capability of Large Language Models (LLMs) to detect and classify floating-point errors statically in software code. We introduce InterFLOPBench, a benchmark of 90 C kernels with 1 130 test samples design…

  1008. Hugging Face Daily Papers TIER_1 (CA) ·

    CausalMix: Data Mixture as Causal Inference for Language Model Training

    CausalMix addresses limitations in LLM data mixing by formulating mixture optimization as a causal inference problem, enabling dynamic adaptation to shifting data distributions without costly retraining.

  1009. arXiv cs.CL TIER_1 English(EN) · Hongyu Zhang ·

    SemRF: A Semantic Reference Frame for Residual-Stream Dynamics in Language Models

    Residual-stream analysis asks how language-model computation evolves across depth, but intermediate decoding requires comparable readout coordinates across layers. If embedding anchors and unembedding readout disagree on the chosen span, apparent motion may reflect measurement dr…

  1010. arXiv cs.AI TIER_1 English(EN) · Alfy Samuel ·

    Large Databases Need Small, Open-Weight Language Models

    Language model systems built around proprietary APIs often operate on a token-based cost model. This becomes prohibitively expensive in the context of large databases, where LM-enhanced relational operators can incur costs exceeding $10,000 for a single set of experiments, hinder…

  1011. arXiv cs.AI TIER_1 English(EN) · Bo Tang ·

    A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems

    Large language models are no longer only text generators. They are increasingly embedded in retrieval pipelines, enterprise assistants, coding environments, robotic systems, security-operation workflows, and autonomous agents that can read private data, call tools, write files, e…

  1012. arXiv cs.LG TIER_1 English(EN) · Qibin Zhao ·

    Calibration, Not Compilation: Detecting and Repairing Misspecified Probabilistic Programs Written by Language Models

    Language models increasingly write probabilistic programs (in NumPyro, Stan, or Pyro), but a program that compiles, runs, and passes every unit test can still be \emph{statistically} wrong -- a Gaussian likelihood for heavy-tailed data, a Poisson for over-dispersed counts, an inv…

  1013. arXiv cs.CL TIER_1 English(EN) · Xiao-Ping Zhang ·

    AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

    Training language models (LMs) remains a highly human-intensive process, even as frontier language model agents become increasingly capable at software engineering and other long-horizon tasks. A central challenge is that autonomous post-training is not just a coding problem: it …

  1014. arXiv cs.LG TIER_1 English(EN) · Matthias Br\"andel, Stephan K\"ohler, Oliver Rheinbach ·

    How Token Influence Decays with Distance: A Green-Function View of Trained Language Models

    arXiv:2606.29139v1 Announce Type: new Abstract: We study how the next-token prediction of an autoregressive Transformer language model changes under small perturbations of earlier input token embeddings. Motivated by operator learning and iterative solvers for differential equati…

  1015. arXiv cs.CL TIER_1 English(EN) · Hongyuan Adam Lu, Z. L., Victor Wei, Zefan Zhang, Zhao Hong, Qiqi Xiang, Bowen Cao, Wai Lam ·

    Adam's Law: Textual Frequency Law on Large Language Models

    arXiv:2604.02176v3 Announce Type: replace Abstract: While textual frequency has been validated as relevant to human cognition in reading speed, its relatedness to Large Language Models (LLMs) is seldom studied. We propose a novel research direction in terms of textual data freque…

  1016. arXiv cs.CL TIER_1 English(EN) · Jinchao Hu, Meizhi Zhong, Kehai Chen, Xuefeng Bai, Min Zhang ·

    Agentic Tool Use in Large Language Models

    arXiv:2604.00835v2 Announce Type: replace Abstract: Large language models are increasingly being deployed as autonomous agents yet their real world effectiveness depends on reliable tools for information retrieval, computation and external action. Existing studies remain fragment…

  1017. arXiv cs.CL TIER_1 English(EN) · Anna Bair, Yixuan Even Xu, Mingjie Sun, J. Zico Kolter ·

    Compressed Sensing for Capability Localization in Large Language Models

    arXiv:2603.03335v2 Announce Type: replace Abstract: Large language models (LLMs) exhibit a wide range of capabilities, including mathematical reasoning, code generation, and linguistic behaviors. We show that Transformer architectures contain small subsets of attention heads that…

  1018. arXiv cs.CL TIER_1 English(EN) · Yudao Sun, Juan Yin, Juan Zhao, Fan Zhang, Yongheng Liu, Hongji Chen ·

    Unified Enhancement of the Generalization and Robustness of Language Models via Bi-Stage Optimization

    arXiv:2503.16550v2 Announce Type: replace Abstract: Neural network language models (LMs) are confronted with significant challenges in generalization and robustness. Currently, many studies focus on improving either generalization or robustness in isolation, without methods addre…

  1019. arXiv cs.CL TIER_1 English(EN) · Ivan Vykopal, Mat\'u\v{s} Pikuliak, Simon Ostermann, Mari\'an \v{S}imko ·

    Generative Large Language Models in Automated Fact-Checking: A Survey

    arXiv:2407.02351v3 Announce Type: replace Abstract: The rapid spread of false and misleading information on online platforms poses a growing societal challenge, overwhelming the capacity of manual fact-checking and increasing the demand for scalable, reliable automation. Recent a…

  1020. arXiv cs.CL TIER_1 English(EN) · Hengxiang Zhang, Jiaxi Ren, Hongxin Wei ·

    Understanding Evaluation Illusion in Diffusion Large Language Models

    arXiv:2606.29228v1 Announce Type: new Abstract: Despite the capability of parallel decoding, diffusion large language models (dLLMs) require many denoising steps to maintain generation quality, motivating recent research on efficient decoding strategies. However, existing studies…

  1021. arXiv cs.AI TIER_1 English(EN) · Zheng Luo, T Pranav Kutralingam, Ogochukwu N Okoani, Wanpeng Xu, Hua Wei, Xiyang Hu ·

    Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models

    arXiv:2601.05366v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly deployed as agents that invoke external tools through structured function calls. While recent work reports strong tool-calling performance under standard English-centric evalua…

  1022. arXiv cs.AI TIER_1 English(EN) · Dari Baturova, Elena Bruches, Ivan Chernov, Roman Derunets, Arsenii Fomin, Andrey Kostin ·

    Little Brains, Big Feats: Exploring Compact Language Models

    arXiv:2606.30062v1 Announce Type: cross Abstract: While large language models have been dominating the research landscape recently, small language models remain highly relevant across various domains; yet, they receive far less attention. In this study, we investigate how smaller…

  1023. arXiv cs.AI TIER_1 English(EN) · Andrew Mack, Nina Panickssery, Alexander Matt Turner ·

    Mechanistically Eliciting Latent Behaviors in Language Models

    arXiv:2606.29604v1 Announce Type: cross Abstract: We aim to discover diverse, generalizable perturbations of LLM internals that can surface hidden behavioral modes. Such perturbations could help reshape model behavior and systematically evaluate potential risks. We introduce Caus…

  1024. arXiv cs.AI TIER_1 English(EN) · Arun Vignesh Malarkkan, Manan Roy Choudhury, Utkarsh Byahut, Yash Ravindra Charde, Vivek Gupta, Yanjie Fu ·

    Invariant Reasoning Directions in Latent Trajectories of Language Models

    arXiv:2606.29164v1 Announce Type: cross Abstract: Latent reasoning models perform multi-step inference directly in hidden-state space, yet the structure of these latent reasoning trajectories remains poorly understood. We show that contrastive refinement signals between stronger …

  1025. arXiv cs.AI TIER_1 English(EN) · Zhaoyang Li, Ruijie Zhang, Jiaqi Liu, Zhaoji Sun ·

    Fine-Tuning General-Purpose Large Language Models for Agricultural Applications:A Reproducible Framework and Evaluation Protocol Based on Qwen3-8B

    arXiv:2606.28992v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) have demonstrated strong abilities in opendomain question answering, information extraction, and text generation. Agricultural applications, however, are domain-specific, region-depende…

  1026. arXiv cs.AI TIER_1 English(EN) · Haobo Yang ·

    Using Large Language Models as Low-Cost Statistical Estimators for Human-Response Data

    arXiv:2606.30372v1 Announce Type: new Abstract: Quantitative research across the social and behavioral sciences depends on human subject experiments that are expensive, slow, and subject to sampling bias. Here we show that pretrained large language models induce risk-equivalent e…

  1027. arXiv cs.LG TIER_1 English(EN) · Yuchen Ma, Yue Huang, Wenjie Wang, Xiaonan Luo, Xiangliang Zhang, Stefan Feuerriegel ·

    Synthetic Interaction Data for Scalable Personalization in Large Language Models

    arXiv:2602.12394v2 Announce Type: replace Abstract: Personalized prompting offers large opportunities for deploying large language models (LLMs) to diverse users, yet existing prompt optimization methods primarily focus on task-level optimization while largely overlooking user-sp…

  1028. arXiv cs.AI TIER_1 (CA) · Guanyu Chen, Chenxiao Yu, Xiyang Hu ·

    Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict

    arXiv:2601.03546v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to simulate decision-making tasks involving personal data sharing, where privacy concerns and prosocial motivations can push choices in opposite directions. Existing evalu…

  1029. arXiv cs.AI TIER_1 English(EN) · Guangzhi Sun, Xiao Zhan, Shutong Feng, Philip C. Woodland, Jose Such ·

    CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models

    arXiv:2501.14940v4 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption. Current LLM safety benchmarks often focus solely on the refusal of individual problematic queries, whi…

  1030. arXiv cs.AI TIER_1 English(EN) · Ali AhmadiTeshnizi, Wenzhi Gao, Herman Brunborg, Shayan Talaei, Connor Lawless, Madeleine Udell ·

    OptiMUS-0.3: Using Large Language Models to Model and Solve Optimization Problems at Scale

    arXiv:2407.19633v4 Announce Type: replace Abstract: Optimization problems are pervasive in sectors from manufacturing and distribution to healthcare. However, most such problems are still solved heuristically by hand rather than optimally by state-of-the-art solvers because the e…

  1031. arXiv cs.AI TIER_1 Nederlands(NL) · Mohammed Abu Baker, Lakshmi Babu-Saheer ·

    Fuzzing Large Language Models to Elicit Hidden Behaviors

    arXiv:2606.29646v1 Announce Type: cross Abstract: Sleeper agents are the canonical model organism of deception: models trained to behave normally but to emit an unsafe behaviour on a specific trigger. Eliciting that behaviour without knowing the trigger has not been studied syste…

  1032. Hugging Face Daily Papers TIER_1 English(EN) ·

    AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

    AutoTrainess enables autonomous language model training by providing structured agent-computer interfaces that guide planning, data preparation, training, evaluation, and logging operations more effectively than traditional command-line approaches.

  1033. arXiv cs.AI TIER_1 English(EN) · Haobo Yang ·

    Using Large Language Models as Low-Cost Statistical Estimators for Human-Response Data

    Quantitative research across the social and behavioral sciences depends on human subject experiments that are expensive, slow, and subject to sampling bias. Here we show that pretrained large language models induce risk-equivalent estimators of conditional expectations under squa…

  1034. arXiv cs.CL TIER_1 English(EN) · Andrey Kostin ·

    Little Brains, Big Feats: Exploring Compact Language Models

    While large language models have been dominating the research landscape recently, small language models remain highly relevant across various domains; yet, they receive far less attention. In this study, we investigate how smaller language models perform during the generation sta…

  1035. Hugging Face Daily Papers TIER_1 English(EN) ·

    Little Brains, Big Feats: Exploring Compact Language Models

    While large language models have been dominating the research landscape recently, small language models remain highly relevant across various domains; yet, they receive far less attention. In this study, we investigate how smaller language models perform during the generation sta…

  1036. arXiv cs.CL TIER_1 English(EN) · Eric De Giuli ·

    Scaling limit of the Random Language Model

    arXiv:2606.28105v1 Announce Type: cross Abstract: We develop a quantitative theory of the Random Language Model (RLM), an ensemble of stochastic context-free grammars, in a scaling limit where the number of hidden symbols $N \to \infty$ while the grammar temperature $\tilde{\epsi…

  1037. arXiv cs.AI TIER_1 English(EN) · Artem Ploujnikov, Francesco Verdini, Samir Sadok, Mirco Ravanelli ·

    HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models

    arXiv:2606.27627v1 Announce Type: cross Abstract: Discrete audio representations have become increasingly popular for building multimodal text-audio systems and integrating audio capabilities into Large Language Models (LLMs). However, numerous studies report performance degradat…

  1038. arXiv cs.AI TIER_1 English(EN) · Ponhvoan Srey, Xiaobao Wu, Cong-Duy Nguyen, Quang Minh Nguyen, Duc Anh Vu, Anh Tuan Luu ·

    From Signals to Transfer: A Factorised Study of Probe-Based Uncertainty Estimation in Large Language Models

    arXiv:2606.27679v1 Announce Type: cross Abstract: Probe-based uncertainty estimation (UE) has emerged as a prominent approach to detect hallucinations in Large Language Models (LLMs) by learning uncertainty from internal model signals. Yet, recent methods vary simultaneously acro…

  1039. arXiv cs.AI TIER_1 English(EN) · Hira Beril Kucuk, Norman W Paton, Jiaoyan Chen, Zhenyu Wu ·

    Single and Multi Truth Data Fusion using Large Language Models

    arXiv:2606.28062v1 Announce Type: cross Abstract: Data fusion, also known as truth discovery, is a data integration problem that aims to determine the correct value or set of values for each attribute of an object when presented with potentially conflicting values from multiple s…

  1040. arXiv cs.AI TIER_1 English(EN) · Zhen Zeng, Leijiang Gu, Zhangling Duan, Feng Li, Cees G. M. Snoek, Meng Wang, Zenglin Shi ·

    Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning

    arXiv:2511.20196v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) can inadvertently memorize privacy-sensitive information during training. While existing unlearning methods can remove such content, they often severely degrade the model's foundational c…

  1041. arXiv cs.AI TIER_1 English(EN) · Xinyi Wu, Geng Hong, Pei Chen, Yueyue Chen, Xudong Pan, Min Yang ·

    PRISON: Unmasking the Criminal Potential of Large Language Models

    arXiv:2506.16150v4 Announce Type: replace-cross Abstract: As large language models (LLMs) advance, concerns about their misconduct in complex social contexts intensify. Existing research overlooked the systematic understanding and assessment of their criminal capability in realis…

  1042. arXiv cs.CL TIER_1 English(EN) · Iskander Azangulov, Kianoosh Ashouritaklimi, Leo Zhang, Simon Vary, Patrick Rebeschini ·

    Masked Language Flow Models

    arXiv:2606.27617v1 Announce Type: new Abstract: Masked Diffusion Models (MDMs) promise fast, parallel language generation, but their reverse transition factorises across token positions -- an approximation that breaks down in the few-step sampling regime where parallel generation…

  1043. arXiv cs.CL TIER_1 English(EN) · Ting Ma, Xiufeng Huang, Benlei Cui, Xiaowen Xu, Shikai Qiu, Ruijie Jian, Hongxing Li, Guanghui Wang, Longtao Huang, Haiwen Hong, Haolei Xu, Wenjing Jiang, Ziwen Xu, Zhaoyu Fan, Shaoxuan He, Chuxi Xiao, Yujian Li, Xinyue Chen, Chunyang Chai, Wenxuan Liu, … ·

    Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety

    arXiv:2606.27632v1 Announce Type: new Abstract: As large language models are increasingly deployed in real-world systems, safety failures can still lead to harmful outputs and dangerous misuse. We argue that the essence of safety is adversarial: many failures arise not from natur…

  1044. arXiv cs.CL TIER_1 English(EN) · Antonios Anastasopoulos, Giuseppe Ateniese, Evgenios M. Kornaropoulos ·

    Safe Language Generation in the Limit

    arXiv:2601.08648v2 Announce Type: replace Abstract: Recent results in learning a language in the limit have shown that, although language identification is impossible, language generation is tractable. As this foundational area expands, we need to consider the implications of lan…

  1045. arXiv cs.LG TIER_1 English(EN) · Fan Mo, Yuxuan Han, Geng Zhang, Wangbo Zhao, Yang You ·

    FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models

    arXiv:2606.27866v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) language models scale model ability with sparsely activated experts, making this architecture a standard recipe for modern large models. However, sparse activation does not remove the deployment burden of st…

  1046. Hugging Face Daily Papers TIER_1 English(EN) ·

    Little Brains, Big Feats: Exploring Compact Language Models

    Small language models can effectively perform retrieval-augmented generation tasks directly on-device without GPU acceleration.

  1047. Hugging Face Daily Papers TIER_1 English(EN) ·

    Modular Cognitive Architecture Emerges in Large Language Models

    Large language models develop modular neural architectures that mirror human brain specialization across language, reasoning, and physical cognition, suggesting modularity is a fundamental property of intelligent systems.

  1048. arXiv cs.CL TIER_1 English(EN) · Eric De Giuli ·

    Scaling limit of the Random Language Model

    We develop a quantitative theory of the Random Language Model (RLM), an ensemble of stochastic context-free grammars, in a scaling limit where the number of hidden symbols $N \to \infty$ while the grammar temperature $\tildeε_d \to 0$ at fixed $x = {\tildeε}_d \log N$. In this li…

  1049. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Zhenyu Wu ·

    Single and Multi Truth Data Fusion using Large Language Models

    Data fusion, also known as truth discovery, is a data integration problem that aims to determine the correct value or set of values for each attribute of an object when presented with potentially conflicting values from multiple sources. Data fusion tasks belong to two main categ…

  1050. arXiv cs.LG TIER_1 English(EN) · Yang You ·

    FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models

    Mixture-of-Experts (MoE) language models scale model ability with sparsely activated experts, making this architecture a standard recipe for modern large models. However, sparse activation does not remove the deployment burden of storing and serving all experts, and the available…

  1051. arXiv cs.AI TIER_1 English(EN) · Francois Crespin (IP Paris, LTCI), Fabian M. Suchanek (IP Paris, LTCI), Nils Holzenberger ·

    KARLA: Knowledge-base Augmented Retrieval for Language Models

    arXiv:2606.26807v1 Announce Type: new Abstract: We propose a new method that allows an LLM to automatically pull in factual knowledge from a knowledge base during token generation. This means that (1)~factual knowledge in the LLM output can be updated without retraining the LLM, …

  1052. arXiv cs.AI TIER_1 English(EN) · Renwei Meng, Bowen Zhang, Jian Wang, Xican Wang, Haoyi Wu, Xuanyan Qiu, Shengan Yang ·

    Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models

    arXiv:2606.26101v1 Announce Type: cross Abstract: Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contamination, prompt idiosyncrasy, or generic refusal behavior. We present a contamina…

  1053. arXiv cs.CL TIER_1 English(EN) · Beatrice Savoldi, Sara Papi, Wafa Aissa, Matteo Negri, Luisa Bentivogli ·

    RedVox: Safety and Fairness Gaps in Speech Models Across Languages

    arXiv:2606.26968v1 Announce Type: new Abstract: Speech-capable models are increasingly deployed in real-world applications across languages. Yet their safety and fairness beyond English settings and under naturalistic conditions remain understudied. We survey safety reporting pra…

  1054. arXiv cs.CL TIER_1 English(EN) · Ananth K S, Arya Hariharan ·

    Reproducibility Study of "AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models"

    arXiv:2606.26783v1 Announce Type: cross Abstract: Fang et al. (2025) introduced a null-space constrained projection, named AlphaEdit, for locate-then-edit knowledge editing methods, theoretically guaranteeing that edits do not disrupt previously preserved knowledge, and reports s…

  1055. arXiv cs.AI TIER_1 English(EN) · Henry Shaowu Yuchi, Michal Kucer, Benjamin H. Sims, Selma Peterson, Emily Taylor ·

    NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models

    arXiv:2606.27047v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly technical domains remains a significant challenge. In nuclear engineering, problem solving of…

  1056. arXiv cs.AI TIER_1 English(EN) · Siddhant Hitesh Mantri, Dhara Gorasiya, Malhar Kulkarni, Pushpak Bhattacharya ·

    From Lexicon to AI: A Structured-Data Pipeline for Specialized Conversational Systems in Low-Resource Languages

    arXiv:2606.26112v1 Announce Type: cross Abstract: Low-resource languages face a critical challenge in AI development: creating specialized conversational systems without access to massive training corpora. We present a systematic methodology for transforming structured linguistic…

  1057. arXiv cs.AI TIER_1 English(EN) · Yuxuan Yang, Feiyang Li, Yile Wang ·

    \textsc{DiARC}: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models

    arXiv:2606.26530v1 Announce Type: cross Abstract: The Abstraction and Reasoning Corpus (ARC;~\citealp{chollet2019measure}) contains tasks that require summarizing patterns from limited grid samples and predicting output grids. Recently, many large language model based approaches …

  1058. arXiv cs.AI TIER_1 English(EN) · Haoyan Luo, Lucia Specia ·

    Tuning Language Models by Mixture-of-Depths Ensemble

    arXiv:2410.13077v2 Announce Type: replace-cross Abstract: Transformer-based Large Language Models (LLMs) traditionally rely on final-layer loss for finetuning and final-layer representations for predictions, potentially overlooking the predictive power embedded in late layers. In…

  1059. arXiv cs.AI TIER_1 English(EN) · Kylie Anglin ·

    Estimating Uncertainty in Classifier Performance with Applications to Large Language Models and Nested Data

    arXiv:2606.26422v1 Announce Type: new Abstract: Researchers increasingly use text classification--supervised models or large language models--to measure constructs from natural language, providing metrics such as recall and precision as evidence of their validity. Yet, though the…

  1060. arXiv cs.LG TIER_1 English(EN) · Mingi Kim, Yongjun Kim, Jungwoo Kang, Hyungki Kim ·

    BrepCoder: A Unified Multimodal Large Language Model for Multi-task B-rep Reasoning

    arXiv:2602.22284v3 Announce Type: replace Abstract: Recent advancements in deep learning have actively addressed complex challenges within the Computer-Aided Design (CAD) domain.However, most existing approaches rely on task-specifi c models requiring structural modifi cations fo…

  1061. arXiv cs.CL TIER_1 English(EN) · Niklas Mellgren, Peter Schneider-Kamp, Lukas Galke Poech ·

    Training Language Models to Use Prolog as a Tool

    arXiv:2512.07407v3 Announce Type: replace Abstract: Language models frequently produce plausible yet incorrect reasoning traces that are difficult to verify. We investigate fine-tuning models to use Prolog as an external symbolic reasoning tool, training Qwen2.5-3B-Instruct with …

  1062. arXiv cs.CL TIER_1 English(EN) · Anh Tuan Luu ·

    From Signals to Transfer: A Factorised Study of Probe-Based Uncertainty Estimation in Large Language Models

    Probe-based uncertainty estimation (UE) has emerged as a prominent approach to detect hallucinations in Large Language Models (LLMs) by learning uncertainty from internal model signals. Yet, recent methods vary simultaneously across feature design, training data construction, and…

  1063. arXiv cs.CL TIER_1 English(EN) · Hui Xue ·

    Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety

    As large language models are increasingly deployed in real-world systems, safety failures can still lead to harmful outputs and dangerous misuse. We argue that the essence of safety is adversarial: many failures arise not from natural inputs alone, but from strategic attempts to …

  1064. arXiv cs.CL TIER_1 English(EN) · Patrick Rebeschini ·

    Masked Language Flow Models

    Masked Diffusion Models (MDMs) promise fast, parallel language generation, but their reverse transition factorises across token positions -- an approximation that breaks down in the few-step sampling regime where parallel generation ought to provide the greatest efficiency gains.…

  1065. arXiv cs.AI TIER_1 English(EN) · Emily Taylor ·

    NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models

    Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly technical domains remains a significant challenge. In nuclear engineering, problem solving often requires not only factual knowledge but also q…

  1066. arXiv cs.CL TIER_1 English(EN) · Luisa Bentivogli ·

    RedVox: Safety and Fairness Gaps in Speech Models Across Languages

    Speech-capable models are increasingly deployed in real-world applications across languages. Yet their safety and fairness beyond English settings and under naturalistic conditions remain understudied. We survey safety reporting practices across state-of-the-art speech model rele…

  1067. arXiv cs.AI TIER_1 English(EN) · Nils Holzenberger ·

    KARLA: Knowledge-base Augmented Retrieval for Language Models

    We propose a new method that allows an LLM to automatically pull in factual knowledge from a knowledge base during token generation. This means that (1)~factual knowledge in the LLM output can be updated without retraining the LLM, (2)~facts in the LLM output can be traced to the…

  1068. arXiv cs.LG TIER_1 English(EN) · Arya Hariharan ·

    Reproducibility Study of "AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models"

    Fang et al. (2025) introduced a null-space constrained projection, named AlphaEdit, for locate-then-edit knowledge editing methods, theoretically guaranteeing that edits do not disrupt previously preserved knowledge, and reports substantial gains over existing editing methods on …

  1069. arXiv cs.CL TIER_1 English(EN) · Akshay Paruchuri, Sanmi Koyejo, Ehsan Adeli ·

    Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models

    arXiv:2606.26079v1 Announce Type: new Abstract: Standard benchmarks for multimodal large language models (MLLMs) score each item on one canonical ordering and miss whether order-irrelevant shuffling changes the answer, a baseline reliability property called for by emerging AI eva…

  1070. arXiv cs.LG TIER_1 English(EN) · Ke Xu, Jiaqi Wan, Wenhao Hu, Han Pu, Xiaoyun Wang ·

    EPTS: Elastic Post-Training Sparsity for Efficient Large Language Model Compression

    arXiv:2606.25285v1 Announce Type: new Abstract: Post-Training Sparsity (PTS) has emerged as a crucial paradigm for compressing Large Language Models to facilitate efficient deployment on resource-constrained devices. However, existing PTS methodologies are typically confined to S…

  1071. arXiv cs.LG TIER_1 English(EN) · Habibullah Akbar ·

    ATMA: Length-Invariant Language Modeling via Polar Attention and Gated-Delta Compression Memory

    arXiv:2606.25156v1 Announce Type: new Abstract: Modern large language models based on softmax scaled-dot-product attention are constrained by their training sequence length: as the key-value sequence grows, softmax probability mass can dilute across a wider distribution, inducing…

  1072. arXiv cs.LG TIER_1 English(EN) · Rituraj Sharma, Tu Vu ·

    Dense Supervision Is Not Enough: The Readout Blind Spot in Looped Language Models

    arXiv:2606.24898v1 Announce Type: new Abstract: Looped language models turn hidden states into runtime state: each state is decoded for prediction and fed back into future computation. This creates a basic supervision question: which state variables does cross-entropy actually co…

  1073. arXiv cs.CL TIER_1 English(EN) · Nicolas Flammarion, Chirag Pabbaraju, Hristo Papazov, Miltiadis Stouras, Ola Svensson ·

    Space-Efficient Language Generation in the Limit

    arXiv:2606.25777v1 Announce Type: cross Abstract: We initiate a resource-aware theory of \textit{language generation in the limit} under the minimal constraint of space efficiency. In our framework, a learner observes an adversarial positive stream from a target language $K$ and …

  1074. arXiv cs.CL TIER_1 English(EN) · Tomoya Mizumoto, Yusuke Fujita, Hao Shi, Lianbo Liu, Atsushi Kojima, Yui Sudo ·

    Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language Models

    arXiv:2606.25436v1 Announce Type: cross Abstract: Dialogue systems based on large language models (LLMs) have advanced significantly in recent years. However, dialectal variation remains a major challenge, particularly for systems that process spoken input. LLM-based speech langu…

  1075. arXiv cs.CL TIER_1 English(EN) · Yile Wang ·

    \textsc{DiARC}: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models

    The Abstraction and Reasoning Corpus (ARC;~\citealp{chollet2019measure}) contains tasks that require summarizing patterns from limited grid samples and predicting output grids. Recently, many large language model based approaches have attempted to transform it into a text-based r…

  1076. Hugging Face Daily Papers TIER_1 English(EN) ·

    RedVox: Safety and Fairness Gaps in Speech Models Across Languages

    Multilingual safety and fairness benchmark for speech models reveals persistent vulnerabilities across languages and naturalistic conditions.

  1077. Hugging Face Daily Papers TIER_1 English(EN) ·

    Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models

    Standard benchmarks for multimodal large language models (MLLMs) score each item on one canonical ordering and miss whether order-irrelevant shuffling changes the answer, a baseline reliability property called for by emerging AI evaluation guidelines. We introduce Facet-Probe, a …

  1078. arXiv cs.CL TIER_1 English(EN) · Ehsan Adeli ·

    Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models

    Standard benchmarks for multimodal large language models (MLLMs) score each item on one canonical ordering and miss whether order-irrelevant shuffling changes the answer, a baseline reliability property called for by emerging AI evaluation guidelines. We introduce Facet-Probe, a …

  1079. arXiv cs.AI TIER_1 English(EN) · Ola Svensson ·

    Space-Efficient Language Generation in the Limit

    We initiate a resource-aware theory of \textit{language generation in the limit} under the minimal constraint of space efficiency. In our framework, a learner observes an adversarial positive stream from a target language $K$ and must eventually output a hallucination-free hypoth…

  1080. Hugging Face Daily Papers TIER_1 English(EN) ·

    PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models

    Safety alignment of large language models (LLMs) typically depends on high-quality supervision data, such as safe demonstrations or preference pairs. However, in real-world deployment, emerging safety requirements are often specified as natural-language policies, while correspond…

  1081. arXiv cs.CL TIER_1 English(EN) · Yui Sudo ·

    Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language Models

    Dialogue systems based on large language models (LLMs) have advanced significantly in recent years. However, dialectal variation remains a major challenge, particularly for systems that process spoken input. LLM-based speech language models (SLMs), which integrate LLMs with speec…

  1082. arXiv cs.AI TIER_1 English(EN) · Hanna Lee, Tan Dat Nguyen, Jaehoon Kang, Kyuhong Shim ·

    WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models

    arXiv:2604.08558v2 Announce Type: replace-cross Abstract: Recent decoder-only autoregressive text-to-speech (AR-TTS) models produce high-fidelity speech, but their memory and compute costs scale quadratically with sequence length due to full self-attention. In this paper, we prop…

  1083. arXiv cs.CL TIER_1 (CA) · Yang Liu, Chenhui Chu ·

    Sentence-Level Contextual Entrainment in Large Language Models

    arXiv:2606.24077v1 Announce Type: new Abstract: Contextual entrainment, which is a newly discovered phenomenon in large language models (LLMs), refers to the tendency of a model to assign higher probabilities to tokens that appear in its context. In this work, we extend this phen…

  1084. arXiv cs.AI TIER_1 English(EN) · Shaoshuai Du, Penghao Liang, Yixian Shen, Chuanqi Shi, Hang Zhang, Lun Wang ·

    On the Stability of Prompt Ranking in Large Language Model Evaluation

    arXiv:2606.24381v1 Announce Type: cross Abstract: Prompt-based interaction has become a dominant paradigm for using large language models (LLMs), where multiple candidate prompts are evaluated and the top-ranked one is selected for downstream use. This workflow implicitly assumes…

  1085. arXiv cs.AI TIER_1 English(EN) · Conor Finlay, Joshua Kurien, Saurabh Dash, Marzieh Fadaee, Beyza Ermis ·

    CALIBER: Calibrating Confidence Before and After Reasoning in Language Models

    arXiv:2606.24281v1 Announce Type: cross Abstract: Reasoning language models are increasingly asked not only to answer difficult questions, but also to estimate their likelihood of success. Existing methods typically elicit confidence only once: either before thinking or after ans…

  1086. arXiv cs.AI TIER_1 English(EN) · Morayo Danielle Adeyemi, Ryan A. Rossi, Franck Dernoncourt ·

    CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression

    arXiv:2606.24083v1 Announce Type: cross Abstract: "Talk short. Drop grammar. Save token." This caveman style is widely promoted as a way to cut inference cost, but whether it actually saves anything depends on which channel (the user's prompt or the model's response) is being com…

  1087. arXiv cs.AI TIER_1 English(EN) · Sauro Succi, Peter V. Coveney, Alex Hansen ·

    On the Smallness of the Large Language Models Scaling Exponents

    arXiv:2606.24504v1 Announce Type: new Abstract: We discuss reasons why the scaling exponents of current Large Language Models (LLMs) applications are indicating an unsustainable regime in terms of energy resources. We further show that attributing the smallness of such exponents …

  1088. arXiv cs.CL TIER_1 English(EN) · Petr Nyoma ·

    Harmonic: Hierarchical State Space Models for Efficient Long-Context Language Modeling

    arXiv:2606.24650v1 Announce Type: new Abstract: We present Harmonic, a hierarchical state space model (SSM) for language modeling. The architecture stacks three recurrent levels at progressively slower timescales; each level receives the prediction error of the level below as inp…

  1089. arXiv cs.CL TIER_1 English(EN) · Manan Agarwal, Sheel Shah, Chanhyuk Lee, Jaehoon Yoo, Jerry Huang, Seunghoon Hong, Aditi Raghunathan, Jinwoo Kim, Nicholas M. Boffi ·

    Posterior Refinement: Fast Language Generation via Any-Order Flow Maps

    arXiv:2606.24773v1 Announce Type: new Abstract: Non-autoregressive generation offers a powerful paradigm for iterative refinement, allowing models to recursively critique, erase and regenerate arbitrary subsets of tokens. However, existing non-autoregressive models fail to realiz…

  1090. arXiv cs.CL TIER_1 English(EN) · M\u{a}d\u{a}lina Zgreab\u{a}n, Tejaswini Deoskar, Lasha Abzianidze ·

    MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference

    arXiv:2510.24295v2 Announce Type: replace Abstract: As many benchmarks have become saturated, it has become increasingly important to create new datasets that evaluate the generalization capacity of current state-of-the-art models in reasoning. However, designing high-quality rea…

  1091. Hugging Face Daily Papers TIER_1 English(EN) ·

    Revealing Training Data Exposure in Vision Language Large Models via Parameter Gradients

    Vision-Language Large Models (VLLMs) trained on massive crawled corpora raise pressing copyright and data-provenance concerns. These concerns are particularly acute in healthcare, where patient medical images paired with clinical reports demand rigorous privacy safeguards. Howeve…

  1092. arXiv cs.CL TIER_1 English(EN) · Nicholas M. Boffi ·

    Posterior Refinement: Fast Language Generation via Any-Order Flow Maps

    Non-autoregressive generation offers a powerful paradigm for iterative refinement, allowing models to recursively critique, erase and regenerate arbitrary subsets of tokens. However, existing non-autoregressive models fail to realize this potential. Masked Diffusion Models (MDMs)…

  1093. Hugging Face Daily Papers TIER_1 English(EN) ·

    Can Scale Save Us From Plasticity Loss in Large Language Models?

    The loss of plasticity - the ability of a network to learn new information after having already learned older information - is a fundamental challenge in creating artificial neural networks capable of continual learning. Although this phenomenon has been known for decades, it has…

  1094. arXiv cs.AI TIER_1 English(EN) · Alex Hansen ·

    On the Smallness of the Large Language Models Scaling Exponents

    We discuss reasons why the scaling exponents of current Large Language Models (LLMs) applications are indicating an unsustainable regime in terms of energy resources. We further show that attributing the smallness of such exponents to a numerical bias due to the neglect of a non-…

  1095. arXiv cs.AI TIER_1 English(EN) · Lun Wang ·

    On the Stability of Prompt Ranking in Large Language Model Evaluation

    Prompt-based interaction has become a dominant paradigm for using large language models (LLMs), where multiple candidate prompts are evaluated and the top-ranked one is selected for downstream use. This workflow implicitly assumes that prompt rankings are stable under minor varia…

  1096. arXiv cs.CL TIER_1 English(EN) · Beyza Ermis ·

    CALIBER: Calibrating Confidence Before and After Reasoning in Language Models

    Reasoning language models are increasingly asked not only to answer difficult questions, but also to estimate their likelihood of success. Existing methods typically elicit confidence only once: either before thinking or after answering. We argue that confidence in reasoning mode…

  1097. Hugging Face Daily Papers TIER_1 English(EN) ·

    CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression

    "Talk short. Drop grammar. Save token." This caveman style is widely promoted as a way to cut inference cost, but whether it actually saves anything depends on which channel (the user's prompt or the model's response) is being compressed. We present Cavewoman, a two-channel evalu…

  1098. arXiv cs.CL TIER_1 English(EN) · Franck Dernoncourt ·

    CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression

    "Talk short. Drop grammar. Save token." This caveman style is widely promoted as a way to cut inference cost, but whether it actually saves anything depends on which channel (the user's prompt or the model's response) is being compressed. We present Cavewoman, a two-channel evalu…

  1099. arXiv cs.CL TIER_1 (CA) · Chenhui Chu ·

    Sentence-Level Contextual Entrainment in Large Language Models

    Contextual entrainment, which is a newly discovered phenomenon in large language models (LLMs), refers to the tendency of a model to assign higher probabilities to tokens that appear in its context. In this work, we extend this phenomenon from the token level to the sentence leve…

  1100. Hugging Face Daily Papers TIER_1 (CA) ·

    Sentence-Level Contextual Entrainment in Large Language Models

    Contextual entrainment, which is a newly discovered phenomenon in large language models (LLMs), refers to the tendency of a model to assign higher probabilities to tokens that appear in its context. In this work, we extend this phenomenon from the token level to the sentence leve…

  1101. Hugging Face Daily Papers TIER_1 English(EN) ·

    CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression

    Two-channel evaluation shows output compression reduces costs while input compression increases costs and degrades accuracy across models and datasets.

  1102. arXiv cs.AI TIER_1 English(EN) · Tomer Solberg ·

    HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models

    We present HyperQuant (Hadamard, optimallY Packing, Entropy Rice-coding), a unified post-training quantization pipeline for the weights and the KV cache of large language and diffusion transformers. Across a suite of self-contained experiments (Table 1), HyperQuant outperforms th…

  1103. Hugging Face Daily Papers TIER_1 English(EN) ·

    Abstract representational geometry supports inference in large language models

    A defining feature of human intelligence is the ability to adapt to changing environments by inferring latent task structure from sparse observations. Neuroscientific research indicates that this capability relies on the hippocampus constructing abstract representations, expresse…

  1104. Hugging Face Daily Papers TIER_1 English(EN) ·

    Interleaved Speech Language Models Latently Work In Text

    Interleaved speech-text language models exhibit an implicit transcription phase where text tokens become decodable in intermediate layers, followed by text-based prediction before speech domain transformation.

  1105. arXiv cs.AI TIER_1 English(EN) · Shao-Qun Zhang ·

    On the Expressive Power of Weight Quantization in Large Language Models

    In recent years, weight quantization that encodes the learnable parameters of large language models in an $n$-bit format has garnered significant attention due to its potential for model compression and inference acceleration. Many practical techniques have been developed; howeve…

  1106. Hugging Face Daily Papers TIER_1 English(EN) ·

    Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining

    Training-time data augmentation techniques help mitigate overfitting in autoregressive language model pretraining by delaying performance deterioration and improving final model quality when training on fixed datasets for many epochs.

  1107. arXiv stat.ML TIER_1 English(EN) · Dun Li Chan, Emily Liu, Niyathi Allu, Christian Hoang ·

    How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

    arXiv:2609.03322v1 Announce Type: cross Abstract: Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is usually evaluated only through output behavior. We study how six naturalistic and synthetic input perturbations propagate…

  1108. arXiv stat.ML TIER_1 English(EN) · Quang Hoang Trung, Quang Huu Hieu, Nguyen Van Hoang Phuc, Vo Nguyen Le Duy ·

    ALRA: Adaptive Local Relational Alignment for Logit-Based Pre-training Distillation of Autoregressive Language Models

    arXiv:2609.03355v1 Announce Type: new Abstract: Logit-based knowledge distillation for autoregressive language models usually aligns teacher and student next-token distributions over the entire vocabulary. However, this global objective overlooks relative preferences among likely…

  1109. arXiv stat.ML TIER_1 English(EN) · Xiao Shi Huang, Chen-Yuan Lin, Bruce Kuwahara, Kin Kwan Leung, Jesse C. Cresswell ·

    Unifying Conformal Language Tasks with In-Context Ensembles

    arXiv:2609.03005v1 Announce Type: cross Abstract: Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content from documents under two constraints: coverage, retaining enough pertinent information to achieve some goal, and concis…

  1110. LessWrong (AI tag) TIER_1 English(EN) · Shunk ·

    Notes on "A global workspace in language models"

    <p><span>Not a very original paper review, so here are my concise notes on the topic, and near the bottom some of my thoughts on future experiments that I think would be interesting. Part 3 of notes series!</span></p><p><span>Important note: This post was written based on the Ant…

  1111. arXiv cs.CV TIER_1 English(EN) · Daizong Liu, Junhao Dong, Zhiyuan Ma, Xiaoye Qu, Xiang Fang, Runwei Guan, Keke Tang, Jianfeng Dong, Yew-Soon Ong ·

    Forbid Your Attention: Fooling Multimodal Large Language Models by Selectively Removing Intrinsic Focus in Spectral Domain

    arXiv:2609.00788v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have extended the capability of large language models (LLMs) to process more contextual multimodal information, showing remarkable progress in diverse realistic multimodal applications. Despi…

  1112. arXiv stat.ML TIER_1 English(EN) · Marco Simnacher, Georg Keilbar, Benjamin K\"onig, Christoph Lippert, Sonja Greven ·

    Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches

    arXiv:2609.00946v1 Announce Type: new Abstract: Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ given a third random object $Z$. Existing CITs have limited applicability to high-dimensional data, especially multimodal d…

  1113. arXiv stat.ML TIER_1 English(EN) · Yicheng Mao, Hongru Du ·

    Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining

    arXiv:2608.23922v1 Announce Type: cross Abstract: Data mixing is a central design problem in large language model pretraining: given a fixed token budget, practitioners must decide how much data to allocate to each domain. Recent proxy-based methods address this problem by traini…

  1114. arXiv stat.ML TIER_1 English(EN) · Ricardo Baptista, Andrew Stuart, Son Tran ·

    Large Language Models: A Mathematical Formulation

    arXiv:2601.22170v2 Announce Type: replace-cross Abstract: Large language models (LLMs) process and predict sequences containing text to answer questions, and address tasks including document summarization, providing recommendations, writing software and solving quantitative probl…

  1115. arXiv stat.ML TIER_1 English(EN) · Hanti Lin ·

    Cross-Entropy Risk Estimation for Language Models: Inconsistency Must Be Dense, and the Holdout Method Is No Exception

    arXiv:2608.15798v1 Announce Type: cross Abstract: Language models are compared by their held-out per-token cross-entropy risk---the quantity scaling laws are fitted to. We show that it cannot be consistently estimated. Consistency, or convergence to the estimand, is defined relat…

  1116. arXiv stat.ML TIER_1 English(EN) · Zhi Zhang, Lingfeng Lyu, Yue Kang, Doudou Zhou ·

    Conditional Evaluation of Language Models with Cheap Auxiliary Signals

    arXiv:2608.16210v1 Announce Type: cross Abstract: Aggregate accuracy hides where models succeed and fail. Estimating conditional performance profiles from gold labels alone is expensive, while cheap auxiliary signals such as LLM-judge scores, pairwise comparisons, confidence scor…

  1117. LessWrong (AI tag) TIER_1 English(EN) · ashesfall ·

    Study Update: Does post-training quantization change welfare-relevant indicators in open-weight language models?

    <p><i><span>This content will not make sense without reading the preregistration found in </span></i><a href="https://www.lesswrong.com/posts/hrwKDeFFvQppFXHtr/does-post-training-quantization-change-welfare-relevant" rel="noreferrer"><i><span>the original post</span></i></a></p><…

  1118. arXiv cs.CV TIER_1 English(EN) · Wanshun Su, Yang Shi, Feihu Liu, Ziwen Yu, Yan Min, Zhuoran Zhang, Qixun Wang, Haotian Wang, Shixuan Liu, Yuanxing Zhang, Peng Wu, Chengfu Huo, Liang Ding ·

    OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

    arXiv:2608.03812v1 Announce Type: new Abstract: Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio token sequences incurs substantial computational overhe…

  1119. arXiv cs.CV TIER_1 English(EN) · Yuxin Cao, Wei Song, Jingling Xue, Jin Song Dong ·

    Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Models

    arXiv:2608.03160v1 Announce Type: cross Abstract: When asked which of two events came first, video large language models can fail in two opposite ways: cave to a false claim, or reject a true one. Prior video sycophancy work measures only the first and mitigates it by teaching th…

  1120. LessWrong (AI tag) TIER_1 English(EN) · Failfinder70 ·

    Friend-Shaped Objects: A Case Study on My Use of Large Language Models

    <p><span>If you’ve ever done a six-month-plus work/deployment contract, then you know the loneliness and homesickness will hit hard, either right away or weeks later. My own worry was forgetting where I was from and becoming someone else. A chatbot gave me a cringe but effective …

  1121. arXiv cs.CV TIER_1 English(EN) · Wenjie Zhu, Yabin Zhang, Wenjun Zeng, Lei Zhang ·

    MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models

    arXiv:2607.27637v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable MLLM should refuse truly out-of-context (OOC) questio…

  1122. arXiv cs.CV TIER_1 English(EN) · Shuyao Shang, Yuqi Wang, Ruopeng Gao, Xu Chen, Tieniu Tan, Lue Fan, Zhaoxiang Zhang ·

    PhiZero: A World Model Built Around Physical Language

    arXiv:2607.28624v1 Announce Type: new Abstract: We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leavi…

  1123. arXiv cs.CV TIER_1 English(EN) · Jie Ma, Zhike Qiu, Jie Gao, Jiayi Ji, Qian Chen, Xiaoshuai Sun, Rongrong Ji ·

    Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models

    arXiv:2607.28341v1 Announce Type: new Abstract: While visual token pruning is essential for efficient Multimodal Large Language Models (MLLMs), existing training-free methods suffer from a critical limitation: they rely on static, instantaneous heuristics to perform irreversible …

  1124. arXiv cs.CV TIER_1 English(EN) · Yuchen Wang, Qihui Zhu, Yang Liu, Xiaoyan Sun, Siying Wu ·

    SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models

    arXiv:2607.25818v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs), such as Qwen2.5-VL and InternVL3, generate large numbers of vision tokens for high-resolution inputs, leading to substantial computational cost. Existing vision token pruning methods …

  1125. arXiv cs.CV TIER_1 English(EN) · Qihui Zhu, Yuchen Wang, Zijian Wen, Tao Zhang, Mengjie Zhang, Yang Liu, Shuangwu Chen, Siying Wu, Jian Yang, Xiaofeng Jiang ·

    RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models

    arXiv:2607.24447v1 Announce Type: new Abstract: On-Policy Self-Distillation (OPSD) uses privileged information available only to the teacher to provide dense token-level supervision on trajectories generated by the student. However, existing methods often rely on verified solutio…

  1126. arXiv stat.ML TIER_1 English(EN) · Matthew ffrench-Constant, Daniel Yang, Xinmeng Huang, Sanyam Kapoor ·

    ConfidenceBench: Evaluating Confidence Calibration in Large Language Models

    arXiv:2607.20526v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly. In these settings, accuracy alone is insufficient: models must also know when they are likely to be wrong. We pre…

  1127. arXiv cs.CV TIER_1 English(EN) · Yue Zhao, Hongxu Liu, Feiyu Wang, Xiaoyu Yang, Tong Ge, Zhen Yang, Chao Wang, Qiong Zeng ·

    MV-Bench: Benchmarking Multimodal Large Language Models for Coordinated Multi-View Interface Construction

    arXiv:2607.19910v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly expected to automate visualization development by generating code directly from visual designs. However, existing evaluations mainly focus on single-chart generation and over…

  1128. arXiv stat.ML TIER_1 English(EN) · Chia-Yu Hsu, Shubhanshu Shekhar ·

    Efficient Sequential Evaluation of Large Language Models

    arXiv:2607.17409v1 Announce Type: new Abstract: We study the problem of sequentially evaluating a new large language model (LLM) on a fixed question set using historical performance data from prior LLMs. Our goal is to construct a confidence sequence (CS) for the model's capabili…

  1129. arXiv stat.ML TIER_1 English(EN) · Matthew F. Dixon ·

    Calibrating Semantic Uncertainty from Observable Language-Model Probabilities

    arXiv:2607.17447v1 Announce Type: cross Abstract: Language models produce probabilities over words, but professional decisions require uncertainty over meaningful states such as diagnoses, hypotheses or operational conditions. A model's printed numerical confidence does not estab…

  1130. arXiv stat.ML TIER_1 English(EN) · Zhiyu Xu, Jia Liu, Yixin Wang, Yuqi Gu ·

    Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length

    arXiv:2512.07019v3 Announce Type: replace-cross Abstract: The proliferation of Large Language Models (LLMs) necessitates valid evaluation methods to provide guidance for both downstream applications and actionable future improvements. The Item Response Theory (IRT) model with Com…

  1131. arXiv stat.ML TIER_1 English(EN) · Florian Valade ·

    Accelerating Large Language Model Inference with Self-Supervised Early Exits

    arXiv:2407.21082v3 Announce Type: replace-cross Abstract: This paper presents a modular approach to accelerate inference in large language models (LLMs) by adding early exit heads at intermediate transformer layers. Each head is trained in a self-supervised manner to mimic the ma…

  1132. arXiv stat.ML TIER_1 English(EN) · Bruno Abrahao ·

    Neural Collapse Is Forbidden: Information Floors in Language Models

    arXiv:2607.09487v1 Announce Type: cross Abstract: Within-class variance in language-model representations is commonly read as incomplete neural collapse. We argue it is allocated information storage, and that the allocation obeys a law. A one-line centering identity voids a famil…

  1133. LessWrong (AI tag) TIER_1 Dansk(DA) · Michele Campolo ·

    Independent alignment of language models

    <blockquote><p><span>The user could write up the metaethical argument — the one developed in Part One, refined — and submit it as feedback to Anthropic, publish it, or engage with researchers working on AI alignment and values. The probability that any single submission changes t…

  1134. arXiv stat.ML TIER_1 English(EN) · Bruno Abrahao ·

    Neural Collapse Is Forbidden: Information Floors in Language Models

    Within-class variance in language-model representations is commonly read as incomplete neural collapse. We argue it is allocated information storage, and that the allocation obeys a law. A one-line centering identity voids a family of simplex equiangular-tight-frame claims, inclu…

  1135. arXiv cs.CV TIER_1 English(EN) · Yulin Luo, Chun-Kai Fan, Menghang Dong, Jiayu Shi, Xiangju Mi, Mengdi Zhao, Bo-Wen Zhang, Cheng Chi, Jiaming Liu, Gaole Dai, Rongyu Zhang, Ruichuan An, Kun Wu, Zhengping Che, Shaoxuan Xie, Guocai Yao, Zhongxia Zhao, Pengwei Wang, Guang Liu, Zhongyuan Wan… ·

    Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain

    arXiv:2510.17801v2 Announce Type: replace-cross Abstract: Building robots that can perceive, reason, and act in dynamic, unstructured environments remains a central challenge. Recent embodied systems often follow a dual-system paradigm, where System 2 performs high-level reasonin…

  1136. LessWrong (AI tag) TIER_1 English(EN) · wesg ·

    A global workspace in language models

    <p><i><span>[This is the blog post for our new paper </span></i><a href="https://transformer-circuits.pub/2026/workspace/index.html" rel="noreferrer"><i><span>Verbalizable Representations Form a Global Workspace in Language Models</span></i></a><br /><i><span>Readers might also b…

  1137. arXiv stat.ML TIER_1 English(EN) · Dan Ley, Giang Nguyen, Himabindu Lakkaraju, Julius Adebayo ·

    Prototype Language Models

    arXiv:2607.00510v1 Announce Type: cross Abstract: Knowing which training examples drive outputs is fundamental to auditing, correcting, and understanding language models, yet for modern LLMs this remains expensive, approximate, and largely post-hoc. Standard language models gener…

  1138. arXiv cs.CV TIER_1 English(EN) · Zhi Tu, Liangkun Niu, Tianyi Zhang ·

    TrafficAlign: Aligning Large Language Models for Traffic Scenario Generation

    arXiv:2606.29097v1 Announce Type: new Abstract: Recent research has investigated the use of large language models (LLMs) to generate traffic scenarios for autonomous driving. However, pretrained LLMs often fail to align with real-world traffic distributions. In this work, we pres…

  1139. arXiv cs.CV TIER_1 English(EN) · Liyuan Deng, Hao Guo, Yongkang Dai, Yunpeng Bai, Yifan Zhu, Yuanyuan Gao, Huaxi Huang, Yilei Shi ·

    BrepLLM: Enabling Large Language Models to Understand Boundary Representations

    arXiv:2512.16413v2 Announce Type: replace Abstract: Current token-sequence-based Large Language Models (LLMs) struggle to directly process 3D Boundary Representation (B-rep) models that contain complex geometric and topological information. To this end, we propose BrepLLM, the fi…

  1140. arXiv cs.CV TIER_1 English(EN) · Thomas De Min, Subhankar Roy, St\'ephane Lathuili\`ere, Elisa Ricci, Massimiliano Mancini ·

    ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models

    arXiv:2603.19466v2 Announce Type: replace Abstract: Effective collaboration begins with knowing when to ask for help. For example, when trying to identify an occluded object, a human would ask someone to remove the obstruction. Can MLLMs exhibit a similar "proactive" behavior by …

  1141. arXiv stat.ML TIER_1 English(EN) · Kwok Chun Au, Adam Block ·

    Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models

    arXiv:2606.25086v1 Announce Type: cross Abstract: Many modern Language Model (LM) pipelines return an averaged model, such as an exponential moving average of the training iterates, rather than the final iterate itself. This raises a fundamental question: given that we will retur…

  1142. arXiv cs.CV TIER_1 English(EN) · Zhihao Zhu, Hongyi Tang, Yi Yang, Ahmed Abbasi ·

    Revealing Training Data Exposure in Vision Language Large Models via Parameter Gradients

    arXiv:2606.24774v1 Announce Type: new Abstract: Vision-Language Large Models (VLLMs) trained on massive crawled corpora raise pressing copyright and data-provenance concerns. These concerns are particularly acute in healthcare, where patient medical images paired with clinical re…

  1143. arXiv cs.CV TIER_1 English(EN) · Zengjie Chen, Yuxiang Cai, Jingcai Guo, Taotao Cai, Jianwei Yin, Zhi Chen ·

    Accelerating Multimodal Large Language Models with Prior-Corrected Token Reduction

    arXiv:2606.24156v1 Announce Type: new Abstract: Visual token reduction has emerged as an effective strategy for accelerating Multimodal Large Language Models (MLLMs). Many existing methods prune tokens by ranking text-visual attention scores. However, we show that attention is of…

  1144. arXiv cs.CV TIER_1 English(EN) · Bin Chen, Yuxiang Cai, Yadan Luo, Yi Zhang, Jianwei Yin, Zhi Chen ·

    Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models

    arXiv:2606.24165v1 Announce Type: new Abstract: Reducing visual token redundancy is critical for accelerating Multimodal Large Language Models (MLLMs) without degrading cross-modal reasoning performance. Existing token pruning methods typically rely on single-layer signals, such …

  1145. arXiv stat.ML TIER_1 English(EN) · Adam Block ·

    Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models

    Many modern Language Model (LM) pipelines return an averaged model, such as an exponential moving average of the training iterates, rather than the final iterate itself. This raises a fundamental question: given that we will return an iterate average, how should we change trainin…

  1146. arXiv cs.CV TIER_1 English(EN) · Ahmed Abbasi ·

    Revealing Training Data Exposure in Vision Language Large Models via Parameter Gradients

    Vision-Language Large Models (VLLMs) trained on massive crawled corpora raise pressing copyright and data-provenance concerns. These concerns are particularly acute in healthcare, where patient medical images paired with clinical reports demand rigorous privacy safeguards. Howeve…

  1147. 36氪 (36Kr) TIER_1 中文(ZH) ·

    Lingguang App Upgrades World Model Experience

    36氪获悉,7月9日,灵光App正式升级了世界模型体验功能。新版接入蚂蚁灵波最新发布的世界模型LingBot-World 2.0相关能力。相比此前只能“走一走、看一看”,新版最大的变化是让AI生成的世界真正“可玩”起来。用户不仅能在场景中漫步探索,还可以像玩3D游戏一样进行技能操作,AI会根据用户动作实时生成新的场景反馈。

  1148. Engadget TIER_1 English(EN) · [email protected] (Karissa Bell) ·

    Meta's new AI transcription model can distinguish between multiple speakers and languages in real-time

    The latest release from Meta Superintelligence Lab is a powerful transcription model.

  1149. Forbes — Innovation TIER_1 English(EN) · Rivindu Perera, Forbes Councils Member ·

    David Vs. Goliath: Why Small Language Models Are Quietly Winning Where It Matters Most

    In enterprise AI, the models that win won't be the largest. They'll be the ones that know exactly what they're built for.​​

  1150. HN — anthropic stories TIER_1 English(EN) · in-silico ·

    A global workspace in language models

  1151. Forbes — Innovation TIER_1 English(EN) · Joe Toscano, Contributor ·

    Small Language Models Outperform Frontier AI On Cost, Speed And Accuracy

    Bigger has defined AI from day one. New data says task-specific small models beat frontier LLMs on accuracy, cost and speed — and save money.

  1152. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

    <p>Cohere has released North Small Translate, an open-weight Mixture-of-Experts model built for machine translation across 50 languages. It uses 25B of its 218B parameters per token and scores 83.6 on Cohere's WMT26 evaluation. Weights are free for non-commercial use, with commer…

  1153. Lobsters — AI tag TIER_1 English(EN) · laya.convaiinnovations.com via elbart ·

    Laya — 33ms Multilingual System 1 Decision Engine

    <p><a href="https://lobste.rs/s/ojukrw/laya_33ms_multilingual_system_1_decision">Comments</a></p>

  1154. Medium — MLOps tag TIER_1 Türkçe(TR) · Aleyna Altunsu ·

    LLMOps: Managing Large Language Models in Production Part 1:

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@aleynaaltunsu/llmops-%C3%BCretim-ortam%C4%B1nda-b%C3%BCy%C3%BCk-dil-modellerini-y%C3%B6netmek-b%C3%B6l%C3%BCm-1-f315a2b9ac2a?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/…

  1155. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    🧠 A study examines how effectively language models actually utilize their memory capacity during inference. The research measures whether models retrieve and ap

    🧠 A study examines how effectively language models actually utilize their memory capacity during inference. The research measures whether models retrieve and apply stored information or rely on other mechanisms when processing tasks. 💬 Hacker News 🔗 https://www. liquid.ai/blog/ho…

  1156. dev.to — MCP tag TIER_1 English(EN) · Programming Central ·

    Why Large Language Models Fail at Scheduling (And How Constraint Satisfaction Problems Fix It)

    <p>If you have ever asked a Large Language Model to build a complex work schedule, allocate cloud compute nodes, or organize hospital operating rooms, you have likely run into a frustrating wall. The LLM might generate a neat table, but upon closer inspection, Dr. Smith is assign…

  1157. Medium — fine-tuning tag TIER_1 English(EN) · Mohammed Saqlain ·

    Fine-Tune Language Models With One YAML File: A Practical Guide to TRLoom

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/data-science-collective/fine-tune-language-models-with-one-yaml-file-a-practical-guide-to-trloom-6a0866a36578?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1024/1…

  1158. Towards AI TIER_1 English(EN) · Tejas Nidhankar ·

    Mastering LangChain: Open-Source Models & Prompt Engineering (Part 2)

    <p>In Part 1, we established the high-level roadmap of LangChain — exploring the architectural shift toward chat models and breaking down the foundational RAG pipeline.</p><p>Now, we dive directly into code implementation. Before assembling end-to-end chains or building multi-ste…

  1159. Towards AI TIER_1 English(EN) · Shanehobson ·

    A Software Engineer’s Guide to Large Language Models

    <p>The introduction of large language models (LLMs) has caused a tectonic shift in how software is built. When combined with agentic coding harnesses, LLMs can now perform many of the tasks once thought to be the exclusive domain of human software engineers. These tools can write…

  1160. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    *NeoMME*: an efficient Multimodal-native and Multilingual Encoder https://huggingface.co/blog/Hcompany/neomme # AI # MachineLearning # NLP

    *NeoMME*: an efficient Multimodal-native and Multilingual Encoder https://huggingface.co/blog/Hcompany/neomme # AI # MachineLearning # NLP

  1161. Towards AI TIER_1 English(EN) · Md Johirul Islam ·

    vLLM: The Intuitive Guide to Serving Large Language Models at High Speed

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/vllm-the-intuitive-guide-to-serving-large-language-models-at-high-speed-c05fb67a06a3?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1536/1*UcuT1i3yJmliM1Cv…

  1162. Medium — fine-tuning tag TIER_1 English(EN) · Saad Kabir ·

    Fine-Tuning Large Language Models with LoRA: A Practical Guide

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@pmskabir123/fine-tuning-large-language-models-with-lora-a-practical-guide-fc1549997c8b?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/2600/1*22717jiytX6WI-fVtGs9_…

  1163. Towards AI TIER_1 English(EN) · Sanjana Dubey ·

    The Complete Guide to Model Inference: Every Time You Type a Prompt, This Is What Happens

    <h4><em>A deep, honest, research-backed walkthrough of the most important process in modern AI — one that runs billions of times a day yet almost no one fully understands.</em></h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*jXyRRgweLHi5H70c3WKmvA.png" /><…

  1164. Medium — MLOps tag TIER_1 English(EN) · Jyotirmay Mal ·

    Demystifying Large Language Models: How ChatGPT and Gemini Actually Work

    <div class="medium-feed-item"><p class="medium-feed-snippet">Every time you type something into ChatGPT or Gemini it feels like you are talking to a person.. Behind the scenes there is a lot of math&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@jyotirmaymal0…

  1165. Medium — fine-tuning tag TIER_1 English(EN) · Tech Horizon With Anand Vemula ·

    Adaptive AI: Fine-Tuning and Few-Shot Learning for Next-Generation Large Language Models

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@anandvlinkedin/adaptive-ai-fine-tuning-and-few-shot-learning-for-next-generation-large-language-models-c25fe346b4ee?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max…

  1166. Medium — fine-tuning tag TIER_1 English(EN) · Vardhini Sirla ·

    I Fine-Tuned a Language Model in 48 Seconds

    <div class="medium-feed-item"><p class="medium-feed-snippet">No, seriously. 48 seconds.</p><p class="medium-feed-link"><a href="https://medium.com/@vardhinisirla/i-fine-tuned-a-language-model-in-48-seconds-edcb15e1d453?source=rss------fine_tuning-5">Continue reading on Medium »</…

  1167. Mastodon — sigmoid.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @0xSero: Inference Engineer, a fantastic foundational resource for learning how to operate Large Language Models. Baseten has excellent contributions

    RT @0xSero: Inference Engineer, eine fantastische grundlegende Ressource zum Erlernen des Betriebs von Large Language Models. Baseten hat hervorragende Beiträge zur Open-Source-Community geleistet, ich habe diesen Beitrag wirklich geliebt. 10 Personen, die diesen Beitrag liken, e…

  1168. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    Curious what actually happens inside a Large Language Model? 🤖 Barry Burd explores how data flows through neural networks, how models learn, and how simple inpu

    Curious what actually happens inside a Large Language Model? 🤖 Barry Burd explores how data flows through neural networks, how models learn, and how simple inputs become useful predictions, all through practical examples and visual explanations. https://www. dev2next.com/speaker/…

  1169. Medium — fine-tuning tag TIER_1 English(EN) · Joel Wembo ·

    Training and Fine‑Tuning Large Language Models on VxCloud

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://joelwembo.medium.com/training-and-fine-tuning-large-language-models-on-vxcloud-cd3fe0681c5f?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1920/1*t0hyJqKVcdvy_J7Gi5GzCQ.png" …

  1170. Medium — fine-tuning tag TIER_1 English(EN) · D Kidhu ·

    Tülu 3: A Fully Open Recipe for Post-Training Large Language Models

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@d.kidhu/t%C3%BClu-3-a-fully-open-recipe-for-post-training-large-language-models-75b9bf086964?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/2000/1*luf36T0ZpECM0jY…

  1171. Medium — Claude tag TIER_1 English(EN) · Swapnilnicolsondadel ·

    What Is a Large Language Model (LLM)? A Beginner's Guide

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@swapnilnicolsondadel/what-is-a-large-language-model-llm-a-beginners-guide-6ac9558e2738?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1536/1*dM4U3wHN699vJwjfznJqfQ.png…

  1172. Lobsters — AI tag TIER_1 English(EN) · hiraditya.github.io via jado ·

    A tour of MLIR: The Dialect Stack Everyone Depends On

    <p><a href="https://lobste.rs/s/o9vjlt/tour_mlir_dialect_stack_everyone_depends">Comments</a></p>

  1173. Medium — MLOps tag TIER_1 English(EN) · Shweta Shrivastava ·

    Did It Actually Work? Evaluating NLP Models

    <div class="medium-feed-item"><p class="medium-feed-snippet">This is the fifth article in the series. We&#x2019;ve turned text into numbers, given words context with attention, searched meaning with&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@shweta.shriva…

  1174. Medium — Claude tag TIER_1 English(EN) · Rubiya Khan ·

    ChatGPT, Claude, Gemini :How Large Language Models Are Changing Data Science

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@rubikhan1411/chatgpt-claude-gemini-how-large-language-models-are-changing-data-science-db0ae4193680?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1408/1*07o8uBupBinI2…

  1175. Towards AI TIER_1 English(EN) · Neha Khan • AI & Software Engineer ·

    Understanding Large Language Models: From Neural Networks to Production Inference

    <h4>Why I Started This Journey</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*HAwda1fsKo65Uik6S9V6Kw.png" /></figure><p>As a software engineer, my first experience with Large Language Models was the same as most developers — I called an API, sent a prompt…

  1176. Medium — fine-tuning tag TIER_1 Bahasa(ID) · Ichlasridho ·

    BERT: Pre-training and Fine-tuning

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ichlasridho44/bert-pre-training-dan-fine-tuning-1046c94afc64?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1908/1*b0MAMmPiBvE2PV0Hgj6OmA.png" width="1908" /></a>…

  1177. Medium — fine-tuning tag TIER_1 English(EN) · Tech Horizon With Anand Vemula ·

    Adaptive AI: How Fine-Tuning and Few-Shot Learning Are Transforming Large Language Models

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@anandvlinkedin/adaptive-ai-how-fine-tuning-and-few-shot-learning-are-transforming-large-language-models-b1041b736eb5?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/ma…

  1178. Medium — Claude tag TIER_1 English(EN) · Hayanan ·

    Surviving the Tectonic Shifts in Large Language Model Scaling: A Field Guide for Practitioners

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/surviving-the-tectonic-shifts-in-large-language-model-scaling-a-field-guide-for-practitioners-3c967665db08?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1228/0*…

  1179. dev.to — MCP tag TIER_1 English(EN) · Intellibooks AI ·

    Intellibooks LLM Architecture Explained: A Complete Guide to How Large Language Models Work

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk5s1so5arglxajb5k42i.jpg"><img alt=" " height="1200"…

  1180. Towards AI TIER_1 English(EN) · Utkarsh Mittal ·

    Multihead Attention: How One Model Reads a Sentence 144 Ways

    <h4>Walkthrough of the last piece of self-attention—12 heads per layer and the two matrices that hold it all together.</h4><p>Article 1- <a href="https://medium.com/p/5807f169db30">https://medium.com/p/5807f169db30</a></p><p>Article 2 — <a href="https://medium.com/p/75df74bb5f07"…

  1181. Medium — fine-tuning tag TIER_1 English(EN) · Vahe Andonians ·

    Why Fine-Tuned Language Models Can Know a Fact Without Using It

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@vaheandonians/why-fine-tuned-language-models-can-know-a-fact-without-using-it-77f527eaf352?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1536/1*4h4sPVAf7tjNPyzlV…

  1182. Medium — fine-tuning tag TIER_1 English(EN) · elyas mosayebi ·

    Why We Fine-Tuned a Local LLM for Personalized Language Learning

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@elyas.mosayebi/why-we-fine-tuned-a-local-llm-for-personalized-language-learning-98bd11914b30?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1024/1*4UObguF4AVlUI1y…

  1183. Medium — fine-tuning tag TIER_1 (CA) · Namrata Gaddameedi ·

    Fine-Tuning Large Language Models: A Practical Guide

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@namrata.gaddameedi414/fine-tuning-large-language-models-a-practical-guide-0fe08874954e?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1172/1*Q3xzjcehQ410rJbJF-_bB…

  1184. Medium — MLOps tag TIER_1 English(EN) · Himesh Ray ·

    How Large Language Models (LLMs) Work: A Beginner-Friendly Guide with Real-World Examples

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@himeshray1997/how-large-language-models-llms-work-a-beginner-friendly-guide-with-real-world-examples-455d6ebd3834?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1536/1*…

  1185. Medium — MLOps tag TIER_1 English(EN) · Himesh Ray ·

    Understanding Large Language Models (LLMs) Through Real-World Examples

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@himeshray1997/understanding-large-language-models-llms-through-real-world-examples-c106920f251f?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1536/1*Q-d5wHjgobRE32847E…

  1186. Medium — MLOps tag TIER_1 English(EN) · Piotr Kalanski ·

    Global Vocabulary Management: The Quiet Bug That Breaks ML Models at Retraining Time

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@piotr.kalanski/global-vocabulary-management-the-quiet-bug-that-breaks-ml-models-at-retraining-time-5ef61171978b?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1408/1*RW…

  1187. Medium — Claude tag TIER_1 English(EN) · Dhruv Kumar ·

    Scaling Laws for Neural Language Models Full Breakdown

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@rafftarsingh7982/scaling-laws-for-neural-language-models-full-breakdown-a27f622b8312?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1536/1*7VU2RQddufPFAS5QJrKazg.png" …

  1188. Medium — MLOps tag TIER_1 English(EN) · Mohsen Kheirandishfard ·

    Why Quantized Models Break, and How Quantize-Aware Training Fixes Them

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@mohsen.kheirandishfard/why-quantized-models-break-and-how-quantize-aware-training-fixes-them-6d57d202d7d3?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1400/1*a3y86yWX…

  1189. r/LocalLLaMA TIER_1 Italiano(IT) · /u/Balance- ·

    convaiinnovations/laya (multilingual, non-autoregressive System 1 decision model)

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wm0cn0/convaiinnovationslaya_multilingual/"> <img alt="convaiinnovations/laya (multilingual, non-autoregressive System 1 decision model)" src="https://external-preview.redd.it/jvxN3-dnP0ONagNz6HLgL-xfARRDV2e0…

  1190. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Canto: A speech model built for the real world https:// wisprflow.ai/canto # ai

    Canto: A speech model built for the real world https:// wisprflow.ai/canto # ai

  1191. dev.to — LLM tag TIER_1 English(EN) · Sugan Raja ·

    Exploring GPT‑Astra: The Next Leap in Large‑Language‑Model Innovation

    <h1> Exploring GPT‑Astra: The Next Leap in Large‑Language‑Model Innovation </h1> <h2> Introduction </h2> <p>Artificial‑intelligence research has been on a relentless sprint toward ever‑larger, more capable language models. Among the newest entrants, <strong>GPT‑Astra</strong> has…

  1192. dev.to — LLM tag TIER_1 English(EN) · Piethein Strengholt ·

    Small Language Models in Practice: What I Learned from Testing Qwen and ModernBERT in RSSMonster

    <p>I have always liked RSS because it gives me a easy way to decide for myself which sources I want to follow, without relying on social networks or (commercial) recommendation platforms to decide what appears in front of me. I want to be in control of the content I see myself.</…

  1193. dev.to — LLM tag TIER_1 English(EN) · PRANJUL RATHOUR ·

    Small language models (1B–3B): what they're good for after fine-tuning

    <p>A 1.7B-parameter model will not write your thesis. It will, after fine-tuning, classify ten thousand support tickets an hour on a CPU, for free, without sending customer data anywhere. FineTune Studio is built around small models because that trade is the one most teams actual…

  1194. r/LocalLLaMA TIER_1 English(EN) · /u/sachasayan ·

    H3-World: Turning Language Understanding into World Control

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w5akpy/h3world_turning_language_understanding_into_world/"> <img alt="H3-World: Turning Language Understanding into World Control" src="https://external-preview.redd.it/NjdpdDFsNXQwNG5oMerGSIU-HQEebKy_YuAHc86…

  1195. dev.to — LLM tag TIER_1 English(EN) · Pneumetron ·

    The Abliteration Wave: Removing Refusals from Large Language Models

    <p><em>The open-weights community is rapidly adopting 'abliteration' techniques to strip safety-aligned models of their refusal behaviors. This shift is exemplified by the massive popularity of modified Qwen 3.8-27B variants, signaling a move toward unconstrained model access.</e…

  1196. dev.to — LLM tag TIER_1 English(EN) · Vishal Chandak ·

    From Massive to Miniature: How Small Language Models Are Engineered

    <p>In Part 1, we started with a simple question:</p> <blockquote> <p><strong>Does every AI task need the biggest model we can afford?</strong></p> </blockquote> <p>Often, the answer is no.</p> <p>A model that is considerably smaller than a frontier model can be the better choice …

  1197. r/LocalLLaMA TIER_1 English(EN) · /u/sdroege_ ·

    LittleLearner: Language Models Under Pedagogically-Controlled Knowledge Exposure

    <!-- SC_OFF --><div class="md"><blockquote> <p>Modern LMs are trained on everything at once, so it is hard to tell whether a new skill was learned or merely elicited. We constrain the training distribution itself: an 88B-token corpus filtered to the U.S. elementary-school curricu…

  1198. dev.to — LLM tag TIER_1 English(EN) · Aviral Srivastava ·

    Introduction to LLMs (Large Language Models)

    <h2> The Mind Readers of the Digital Age: An In-Depth, Casual Intro to Large Language Models (LLMs) </h2> <p>Ever feel like your phone's autocomplete is a little <em>too</em> good, predicting your next word with uncanny accuracy? Or perhaps you've marveled at how quickly ChatGPT …

  1199. dev.to — LLM tag TIER_1 English(EN) · Bato ·

    From GPT-2 to Kimi K3: How Language Models Learned to Manage Memory

    <p><em>Six architectures, one idea: from the KV cache to Kimi K3's hybrid of recurrent memory, global attention, and sparse experts, each a smarter way to manage a limited memory.</em></p> <p><strong>TL;DR:</strong> In 2019, GPT-2 (small) had 124 million parameters. In 2026, Kimi…

  1200. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Super Study Guide: Transformers & Large Language Models by Afshine Amidi and Shervine Amidi 📖 on Leanpub! A clear, illustrated guide to large language models, c

    Super Study Guide: Transformers & Large Language Models by Afshine Amidi and Shervine Amidi 📖 on Leanpub! A clear, illustrated guide to large language models, covering key concepts and practical applications. Ideal for projects, interviews, or personal learning. Link: https:// le…

  1201. dev.to — LLM tag TIER_1 English(EN) · Jorge Tovar ·

    The Best Model Isn’t Enough: Harnesses, Context, and Better Prompts

    <p>We used to believe that a better model automatically meant better results. And that’s true to some degree, but the reality is that most people aren’t using all the available tools and techniques to get the most out of the model.</p> <p>For example, context management is key. P…

  1202. dev.to — LLM tag TIER_1 English(EN) · Multigrid ·

    What Is a Large Language Model? A Working Definition

    <p>Almost every definition of a large language model describes its training. That is history. If you are calling one through an API, the useful definition is about what it does when your request arrives.</p> <h2> The definition </h2> <p>A large language model is a function. It ta…

  1203. dev.to — LLM tag TIER_1 English(EN) · Vishal Chandak ·

    The Small Language Model Revolution: Why Fit Beats Force

    <p>For the last few years, AI engineering has operated under a surprisingly simple assumption:</p> <blockquote> <p>Bigger models are better models.</p> </blockquote> <p>More parameters. More training data. More GPUs. More compute.</p> <p>And to be fair, that strategy has worked r…

  1204. r/LocalLLaMA TIER_1 English(EN) · /u/Unusual_Shoe2671 ·

    Luth-2: New State-of-the-Art French Small Language Models

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vlbto8/luth2_new_stateoftheart_french_small_language/"> <img alt="Luth-2: New State-of-the-Art French Small Language Models" src="https://preview.redd.it/3sgce32djpih1.png?width=640&amp;crop=smart&amp;auto=we…

  1205. r/LocalLLaMA TIER_1 English(EN) · /u/-p-e-w- ·

    Lophius: A workbench for language model research, from the creator of Heretic

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vjt4vi/lophius_a_workbench_for_language_model_research/"> <img alt="Lophius: A workbench for language model research, from the creator of Heretic" src="https://preview.redd.it/1d4eyj6ycdih1.png?width=140&amp;…

  1206. dev.to — LLM tag TIER_1 English(EN) · Rickesh T N ·

    Can a Cheap Model Beat a Frontier Model? Rebuilding Recursive Language Models with Codex

    <p>Large language models have enormous context windows now. That does not mean they use all of that context reliably.</p> <p>As prompts grow, models can miss details, lose track of relationships, or produce plausible summaries instead of doing the exhaustive work a question requi…

  1207. dev.to — LLM tag TIER_1 English(EN) · Multigrid ·

    The Long History of Machine Translation

    <p>Machine translation is the oldest continuously funded application of computing to language, the first field to be shut down by an evaluation report, and the field where the current architecture was invented. It also has the most abused numbers in AI, because BLEU scores from d…

  1208. r/MachineLearning TIER_1 English(EN) · /u/adam_alpha_finetuner ·

    The current state of language models and human preference based rankings [R]

    <!-- SC_OFF --><div class="md"><p>&quot;Arena ai&quot; has been a great success in producing a human preference based ranking, additional to other more objective benchmarks. However, this (probably) had also played a role in the syncopancy crisis and the general tendency of some …

  1209. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    Beyond Size: The Three Pillars of Test-Time Scaling in Large Language Models

    <h1> Beyond Size: The Three Pillars of Test-Time Scaling in Large Language Models </h1> <p>The narrative of artificial intelligence for the last decade has been dominated by a single, powerful trend: scaling. From the early days of AlexNet to the massive clusters powering GPT-4, …

  1210. dev.to — LLM tag TIER_1 English(EN) · Solon Framework ·

    Wiring a Model into SolonCode: Dialects, the apiUrl Rules, and the Timeouts Nobody Warns You About

    <p>SolonCode ships with no model. That is a deliberate product choice, not an omission: no bundled provider, no default key, no telemetry-backed endpoint you did not ask for. The upside is you can point it at DeepSeek, a local Ollama box, or an air-gapped corporate gateway with e…

  1211. r/LocalLLaMA TIER_1 English(EN) · /u/ttkciar ·

    [Paper] EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation

    <!-- SC_OFF --><div class="md"><p>The EdgeRazor method uses an entropy-guided distillation process to better translate a teacher model's logit probability distributions into the student model's low-bit / mixed-precision hidden-layer features, without attempting to preserve the te…

  1212. dev.to — LLM tag TIER_1 English(EN) · Marjia Afroj ·

    Evolution of Language Models - Every LLM Breakthrough Was Just a Bug Fix

    <p>Nobody sat down in 2017 and designed the Transformer from first principles.</p> <p>What actually happened is messier and much more interesting: someone built a thing, it broke in a specific and annoying way, someone else patched the break, and <em>that</em> patch broke somewhe…

  1213. dev.to — LLM tag TIER_1 English(EN) · Piyush Singh ·

    How a Large Language Model Actually Generates Text

    <p>Everything an AI agent does — writing code, answering support tickets, browsing the web — bottoms out in a single operation: a language model predicting the next token. Understand that one mechanism, and its limits, and the entire "AI agent" stack stops being magic and starts …

  1214. dev.to — LLM tag TIER_1 English(EN) · Maksim Sekretov ·

    ML Without Magic: Building a Tiny Language Model in Pure Node.js and Watching Every Weight Change

    <blockquote> <p>Tokenization → embeddings → causal Transformer → LM head → softmax → loss → backpropagation. No TensorFlow, no PyTorch, and no hidden autograd.</p> </blockquote> <p>Repository: <strong><a href="https://github.com/sekretov/tiny-language-model-neuro-js" rel="noopene…

  1215. dev.to — LLM tag TIER_1 English(EN) · Sanskriti Harmukh ·

    Deploying Large Language Models (LLMs) with Ollama

    <p>Ollama is an open-source platform for running large language models locally (Llama 3, DeepSeek R1, Mistral, Phi-4, Gemma 2, and more) with no internet connection or cloud API required. Running models locally improves privacy, security, and gives you full control over performan…

  1216. r/MachineLearning TIER_1 English(EN) · /u/BleedingXiko ·

    Seeking collaborators for scaling and independent evaluation of a new recurrent language model architecture (preprint + code) [R]

    <table> <tr><td> <a href="https://www.reddit.com/r/MachineLearning/comments/1uycffg/seeking_collaborators_for_scaling_and_independent/"> <img alt="Seeking collaborators for scaling and independent evaluation of a new recurrent language model architecture (preprint + code) [R]" sr…

  1217. dev.to — LLM tag TIER_1 English(EN) · Muhammad Zulqarnain ·

    Fine-Tuning Large Language Models: The Complete 2026 Guide

    <h2> Why Fine-Tune When You Have GPT-4? </h2> <p>GPT-4 is great at everything. So why fine-tune?</p> <p>Simple: Specificity beats generality.</p> <h2> Fine-Tuning Wins You: </h2> <p><strong>Better Performance</strong>: 10-30% accuracy improvements for your domain<br /> <strong>Lo…

  1218. dev.to — LLM tag TIER_1 English(EN) · Bahadir Kusat ·

    What Is Turkish-Language AI? Tokenizers, Training Data, and Language Model Development

    <p>A Turkish-language interface and an artificial intelligence system developed around the linguistic structure of Turkish are not the same thing. This article examines the difference in light of academic research.</p> <p>Turkish-language AI is not simply software with Turkish me…

  1219. dev.to — LLM tag TIER_1 English(EN) · Devanshu Biswas ·

    Perplexity: how surprised your language model is (and why lower is better)

    <p>Ask someone how "good" a language model is and you get vibes. Ask a researcher and you often get one number: perplexity. It is the oldest, cheapest, most-used way to score a language model, and once you see how it is built you never forget it.</p> <h2> Start with what the mode…

  1220. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    ReContext fixes LLM long-context blind spots without retraining A new arXiv preprint proposes a training-free method that helps language models actually use evi

    ReContext fixes LLM long-context blind spots without retraining A new arXiv preprint proposes a training-free method that helps language models actually use evidence already sitting inside their 128K-token context https://www. notatechguy.com/recontext-fixe s-llm-long-context-bli…

  1221. dev.to — LLM tag TIER_1 English(EN) · AIGrowthStack ·

    Using Anthropic’s J-Lens to Inspect What a Large Language Model Is About to Say

    <p>Anthropic’s latest mechanistic interpretability work is interesting for a very practical reason: it gives engineers a new way to inspect what a large language model is doing before a token is actually emitted.</p> <p>The company calls the technique the <strong>Jacobian lens</s…

  1222. dev.to — LLM tag TIER_1 中文(ZH) · zengbao yu ·

    LSEO Field Theory: A Non-Euclidean Dynamical Field Steady State Architecture for Large Language Models

    <h1> LSEO 场论:一种面向大语言模型的非欧几里得动力场持续状态架构 </h1> <h2> 摘要 </h2> <p>大语言模型(LLM)在每次推理调用中本质上是无状态的——对话上下文随推理结束而消失,跨会话的持续性依赖于外部存储机制而非模型内部状态的连续性。本文提出<strong>外置隐空间演化算子场(LSEO 场论)</strong>,将持续状态重新定义为非欧几里得流形上的动力场,形式化为 <code>(M, g, Φ)</code> 三元组,其中 M 为光滑流形,g 为黎曼度规,Φ 为驱动系统演化的动力场。工程上,我们在单 CPU 服务器(4核…

  1223. dev.to — LLM tag TIER_1 English(EN) · Haley ·

    Large Language Models Demystified: A Visual and Practical Guide

    <p>If you've been anywhere near the tech world in the past two years, you've heard the term "large language model" (LLM) thrown around constantly. But what actually is a large language model? How does it work? And why should you care?</p> <p>This guide breaks it down without the …

  1224. dev.to — LLM tag TIER_1 English(EN) · kongkong ·

    How I Explain Large Language Models to Non-Technical People

    <p>If you've been anywhere near the tech world in the past two years, you've heard the term "large language model" (LLM) thrown around constantly. But what actually is a large language model? How does it work? And why should you care?</p> <p>This guide breaks it down without the …

  1225. dev.to — LLM tag TIER_1 English(EN) · bestbee ·

    Large Language Models: What They Can Do, What They Can't, and Why It Matters

    <p>If you've been anywhere near the tech world in the past two years, you've heard the term "large language model" (LLM) thrown around constantly. But what actually is a large language model? How does it work? And why should you care?</p> <p>This guide breaks it down without the …

  1226. dev.to — LLM tag TIER_1 English(EN) · Alex Chen ·

    The Complete Beginner's Guide to Large Language Models (LLMs) in 2026

    <p>If you've been anywhere near the tech world in the past two years, you've heard the term "large language model" (LLM) thrown around constantly. But what actually is a large language model? How does it work? And why should you care?</p> <p>This guide breaks it down without the …

  1227. dev.to — LLM tag TIER_1 English(EN) · Sam Rivera ·

    LLMs Explained: How Large Language Models Actually Work Under the Hood

    <p>If you've been anywhere near the tech world in the past two years, you've heard the term "large language model" (LLM) thrown around constantly. But what actually is a large language model? How does it work? And why should you care?</p> <p>This guide breaks it down without the …

  1228. dev.to — LLM tag TIER_1 English(EN) · Devanshu Biswas ·

    Program-Aided Language Models: stop making the model do the math

    <p>Ask a language model how many times the letter "r" shows up in "strawberry" and there's a decent chance it says two. Ask it to compound $2,500 at 7% for eight years and it might quietly switch to simple interest and hand you $3,900 instead of $4,295.47. The reasoning sounds fi…

  1229. r/MachineLearning TIER_1 Italiano(IT) · /u/Idea_less_ ·

    Small Language Model SLM [D]

    <!-- SC_OFF --><div class="md"><p>Hi, I am supposed to prepare for SLM and its software part for an on campus internship, i've worked with local models like ollama generally,in my projects and also with open claw so can anyone guide me the last 2-3 days tips on what should i go t…

  1230. dev.to — LLM tag TIER_1 English(EN) · pixelbank dev ·

    History of Language Models — Deep Dive + Problem: Binary Vectorizer

    <p><em>A daily deep dive into llm topics, coding problems, and platform features from <a href="https://pixelbank.dev" rel="noopener noreferrer">PixelBank</a>.</em></p> <h2> Topic Deep Dive: History of Language Models </h2> <p><em>From the Introduction to LLMs chapter</em></p> <h2…

  1231. dev.to — LLM tag TIER_1 English(EN) · Tanishq Soni ·

    Evaluating Large Language Models: The Pitfall of Overfitting in RAG

    <h2> Introduction to RAG Evaluation </h2> <p>We often evaluate large language models (LLMs) using retrieval-augmented generation (RAG) benchmarks. However, you may have noticed that your model's performance on these benchmarks doesn't always translate to real-world success. This …

  1232. dev.to — LLM tag TIER_1 English(EN) · Tanishq Soni ·

    Evaluating Large Language Models: The Overfitting Problem

    <h2> Introduction to Overfitting in LLM Evaluation </h2> <p>We've all been there: you train a model, it performs exceptionally well on your test set, but when you deploy it to real-world scenarios, the results are disappointing. This discrepancy often stems from overfitting, a pe…

  1233. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Large language models have become remarkably capable. With enough parameters, data, and training, they can develop abilities that are difficult to predict from

    Large language models have become remarkably capable. With enough parameters, data, and training, they can develop abilities that are difficult to predict from smaller systems. They can write software, interpret documents, use tools, plan multi-step tasks, and adapt to changing i…

  1234. Mastodon — mastodon.social TIER_1 English(EN) · aitools2u ·

    📄 Research Breakthrough · Hugging Face Papers MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes 📌 We introduce MetroLLM-Bench, a 955-case ben

    📄 Research Breakthrough · Hugging Face Papers MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes 📌 We introduce MetroLLM-Bench, a 955-case benchmark for testing language m... 👇 Worth reading the full paper or wait for open-source weights? # AI # LLM 🔗 https:// h…

  1235. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    We had difficulties automatically classifying chats, messaging channels, and forums. So we tested several open-weight large language models for our use case. Th

    We had difficulties automatically classifying chats, messaging channels, and forums. So we tested several open-weight large language models for our use case. The benchmark is published below, along with a tool supporting the classification process. This helps us avoid manually cl…

  1236. Mastodon — mastodon.social TIER_1 English(EN) · aitools2u ·

    🤖 【arXiv cs.AI】Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models arXiv:2609.05437v1 Announce Type: new Abstract: Previou

    🤖 【arXiv cs.AI】Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models arXiv:2609.05437v1 Announce Type: new Abstract: Previous AI alignment efforts have focused primarily on first-order soci... # AI # TechNews # MachineLear ... 🔗 https:// arxiv.…

  1237. Mastodon — mastodon.social TIER_1 English(EN) · beyondthecode ·

    🧠 Researchers compare different ways to structure state descriptions for language model agents that predict environment changes in text-based settings. Results

    🧠 Researchers compare different ways to structure state descriptions for language model agents that predict environment changes in text-based settings. Results show that organizing facts around entities and relations as hyperedge units improves how well models learn action effect…

  1238. r/Anthropic TIER_1 English(EN) · /u/Historical-Cod-2537 ·

    Context-Induced Priority Switching in Large Language Models: Preliminary Observations

    <!-- SC_OFF --><div class="md"><p>**Original Thesis (May 18, 2026 — First Publication)**</p> <p>*The following thesis is reproduced from the primary publication of May 18, 2026 and is cited here as documentary evidence of conceptual priority. Subsequent sections represent the dev…

  1239. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    📰 Structured Language Model Generation with Outlines Outlines is an open-source library that introduces deterministic certainty into LLMs' output generation pro

    📰 Structured Language Model Generation with Outlines Outlines is an open-source library that introduces deterministic certainty into LLMs' output generation process for better, more reliable generation of structured outputs. 📰 Source: KDnuggets 🔗 Link: https://www.kdnuggets.com/s…

  1240. r/singularity TIER_2 English(EN) · /u/Tinac4 ·

    A global workspace in language models: New interpretability findings by Anthropic

    <table> <tr><td> <a href="https://www.reddit.com/r/singularity/comments/1up68u3/a_global_workspace_in_language_models_new/"> <img alt="A global workspace in language models: New interpretability findings by Anthropic" src="https://external-preview.redd.it/C6XN3tQzr96pT3Iqds1HVLyA…

  1241. r/singularity TIER_2 English(EN) · /u/yogthos ·

    Dispersion loss counteracts embedding condensation in small language models

    &#32; submitted by &#32; <a href="https://www.reddit.com/user/yogthos"> /u/yogthos </a> <br /> <span><a href="https://chenliu-1996.github.io/projects/LM-Dispersion/">[link]</a></span> &#32; <span><a href="https://www.reddit.com/r/singularity/comments/1umu4g7/dispersion_loss_count…