PulseAugur
EN
LIVE 07:40:02

New research explores LLM efficiency, safety, and multilingual capabilities

Researchers are exploring various methods to enhance the efficiency and capabilities of large language models (LLMs). Apple Inc. has published research on improving multilingual speech models by enhancing language discrimination. Other studies focus on parameter-efficient fine-tuning techniques like DyPAM and GRADE, methods for understanding internal model mechanisms such as Massive Activation Gating Channel (MAGC) and COMPASS, and efficient inference strategies like VALSE and Hybrid Latent Attention (HLA). Additionally, frameworks like SAFESHIELD are being developed for deployment-time safety of smaller language models, and novel approaches like Wasserstein-based knowledge distillation (WASD) and zero-knowledge proof of training (zkLLMPoT) are being investigated to optimize LLM performance and verification. AI

IMPACT Advances in parameter-efficient fine-tuning, model interpretability, and inference efficiency are crucial for broader LLM adoption and deployment.

RANK_REASON Cluster consists of multiple research papers published on arXiv covering various aspects of LLM development and optimization.

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 111 sources. How we write summaries →

New research explores LLM efficiency, safety, and multilingual capabilities

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Cluster consists of multiple research papers published on arXiv covering various aspects of LLM development and optimization.
Source corroboration
111 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
24 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+15 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [111]

  1. Apple Machine Learning Research TIER_1 English(EN) ·

    Language Discrimination Improves Linguistic Learning in Multilingual Speech Models

    Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model’s ability to discriminate languages during pretrai…

  2. arXiv cs.CL TIER_1 English(EN) · Mingyan Liu, Min Huang ·

    SemanticFold: Latent Sequence Compression SeparatesLanguage Modeling, Decodability, and Reasoning

    arXiv:2610.10304v1 Announce Type: cross Abstract: We study whether latent sequence compression of prompt prefixes preserves the capabilities that large language models rely on during inference. We introduce SemanticFold, a compression scheme that folds prefix hidden states at lea…

  3. arXiv cs.CL TIER_1 English(EN) · Maverick Morales, Tom\'a\v{s} Dominik, Vermut Gao, Katrina Shirey, Paulius Rimkevi\v{c}ius, Aaron Schurger, Uri Maoz ·

    Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models

    arXiv:2610.10405v1 Announce Type: cross Abstract: Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models. However, semantic chain-of-thought monitoring depends on…

  4. arXiv cs.CL TIER_1 English(EN) · Javier Mar\'in ·

    APE: Selective Fine-tuning with Acceptance Criteria for Language Model Adaptation

    arXiv:2505.19912v3 Announce Type: replace Abstract: We present Adjacent Possible Exploration (APE), a selective fine-tuning method for adapting large language models that systematically explores parameter modifications while maintaining model stability. Inspired by evolutionary o…

  5. arXiv cs.CL TIER_1 English(EN) · Kola Tubosun, Aanuoluwapo Aremu, Tolulope Ogunremi, Iroro Orife, David Ifeoluwa Adelani ·

    The \`{I}r\`{o}y\`{i}nSpeech Text Corpus: 24,905 Curated Yor\`ub\'a Sentences for Speech and Language Technology

    arXiv:2610.05366v2 Announce Type: replace Abstract: \`{I}r\`{o}y\`{i}nSpeech is a 42-hour, 80-speaker Yor\`ub\'a read-speech corpus whose audio has been distributed by ELRA since 2024. This paper describes the release of its text component: 24,905 unique, hand-verified, tone-mark…

  6. arXiv cs.LG TIER_1 English(EN) · Nicolas Lacroix, Frederic Precioso, Mireille Blay-Fornarino, Sebastien Mosser ·

    Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures

    arXiv:2610.10261v1 Announce Type: cross Abstract: Context: Once defined a taxonomy of stages structuring Machine Learning (ML) pipelines (e.g. Data Preprocessing, Modeling...), extracting these stages from source code is key for better understanding ML practices. However, the div…

  7. arXiv cs.LG TIER_1 English(EN) · Joe Dwyer ·

    Scaling Down the Scaling Laws: Parameter Efficiency and Compute-Optimal Training in Resource-Constrained Large Language Models

    arXiv:2610.06387v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved substantial performance gains through increases in model size, training data, and computational resources. However, traditional scaling approaches produce diminishing returns, rising fi…

  8. arXiv cs.CL TIER_1 English(EN) · James Ravi Kirkpatrick, Alexandru Radulescu, Rachel Katharine Sterken ·

    Talking with Language Models

    arXiv:2610.09064v1 Announce Type: cross Abstract: When we interact with large language models (LLMs), are we having a conversation? They are designed to invite us to treat them as intelligent interlocutors who remember, act, and make commitments. But appearances deceive. We intro…

  9. arXiv cs.CL TIER_1 English(EN) · Obada Kraishan ·

    The Long Road to the Same Answer: Cognitive Bias Under Escalating Reasoning Budgets in Large Language Models

    arXiv:2610.10049v1 Announce Type: new Abstract: Reasoning models allocate extra computation at inference time and present their answers as the product of deliberate thought. If this deliberation works the way dual-process accounts of human cognition suggest, longer thinking shoul…

  10. arXiv cs.CL TIER_1 English(EN) · Wenjun Wang, Heng Li, Yanggan Gu, Hongxia Yang ·

    OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models

    arXiv:2610.09346v1 Announce Type: new Abstract: Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated…

  11. arXiv cs.CL TIER_1 English(EN) · Marek \v{S}uppa, Ivan Vykopal, Andrej Ridzik, Kristi\'an Sopkovi\v{c}, Nat\'alia K\v{n}a\v{z}ekov\'a, Jaroslav Kop\v{c}an, Miroslav Bl\v{s}t\'ak, Vikt\'oria Ondrejov\'a, Daniel Hl\'adek, Michal Gregor, Martin Tamajka, Mari\'an \v{S}imko ·

    sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak

    arXiv:2610.09152v1 Announce Type: new Abstract: Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored…

  12. arXiv cs.CL TIER_1 English(EN) · Tobias Braun, Nils Loose, Alexander Herzog, Virginia Ceccatelli, Marcus Rohrbach, Thomas Eisenbarth, Lorenzo Cavallaro ·

    U-Space: Uncovering When and Why Uncertainty Arises in Language Models

    arXiv:2610.09087v1 Announce Type: new Abstract: Large language models are informing decisions with ever-higher stakes. As the consequences of their errors grow, a central question becomes harder to ignore: how much can we trust an individual answer? Yet recognizing when to defer …

  13. arXiv cs.CL TIER_1 English(EN) · Pavan Maddula ·

    Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs

    arXiv:2610.09033v1 Announce Type: new Abstract: Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and …

  14. arXiv cs.AI TIER_1 English(EN) · Xingru Zhou, Luis Sentis, Aarti Choudhary ·

    SAFESHIELD: A Decision-Organization Framework for Deployment-Time Safety of Small Language Models

    arXiv:2610.07276v1 Announce Type: cross Abstract: Deployment-time safety of language models is commonly implemented through runtime guardrails such as input moderation, routing, retrieval verification, and output filtering. Existing deployment frameworks provide increasingly capa…

  15. arXiv cs.CL TIER_1 English(EN) · Liwei Jiang, Yuanjun Chai, Margaret Li, Mickel Liu, Raymond Fok, Nouha Dziri, Yulia Tsvetkov, Maarten Sap, Alon Albalak, Yejin Choi ·

    Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)

    arXiv:2510.22954v2 Announce Type: replace Abstract: Language models (LMs) often struggle to generate diverse, human-like creative content, raising concerns about the long-term homogenization of human thought through repeated exposure to similar outputs. Yet scalable methods for e…

  16. arXiv cs.CL TIER_1 English(EN) · Seobin Song, Geonho Lee, Janghwan Lee, Jungwook Choi ·

    Align, Then Correct: Training-Free Two-Stage Low-Rank Compensation for Extremely Quantized Large Language Models

    arXiv:2610.08164v1 Announce Type: cross Abstract: Low-rank quantization error compensation (LQEC) recovers the accuracy lost under aggressive weight quantization by attaching a closed-form rank-$r$ adapter beside each frozen quantized weight, without any training. We show that ex…

  17. arXiv cs.CL TIER_1 English(EN) · Avichal Sahai (Ofbusiness), Nishant Raj (Ofbusiness), Animesh Srivastava (Ofbusiness) ·

    Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changes

    arXiv:2610.07853v1 Announce Type: cross Abstract: Ternary language models such as BitNet b1.58, Falcon-E and BitCPM are fine-tuned with higher-precision latent weights and deployed as ternary codes produced by an export step that, in the labs' documented pipelines, first casts th…

  18. arXiv cs.CL TIER_1 English(EN) · Haoyu Huang, Zhongwei Xie, Jiaxin Bai, Yisen Gao, Hong Ting Tsang, Wuganjing Song, Huihao Jing, Yufei Li, Yangqiu Song ·

    Towards In-Parameter Memory Augmentation for Large Language Models

    arXiv:2610.08630v1 Announce Type: new Abstract: Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, user preferences, documents, and interaction experience. In-context learning (ICL)…

  19. arXiv cs.CL TIER_1 English(EN) · Shuqing Shi, Ziyan Wang, Milind Tambe, Yali Du ·

    Large Language Model Orchestration under Heterogeneous Preferences via Explicit Persona Inference

    arXiv:2610.07587v1 Announce Type: new Abstract: LLM orchestration investigates how an orchestrator coordinates a group of autonomous agents to achieve common goals or maximize collective welfare. The agents are typically heterogeneous, each holding a private preference that it pu…

  20. arXiv cs.CL TIER_1 English(EN) · Huan Li, Zhe Cao, Qinlei Xie, Fushun Cui, Xuechen Liang ·

    Stabilizing language models under continual learning via condition-anchored distillation

    arXiv:2610.06940v1 Announce Type: new Abstract: Continual adaptation of language models can change their output distribution on prompts learned earlier, while retaining every old prompt-answer pair may be undesirable or impossible. We study condition-anchored generative distillat…

  21. arXiv cs.AI TIER_1 English(EN) · Rhitabrat Pokharel, Ameeta Agrawal, Tanay Nagar ·

    Cross-Lingual Activation Steering for Multilingual Language Models

    arXiv:2601.16390v2 Announce Type: replace-cross Abstract: Large language models exhibit strong multilingual capabilities, yet significant performance gaps persist between dominant and non-dominant languages. Prior work attributes this gap to imbalances between shared and language…

  22. arXiv cs.AI TIER_1 English(EN) · Chuan Li, Qianyi Zhao, Fengran Mo, Cen Chen ·

    FedCoT: Communication-Efficient Federated Reasoning Enhancement for Large Language Models

    arXiv:2508.10020v2 Announce Type: replace-cross Abstract: Enhancing LLM reasoning in federated settings is nontrivial due to stringent computational, communication, and privacy constraints, especially in healthcare, where clinically consequential decisions require not only accura…

  23. arXiv cs.AI TIER_1 English(EN) · Mingyuan Zhang, Yue Bai, Huan Wang, Yizhou Wang, Qihua Dong, Yitian Zhang, Yun Fu ·

    Boosting Large Language Models with Mask Fine-Tuning

    arXiv:2503.22764v3 Announce Type: replace-cross Abstract: The large language model (LLM) is typically integrated into the mainstream optimization protocol. However, it remains underexplored whether maintaining the model integrity is \textit{indispensable} for promising performanc…

  24. arXiv cs.AI TIER_1 English(EN) · Yuto Suzuki, Farnoush Banaei-Kashani ·

    Universe of Thoughts: A Computational Framework for Creative Reasoning in Large Language Models

    arXiv:2511.20471v3 Announce Type: replace Abstract: Recent advances in Large Language Model (LLM) reasoning have improved conventional problem solving, but creative reasoning remains comparatively underexplored. Inspired by cognitive science, we formalize combinational, explorato…

  25. arXiv cs.AI TIER_1 English(EN) · Yuhan Chen, Siyuan Zhang, Nan Wang, Feiyang Kang, Ruoxi Jia ·

    Hybrid Latent Attention for Looped Language Models

    arXiv:2610.07940v1 Announce Type: cross Abstract: Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T. The larger cache limits how many sequences a GPU can deco…

  26. arXiv cs.AI TIER_1 English(EN) · Dayan Pan, Jingyuan Wang, Xie Yu ·

    Dynamic Positional Attention Modulation for Parameter-Efficient Fine-Tuning of Large Language Models

    arXiv:2610.07848v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning (PEFT) has become a standard approach for adapting large language models to downstream tasks. However, most existing PEFT methods rely on uniform and static adaptations, without accounting for the s…

  27. arXiv cs.AI TIER_1 English(EN) · Byeonghu Na, Donghyeok Shin, Yeongmin Kim, Mina Kang, Il-Chul Moon ·

    WASD: Wasserstein-based Knowledge Distillation for Large Language Models

    arXiv:2610.07706v1 Announce Type: cross Abstract: Autoregressive large language models (LLMs) have rapidly advanced in capability, but their increasing scale comes with substantial computational and memory costs at inference time. Knowledge distillation (KD) offers a practical so…

  28. arXiv cs.AI TIER_1 English(EN) · Hongyu Cao, Yanchi Liu, Kunpeng Liu, Xujiang Zhao, Wei Cheng, Zhengzhang Chen, Yanjie Fu, Haifeng Chen ·

    Which and When to Admit: Gradient Admission for Data-Centric Small Language Model Finetuning

    arXiv:2610.07553v1 Announce Type: cross Abstract: LoRA fine-tuning adapts small language models (SLMs) to heterogeneous instruction data within a low-rank update subspace, making it vulnerable to three structural problems: conflicting gradients that cancel, static data selection …

  29. arXiv cs.AI TIER_1 English(EN) · Junkai Liang, Zhanpeng Guo, Pengfei Wu, Qingni Shen, Jiaheng Zhang, Zhonghai Wu, Haiyang Xue, Shengfang Zhai ·

    zkLLMPoT: Efficient Zero Knowledge Proof of Training for Large Language Models

    arXiv:2610.08258v1 Announce Type: new Abstract: Auditing the claimed outcomes of large language model (LLM) training is challenging when model weights and training data are private, while cryptographically proving the full training process is prohibitively expensive at Transforme…

  30. arXiv cs.AI TIER_1 English(EN) · Minjia Mao, Shi Chen, Bowen Yin, Xiao Fang ·

    Massive Activation Gating Channel in Large Language Models

    arXiv:2610.07661v1 Announce Type: new Abstract: Massive activations, a phenomenon in which a small number of hidden channels exhibit exceptionally large magnitudes, are pervasive in large language models (LLMs). However, the mechanism by which a token develops massive activations…

  31. arXiv cs.AI TIER_1 English(EN) · Jia-Dong Zhang ·

    VALSE: Vertical Adaptive Layer Skipping for Efficient Inference in Large Language Models

    arXiv:2610.07606v1 Announce Type: new Abstract: This paper establishes a theoretical framework for vertical adaptive layer skipping, proving three foundational results: (i) an Expected FLOPs formula (theorem 2) giving a closed-form expression for the computational cost of arbitra…

  32. arXiv cs.AI TIER_1 English(EN) · Pratyay Dutta, Kowshik Thopalli, Vivek Narayanaswamy ·

    COMPASS: Finding Where Reasoning Lives in Language Models

    arXiv:2610.07469v1 Announce Type: new Abstract: Explicitly eliciting reasoning substantially improves LLM performance. Existing approaches require a predefined characterization of reasoning, whether through CoT prompt design, contrastive CoT directions, or via SAE derived reasoni…

  33. arXiv cs.LG TIER_1 English(EN) · Shuo Yang, Changbai Li, Linlin Yang, Huobin Tan, Rongyu Chen, Tongfei Chen, Tian Wang, Sheng Xu, Baochang Zhang ·

    DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models

    arXiv:2610.08341v1 Announce Type: cross Abstract: Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic de…

  34. arXiv cs.LG TIER_1 English(EN) · Heng Liang, Xinwen Zhang, Hongchang Gao ·

    Learning What to Distill: Bilevel Top-K Token Selection for Self-Distillation in Large Language Models

    arXiv:2610.07247v1 Announce Type: new Abstract: Large language models have shown strong reasoning capabilities, but their high inference costs make knowledge distillation an important approach for transferring such capabilities to compact models in resource-constrained scenarios.…

  35. Hugging Face Daily Papers TIER_1 English(EN) ·

    Align, Then Correct: Training-Free Two-Stage Low-Rank Compensation for Extremely Quantized Large Language Models

    Low-rank quantization error compensation (LQEC) recovers the accuracy lost under aggressive weight quantization by attaching a closed-form rank-$r$ adapter beside each frozen quantized weight, without any training. We show that existing compensators are limited by two shared simp…

  36. Hugging Face Daily Papers TIER_1 English(EN) ·

    VALSE: Vertical Adaptive Layer Skipping for Efficient Inference in Large Language Models

    This paper establishes a theoretical framework for vertical adaptive layer skipping, proving three foundational results: (i) an Expected FLOPs formula (theorem 2) giving a closed-form expression for the computational cost of arbitrary per-sample skip schedules as a function of la…

  37. Hugging Face Daily Papers TIER_1 English(EN) ·

    Towards In-Parameter Memory Augmentation for Large Language Models

    Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, user preferences, documents, and interaction experience. In-context learning (ICL) and ICL-based agent harness remain flexible, bu…

  38. arXiv cs.AI TIER_1 English(EN) · Linkai Ma, Xinyu Luo, Mengbo Wang, Ananth Grama, Petros Drineas, Brian Bullins ·

    MuonIO: Principled Norm-Aware Descent for Embedding Tables and Language Model Heads

    arXiv:2610.02705v1 Announce Type: cross Abstract: The Muon optimizer derives its update rule for hidden linear layers by solving a local linearization of the loss penalized by the spectral norm, motivated by an RMS-stability argument for dense linear layers. Standard Muon impleme…

  39. arXiv cs.AI TIER_1 English(EN) · Mohsen Larni (Department of Computer Science, University of Nevada, Las Vegas), Sobhan Ebrahimi Azar (Department of Computer Science, University of Nevada, Las Vegas), Pouyan Nahed (Department of Computer Science, University of Nevada, Las Vegas), Kazem … ·

    SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models

    arXiv:2610.03329v1 Announce Type: cross Abstract: Large language models are increasingly used where small syntactic errors matter, yet character-level reasoning is still evaluated mostly through isolated probes and aggregate accuracy. We introduce SyntaxBench, a diagnostic benchm…

  40. arXiv cs.AI TIER_1 English(EN) · Shuyang Cao, Karthik Radhakrishnan, David Rosenberg, Steven Lu, Pengxiang Cheng, Lu Wang, Shiyue Zhang ·

    Evaluating the Retrieval Robustness of Large Language Models

    arXiv:2505.21870v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) generally enhances large language models' (LLMs) ability to solve knowledge-intensive tasks. But RAG could also lead to performance degradation due to imperfect retrieval and the model'…

  41. arXiv cs.LG TIER_1 English(EN) · Shivam Gupta ·

    Exact Memory-Time Optimization for Prefix-Cached Language Model Serving

    arXiv:2610.02766v1 Announce Type: new Abstract: Retaining language-model prefix states trades recomputation against storage time. Optimizing each cached block independently can overcount savings: a resident block is usable only when the required preceding prefix is also available…

  42. arXiv cs.LG TIER_1 English(EN) · Nicolas Martorell, Wendy Brau, Gonzalo A. Heredia, Tom\'as Pablo Korenblit, Gaspar Labasti\'e, Tom\'as Gimenez Molina ·

    PowerBench: Measuring Language Model Bias in Power-shifting Requests

    arXiv:2610.02303v1 Announce Type: new Abstract: Language models increasingly assist people with power-related requests, so systematic differences in whom they help could shift the distribution of power at scale, or be exploited by users who learn which identities are refused less…

  43. arXiv cs.CL TIER_1 English(EN) · Jungseob Lee, Seungyoon Lee, Sugyeong Eo, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim ·

    Predicting and Repairing Merge Collapse in Large Language Models

    arXiv:2610.03199v1 Announce Type: cross Abstract: Large language models fine-tuned from a shared base can be merged by averaging their task vectors, but some merges collapse far below the base model, and common merge operators give no warning before evaluation. We show that one s…

  44. arXiv cs.CL TIER_1 English(EN) · Taiheng Pan ·

    Probe the Harness: Setup Checks for Stale-Data RL Comparisons in Language Models

    arXiv:2610.02911v1 Announce Type: cross Abstract: Methods for training language models on stale samples are judged by comparisons against importance-corrected baselines. We show that details of the experimental harness can reverse the observed ranking of methods, and we introduce…

  45. arXiv cs.LG TIER_1 English(EN) · Erfan Hajihashemi, Yanning Shen ·

    Online Verification of Language Model Responses Under Cost Constraints

    arXiv:2610.02632v1 Announce Type: new Abstract: As large language models are increasingly deployed for multi-step reasoning, verifying the correctness of their outputs has become essential for maintaining reliability at scale. Verifying the correctness of large language model out…

  46. arXiv cs.CL TIER_1 English(EN) · Narek Maloyan ·

    Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations

    arXiv:2610.02432v1 Announce Type: cross Abstract: Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM r…

  47. arXiv cs.CL TIER_1 English(EN) · Tao Shi, Chaoyi Xiang, Qiongkai Xu, Jey Han Lau ·

    OLMo-Detect: A Multi-Stage, Confounder-Controlled Benchmark for Membership Inference on Large Language Models

    arXiv:2610.02986v1 Announce Type: new Abstract: Membership inference on large language models (LLMs) aims to determine whether a given text sample was included in an LLM's training data, without access to its training corpus. Despite recent progress, existing benchmarks suffer fr…

  48. arXiv cs.CL TIER_1 English(EN) · Baohang Li, Xiaocheng Feng, Yichong Huang, Chengpeng Fu, Wenshuai Huo, Zekun Zhou, Zekun Yuan, Tingjia Zhang, Bing Qin ·

    Adaptive Mutual Distillation for Balanced Multi-Task Post-Training of Large Language Models

    arXiv:2610.02856v1 Announce Type: new Abstract: Multi-task post-training of large language models (LLMs) aims to improve performance across tasks with unequal amounts of training data. Existing methods focus primarily on balancing task contributions during single-model training. …

  49. arXiv cs.LG TIER_1 English(EN) · Zhenghao Zhao, Gaowen Liu, Zhiling Lan, Yan Yan ·

    Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models

    arXiv:2610.02593v1 Announce Type: new Abstract: Data selection is already a central bottleneck in large-language-model training, where web-scale corpora are noisy and token budgets are finite. In continual pre-training (CPT), it becomes a forgetting-control problem: a poorly chos…

  50. arXiv cs.AI TIER_1 English(EN) · Han Wang, Ishwar B Balappanawar, Huan Zhang ·

    On the Chain-of-Thought Monitorability of Looped Language Models

    arXiv:2610.02741v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring provides a promising approach for detecting undesirable model behavior. Looped language models (LoopLMs) repeatedly apply shared transformer layers, increasing effective computational depth and enab…

  51. arXiv cs.CL TIER_1 English(EN) · Fahmid Shahriar Iqbal, Ritam Dutt, Soumitra Das, Arnav Verma, Sagnik Ray Choudhury ·

    Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks

    arXiv:2610.02549v1 Announce Type: new Abstract: Time and event expression extraction are fundamental temporal reasoning tasks, but the problem remains difficult due to annotation ambiguity, domain sensitivity, and unstable model behavior. Existing evaluations focus on in-domain p…

  52. arXiv cs.AI TIER_1 English(EN) · Zizhuo Zhang, Xiong Peng, Jingwei Sun, Rong Yao, Shixiong Kai, Mingxuan Yuan, Bo Han ·

    Test-time Calibration Learning for Large Language Model Reasoning

    arXiv:2610.02695v1 Announce Type: cross Abstract: Reliable large language models (LLMs) must not only produce accurate answers but also express confidence that faithfully reflects their probability of being correct. Such calibration is essential for identifying uncertain predicti…

  53. arXiv cs.LG TIER_1 English(EN) · Yuxiang Wang, Kunyu Feng, Yuancheng Wang, Zihang Liu, Shengbo Cai, Qinke Ni, Wan Lin, Tao Feng, Yingda shen, Ming-Hao Hsu, Zhixian Zhao, Liqiang Zhang, Teddy Sun, Steve Yves, Zhizheng Wu ·

    AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

    arXiv:2610.01560v1 Announce Type: cross Abstract: Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but g…

  54. arXiv cs.LG TIER_1 English(EN) · Jian Xu ·

    Large Language Bayes Is Not Reparameterisation-Invariant

    arXiv:2610.00265v1 Announce Type: new Abstract: Large Language Bayes (LLB) answers an informal modelling question by sampling candidate probabilistic programs from a language model, running approximate inference on each, and averaging them with weights proportional to an exponent…

  55. arXiv cs.LG TIER_1 English(EN) · Bangji Yang, Jingyuan Li, Jiajun Fan, Yi Evie Zhang, Ruihan Guo, Hongba Ma, Neil He, Chumeng Liang, Qinglong Zheng, Zhanghan Ni, Ge Liu ·

    How Much Can Language Models Gain from Test-Time Computation?

    arXiv:2610.01110v1 Announce Type: new Abstract: How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge s…

  56. arXiv cs.LG TIER_1 English(EN) · Shengye Tao, Yinzhu Cheng, Haihua Xie ·

    Persistent Depth Ordering amid Shifting Block-Bypass Responses in Language Model Pretraining

    arXiv:2610.01165v1 Announce Type: new Abstract: Layer interventions are widely used to probe the internal organization of language models, yet most analyses examine a single training checkpoint even though model representations and computations evolve throughout pretraining. This…

  57. arXiv cs.LG TIER_1 English(EN) · Alessandro Turrin, Patrik Okanovic, Torsten Hoefler, Nezihe Merve G\"urel ·

    Which LLM to pick? Online Active Model Selection for Large Language Models

    arXiv:2610.01592v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly applied to process streaming data, with practitioners relying on benchmarks to select the best model even though these signals only approximate real performance. While oracle annotatio…

  58. arXiv cs.LG TIER_1 English(EN) · Roshan Balaji, Kamran Chitsaz, Quentin Fournier, Nirav Pravinbhai Bhatt, Sarath Chandar ·

    A Large Scale Investigation of Scaling Limits in Chemical Language Models

    arXiv:2508.13408v3 Announce Type: replace Abstract: Chemical Language Models (CLMs) are increasingly used in de novo drug design, driven by recent growth in model scale, compute, and dataset size. However, the relationship between design choices, training dynamics, and downstream…

  59. arXiv cs.LG TIER_1 English(EN) · Chayne Thrash, Kevin Chen, Soheil Kolouri ·

    Output-aware Residual Stream Pruning for Large Language Models

    arXiv:2609.35579v2 Announce Type: replace Abstract: Residual stream pruning methods reduce inference cost by shrinking the model's hidden dimension, but existing approaches typically choose these dimensions by minimizing activation reconstruction error. This criterion implicitly …

  60. Hugging Face Daily Papers TIER_1 English(EN) ·

    The Numerical Linear Algebra of Large Language Models

    Numerical Linear Algebra (NLA) has consistently played a vital role in advancing science by providing tools to solve fundamental problems encountered in scientific and engineering applications. Over the decades, it has continually evolved to meet the demands driven by successive …

  61. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Naeemullah Khan ·

    Periscope: Extending Frozen Language Models Beyond Their Context Window

    A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which …

  62. arXiv cs.AI TIER_1 English(EN) · Beatriz Almeida Felicio ·

    How Divergence Becomes Decision Flips in Compressed Language Models

    arXiv:2610.00694v1 Announce Type: cross Abstract: Compression reports summarize how far a compressed language model moved from the dense one, usually by a KL divergence; a deployment that relies on the dense model's outputs needs to know how many of its decisions changed. We show…

  63. arXiv cs.AI TIER_1 English(EN) · Andrei Marchenko, Viacheslav Bezrukov, Oleg Kashurin, Inessa Fedorova, Dmitry Bocharov, Yuliana Shakhvalieva, Maria Tikhonova, Valerii Ternovskii ·

    Closing the Loop: Practical Training Recipes for Looped Language Models

    arXiv:2610.00673v1 Announce Type: cross Abstract: Looped language models increase effective depth by repeatedly applying a shared block of layers, but existing large-scale recipes require multi-stage training over trillions of tokens, while the benefits of recurrence remain diffi…

  64. arXiv cs.AI TIER_1 English(EN) · Pardis Sadat Zahraei, Janvijay Singh, Gokhan Tur, Dilek Hakkani-Tur ·

    Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness

    arXiv:2610.00568v1 Announce Type: cross Abstract: Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tens…

  65. arXiv cs.AI TIER_1 English(EN) · Timoth\'ee Lesort, Alejandra L\'opez de Aberasturi G\'omez, Tristan Karch, Tom Veniat, Philippe Modard, Karl Tuyls, Ludovic Denoyer ·

    Benchmarking Prompt Optimization of Large Language Models With Chess

    arXiv:2610.00416v1 Announce Type: new Abstract: Evaluating large language models becomes increasingly challenging as their capabilities advance: benchmarks can saturate, public test sets risk contamination, and assessing harder tasks can require expensive grading or execution inf…

  66. arXiv cs.AI TIER_1 English(EN) · Oguzhan Baser, Elahe Sadeghi, Eric Wang, Nico Vergauwen, Sam Kazemian, Hong Kang, Sandeep P. Chinchali, Sriram Vishwanath ·

    TensorCommitments: A Lightweight Verifiable Inference for Language Models

    arXiv:2602.12630v2 Announce Type: replace-cross Abstract: Most large language models (LLMs) run on external clouds: users send a prompt, pay for inference, and must trust that the remote GPU executes the LLM without any adversarial tampering. We critically ask how to achieve veri…

  67. arXiv cs.AI TIER_1 English(EN) · Jiangfeng Chen, Xinyu Wang, Tianshuo Yan, Hanwei Wu, Xiao-Wen Chang, Yang Zhang, Lei Ding ·

    Sequential Functional Structured Tucker Compression for Large Language Model Attentions

    arXiv:2610.00717v1 Announce Type: cross Abstract: Post-training compression of LLM attention is often formulated as independent matrix approximation, ignoring both the shared structure among attention projections and the representation shift introduced by earlier compression. We …

  68. arXiv cs.AI TIER_1 English(EN) · Dongyub Jude Lee, Jungseob Lee, Seungyoon Lee, Seongtae Hong, Suhyune Son, Sugyeong Eo, Jaehyung Seo, Heuiseok Lim ·

    Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations

    arXiv:2606.22676v2 Announce Type: replace Abstract: Refusal on a safety benchmark does not reveal how stable that behavior will remain after model updates. Benign downstream fine-tuning can weaken refusal, yet behavioral evaluations typically expose this fragility only after an i…

  69. Hugging Face Daily Papers TIER_1 English(EN) ·

    Periscope: Extending Frozen Language Models Beyond Their Context Window

    A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which …

  70. Hugging Face Daily Papers TIER_1 English(EN) ·

    Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale

    A trainable input embedding table assigns each vocabulary item an independently adjustable vector. We investigate whether this token-specific parameterization is required for substantial language-modeling capability, or whether a shared Transformer can learn from fixed token iden…

  71. arXiv cs.AI TIER_1 English(EN) · Wanda Hou, Leon Zhou, Hong-Ye Hu, Yubei Chen, Yi-Zhuang You, Xiao-Liang Qi ·

    A Quantitative Study of Sustained Focus in Large Language Models via Repetitive Deterministic Prediction Tasks

    arXiv:2511.00763v3 Announce Type: replace Abstract: We investigate the performance of large language models (LLMs) on repetitive deterministic prediction tasks and study how the sequence accuracy rate (SAR) scales with output length. Each such task involves the repetition of the …

  72. arXiv cs.AI TIER_1 English(EN) · Gabriele Tuccio, Antonino Furnari, Aldo Gangemi, Misael Mongiov\`{\i} ·

    GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning

    arXiv:2609.39869v1 Announce Type: new Abstract: Grammar-constrained generation guarantees syntactic validity, but can substantially degrade semantic quality when the model's preferred outputs are poorly aligned with the imposed grammar. This trade-off is particularly severe when …

  73. arXiv cs.AI TIER_1 English(EN) · Takanori Kotama, Shun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino, Takahiro Katagiri ·

    NinaXander: Feasibility and Limits of Composing Frozen Language Models Across Architecture Families via a Shared Latent Space

    arXiv:2609.38261v1 Announce Type: cross Abstract: In this paper we propose NinaXander, a series of composed language models obtained by connecting layers of frozen language models from different architecture families with a single trained shared-latent adapter. A composed model r…

  74. arXiv cs.AI TIER_1 English(EN) · Qiuyu Ren, Sudipta Paria, Aritra Dasgupta, Swarup Bhunia ·

    Security-Enhanced Seed-Based Weight Quantization for Large Language Models

    arXiv:2609.38477v1 Announce Type: cross Abstract: Large language models (LLMs) incur substantial storage, memory-bandwidth and energy costs, motivating compact weight representations. Existing seed-based compression methods reconstruct weights from compact pseudo-random represent…

  75. arXiv cs.AI TIER_1 English(EN) · Rajat Ghosh, Vaishnavi Bhargava, Henry Wong, Aryan Singhal, Debojyoti Dutta ·

    GRPO Training Dynamics for Small Language Models

    arXiv:2609.39321v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) has emerged as a memory-efficient reinforcement fine-tuning (RFT) technique for reasoning-intensive tasks. How- ever, GRPO training dynamics on small language models (SLMs) remain poorly u…

  76. arXiv cs.AI TIER_1 English(EN) · Zhen Liang, Hai Huang, Wentao Chen ·

    CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion

    arXiv:2609.39902v1 Announce Type: cross Abstract: Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety gen…

  77. arXiv cs.AI TIER_1 English(EN) · Hoa Quynh Nhung Nguyen, Jacopo Staiano, Michael Sullivan ·

    On the (In)effectiveness of AMR Augmentation for Large Language Models

    arXiv:2609.40121v1 Announce Type: cross Abstract: While Abstract Meaning Representation (AMR) has historically improved performance on a range of NLP tasks, the benefit---or lack thereof---of AMR augmentation for modern LLMs is thus far unclear. In this paper, we attempt to repro…

  78. arXiv cs.CL TIER_1 English(EN) · Hao Liang, Zhengyang Zhao, Mingrui Chen, Meiyi Qiang, Lu Ma, Rongyi Yu, Hengyi Feng, Shixuan Sun, Zimo Meng, Xiaochen Ma, Xuanlin Yang, Qifeng Cai, Ruichuan An, Bohan Zeng, Zhen Hao Wong, Chengyu Shen, Runming He, Zhaoyang Han, Yaowei Zheng, Fangcheng Fu… ·

    DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models

    arXiv:2603.26164v2 Announce Type: replace-cross Abstract: Data-centric training has emerged as a promising direction for improving large language models (LLMs) by optimizing not only model parameters but also the selection, composition, and weighting of training data during optim…

  79. arXiv cs.CL TIER_1 English(EN) · Saatvik Kher, Shang Wu, Rachel Longjohn, Catarina Bel\'em, Padhraic Smyth ·

    Sequential Bayesian Evaluation of Large Language Model Behavior

    arXiv:2511.10661v2 Announce Type: replace Abstract: It is increasingly important to evaluate the characteristics of systems based on large language models (LLMs). Evaluations in this context often rely on a curated benchmark set of input prompts provided to the LLM, where the out…

  80. arXiv cs.CL TIER_1 English(EN) · Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov ·

    Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models

    arXiv:2410.02343v2 Announce Type: replace Abstract: Large language models (LLMs) routinely fail to output the correct option in multiple-choice question answering (MCQA) while encoding the answer internally. We expose this latent knowledge via the Query--Key (QK) score, defined f…

  81. arXiv cs.CL TIER_1 English(EN) · Jialiang Sun, Kuldeep Meel ·

    Provably Tractable NFA-Constrained Language Generation via HMMs

    arXiv:2609.40185v1 Announce Type: new Abstract: Constrained generation aims to sample from language models (LMs) conditioned on hard constraints. Existing constrained-generation techniques for nondeterministic finite automaton (NFA) constraints either distort the distribution or …

  82. arXiv cs.CL TIER_1 English(EN) · Jhen-Ke Lin, Chung Chun Wang ·

    StreamDecisionBench: Evaluating Decisions in Force on Evolving Language Streams

    arXiv:2609.38612v1 Announce Type: new Abstract: As natural language drives more applications, language models increasingly run inside programs as decision components: the program sends them the current state and acts on the returned decision until a newer one arrives. When eviden…

  83. arXiv cs.CL TIER_1 English(EN) · Parisa Salmani, Peter R. Lewis ·

    Evaluating Language Model Safety Across Long Adversarial Conversations

    arXiv:2609.38357v1 Announce Type: new Abstract: Conversational safety evaluations often test language models with a single harmful prompt, even though real-world systems interact with users through long, adaptive conversations. This study examines whether models continue to respo…

  84. arXiv cs.CL TIER_1 (CA) · Juan M Zambrano Chaves, Peniel Argaw, Risa Ueno, Carlo Bifulco, Kristina Young, Rom Leidner, Tristan Naumann, Hoifung Poon ·

    Large Language Models are Approximate Survival Estimators

    arXiv:2609.38181v1 Announce Type: new Abstract: Survival analysis estimates time-to-event outcomes from patient covariates and is widely used for medical risk assessment. Patients seeking prognostic information after a diagnosis may turn to large language models (LLMs), now readi…

  85. arXiv cs.AI TIER_1 English(EN) · Kenan Alkiek, David Jurgens, Vinod Vydiswaran ·

    Instruction Retrieval at Inference Time for Small Language Models

    arXiv:2510.13935v3 Announce Type: replace-cross Abstract: The facts a language model stores are tied to its parameter count, so small models that fit on edge devices fail on expert problems, which need specialized knowledge and follow multi-step procedures. Fine-tuning for a spec…

  86. arXiv cs.AI TIER_1 English(EN) · Mansi Sakarvadia, Aswathy Ajith, Arham Khan, Nathaniel Hudson, Caleb Geniesse, Kyle Chard, Yaoqing Yang, Ian Foster, Michael W. Mahoney ·

    Mitigating Memorization In Language Models

    arXiv:2410.02159v3 Announce Type: replace-cross Abstract: Language models (LMs) can "memorize" information, i.e., encode training data in their weights in such a way that inference-time queries can lead to verbatim regurgitation of that data. This ability to extract training data…

  87. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Roberto Di Pietro ·

    Consensus and Factual Dynamics in Large Populations of Interacting Language Models

    Large Language Model (LLM) agents are increasingly deployed as populations of interacting entities, in which consensus --agreement on a shared answer-- emerges as a collective, unengineered behaviour. Prior work on LLM consensus shows that agents can cross-verify their answers an…

  88. arXiv cs.CL TIER_1 English(EN) · Elsayed Eshra, Ali Al-Lawati, Dongwon Lee, Suhang Wang ·

    Quantifying Behavioral Tails in Black-Box Language Models

    arXiv:2609.33638v2 Announce Type: replace-cross Abstract: We introduce RareTrap, a framework for estimating the probability of severe behaviors in black box large language models (LLMs). A key challenge for probability estimation is defining a tractable distribution over the inpu…

  89. arXiv cs.LG TIER_1 English(EN) · Puning Yang, Qizhou Wang, Junchi Yu, Bo Han, Xiuying Chen ·

    UnlearningSoup: Is Repeated Tuning Necessary for Large Language Model Unlearning?

    arXiv:2609.37076v1 Announce Type: new Abstract: Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this issue, existing unlearning methods typically rely on training-based parameter updat…

  90. arXiv cs.CL TIER_1 English(EN) · Prosper Arineitwe Asiimwe, Francois Meyer, Jan Buys ·

    RunyaNER: Auxiliary Language Selection for Runyankore NER

    arXiv:2609.37543v1 Announce Type: new Abstract: Cross-lingual zero-shot transfer and multilingual fine-tuning are promising approaches for NLP tasks such as Named Entity Recognition (NER) in low-resource languages, but in the absence of target language benchmarks, it is unclear w…

  91. arXiv cs.AI TIER_1 English(EN) · Sin-Han Yang, Cheng-Kuang Wu, Chieh-Yen Lin, Yun-Nung Chen, Hung-yi Lee, Shao-Hua Sun ·

    On Calibration of Large Language Models: From Response To Capability

    arXiv:2602.13540v2 Announce Type: replace-cross Abstract: Accurate confidence estimation is critical for reliable use of large language models (LLMs). Prior work on LLM calibration largely focuses on response-level confidence, which estimates the correctness of a single generated…

  92. arXiv cs.AI TIER_1 English(EN) · Jiaqi Xue, Mengxin Zheng, Yebowen Hu, Fei Liu, Xun Chen, Qian Lou ·

    BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models

    arXiv:2406.00083v3 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by retrieving relevant information from external knowledge bases to provide more accurate, contextually informed, and up-to-date responses. However…

  93. arXiv cs.AI TIER_1 English(EN) · Rebecca Ramnauth, Brian Scassellati ·

    Compiling Learning Problems into Adaptation Programs for Language Models

    arXiv:2609.37371v1 Announce Type: cross Abstract: Model adaptation is typically governed by a fixed recipe, even though different update programs can produce substantially different behavioral outcomes. We introduce adaptation compilation, which reframes where, how, and to what e…

  94. arXiv cs.AI TIER_1 English(EN) · Kenan Alkiek, Moontae Lee, David Jurgens, V. G. Vinod Vydiswaran ·

    HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment

    arXiv:2609.38006v1 Announce Type: new Abstract: Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud mo…

  95. arXiv cs.AI TIER_1 Deutsch(DE) · Yuansen Liu, Yixuan Tang, Anthony Kum Hoe Tung ·

    Locating Answer-Correctness Signals in Frozen Large Language Models

    arXiv:2609.37700v1 Announce Type: new Abstract: Language models expose internal signals that predict whether an answer is correct, readable from a single forward pass of a frozen model without additional generations. Yet existing probes often commit to one signal family or layer …

  96. Hugging Face Daily Papers TIER_1 Deutsch(DE) ·

    Locating Answer-Correctness Signals in Frozen Large Language Models

    Language models expose internal signals that predict whether an answer is correct, readable from a single forward pass of a frozen model without additional generations. Yet existing probes often commit to one signal family or layer and can be brittle under distribution shift; in …

  97. arXiv cs.AI TIER_1 English(EN) · Tong Xiao, Jingbo Zhu ·

    Foundations of Large Language Models

    arXiv:2501.09223v3 Announce Type: replace-cross Abstract: This is a book about large language models. As indicated by the title, it primarily focuses on foundational concepts rather than comprehensive coverage of all cutting-edge technologies. The book is structured into six main…

  98. arXiv cs.AI TIER_1 English(EN) · Andrei Chetvergov, Mikhail Solovev, Timofei Sivoraksha, Stepan Ukolov, Valeriia Kuschenko, Alexander Evseev, Sergey Bolovtsov ·

    PROOF: Profiling Reliability of Object-Level Facts in Large Language Models

    arXiv:2609.29504v1 Announce Type: cross Abstract: Aggregate factuality scores hide where a language model succeeds, which relations it confuses, and whether an answer survives innocuous changes to the question or decoder. We introduce PROOF, a profile-oriented benchmark for factu…

  99. Hugging Face Daily Papers TIER_1 English(EN) ·

    Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation

    We study a practical question: can a small correction module fix errors in a frozen language model's outputs without degrading its base capabilities? We propose CRN v2, a lightweight logit-level correction module (~34M trainable parameters, 0.73% of the 4.65B text module) that si…

  100. arXiv cs.CV TIER_1 English(EN) · Yupeng Zhang, Ziyi Zhao, Juntao Cheng, Sheng Wang, Ningnan Guo, Ruize Han, Liang Wan ·

    RT-DETR-World: Transferring Rich LLM Semantics to Real-Time Open-Vocabulary Detection

    arXiv:2610.09502v1 Announce Type: new Abstract: Open-vocabulary detection (OVD) recognizes categories unseen during training through textual category queries, yet achieving strong generalization with real-time efficiency remains challenging. Beyond vocabulary scaling, zero-shot g…

  101. arXiv cs.CV TIER_1 English(EN) · H M Dipu Kabir, Subrota Kumar Mondal, Mohammad Ali Moni ·

    Batch Augmentation with Unimodal Fine-tuning for Multimodal Fusion of Large Language Models

    arXiv:2505.06592v2 Announce Type: replace Abstract: In this paper, we propose batch augmentation with unimodal fine-tuning for multimodal learning. We start with pre-trained unimodal models. We fine-tune the unimodal models with the application data. After that, we form a Multi-L…

  102. arXiv cs.CV TIER_1 English(EN) · Pengcheng Zheng, Chaoning Zhang, Jiaxin Yan, Sihan Cao, Jianwei Zhang, Xudong Wang, Jiaquan Zhang, Jewon Lee, Tae-Ho Kim, Yang Yang, Heng Tao Shen ·

    MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models

    arXiv:2610.08830v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities across vision and language tasks. However, their massive computational and memory demands hinder real-world deployment. While recent effort…

  103. Latent Space (podcast video) TIER_1 English(EN) · Latent Space ·

    Recursive Language Models — Alex Zhang, MIT PhD

    From GPU kernels and KernelBench to Recursive Language Models, agent harnesses, and massive multi-agent swarms, Alex Zhang is exploring how much capability we’re leaving on the table by wrapping increasingly powerful models in primitive systems. In this episode, the MIT researche…

  104. Mastodon — sigmoid.social TIER_1 Italiano(IT) · [email protected] ·

    Hallucinengine: n, large language model (LLM)

    Hallucinengine: n, large language model (LLM) # LLM # LLMs # AI # Hallucination # Hallucinations

  105. Medium — MLOps tag TIER_1 English(EN) · Emin Mammadov ·

    The Model Is the Easy Part: What We Learned From Self-Hosting Large Language Models

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/geotab/the-model-is-the-easy-part-what-we-learned-from-self-hosting-large-language-models-5c4d76eaf993?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/2242/1*Wv3EHW06mzDv…

  106. Medium — fine-tuning tag TIER_1 English(EN) · Tolulade Ademisoye ·

    How I Would Fine-Tune a Small Language Model From Scratch

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://tolulade-ademisoye.medium.com/how-i-would-fine-tune-a-small-language-model-from-scratch-bce869b2b79e?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1600/1*Yfk2_5KjKpQiL7K1dDd…

  107. dev.to — LLM tag TIER_1 English(EN) · developerz.ai ·

    Practical Tips for Deploying Large Language Models in Production

    <h1> Introduction </h1> <p>Deploying large language models (LLMs) in a production environment presents a different set of challenges than running them in a notebook. Engineers need to balance latency, cost, and reliability while keeping the model up to date with the latest data. …

  108. r/LocalLLaMA TIER_1 English(EN) · /u/pmttyji ·

    [Paper] FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wz59xj/paper_factorengram_factorized_ngram_memory_with/"> <img alt="[Paper] FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models" src="https://preview.redd.it/6439m6eg5vth1.png?w…

  109. dev.to — LLM tag TIER_1 Português(PT) · Marcelo Silva ·

    AI Engineering - Study on How Language Models Work

    <p>Há um tempo estudei Engenharia de IA, usando como base o livro Engenharia de IA da Chip Huyen e fui escrevendo alguns artigos ..<br /> Agora que a série está completa, achei que valia a pena compartilhar por aqui.</p> <p>Se você está estudando IA, LLMs ou só quer entender melh…

  110. r/MachineLearning TIER_1 English(EN) · /u/cbl007 ·

    Learning to Learn a Language: in-context learning of natural language from a synthetic non-linguistic prior [R]

    <!-- SC_OFF --><div class="md"><p>Learning from data as we observe it is easy for humans, but most machine learning models have limited ability to learn from new data that they have not seen during training. Prior-fitted networks (the idea behind TabPFN) showed that a model train…

  111. dev.to — LLM tag TIER_1 English(EN) · techaiwire ·

    Aleph Alpha Kolibri 1: open German-English MoE model

    <p>German AI company Aleph Alpha released Kolibri 1 on October 3, 2026, an open-weight language model built for German and English. It has 78.1 billion parameters, but only 3.46 billion do work on each token, which keeps it fast for its size. The weights are free to download and …