English(EN)RT-DETR-World: Transferring Rich LLM Semantics to Real-Time Open-Vocabulary Detection
新研究探讨LLM的效率、安全性和多语言能力
作者PulseAugur 编辑部·[111 个来源]·
研究人员正在探索各种方法来提高大型语言模型(LLM)的效率和能力。Apple Inc. 发布了关于通过增强语言辨别能力来改进多语言语音模型的研究。其他研究侧重于参数高效微调技术,如DyPAM和GRADE;理解模型内部机制的方法,如Massive Activation Gating Channel (MAGC)和COMPASS;以及高效推理策略,如VALSE和Hybrid Latent Attention (HLA)。此外,正在开发像SAFESHIELD这样的框架,用于小型语言模型的部署时安全,并正在研究像Wasserstein-based knowledge distillation (WASD)和zero-knowledge proof of training (zkLLMPoT)这样的新颖方法,以优化LLM性能和验证。
AI
Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model’s ability to discriminate languages during pretrai…
arXiv cs.CL
TIER_1English(EN)·Mingyan Liu, Min Huang·
arXiv:2610.10304v1 Announce Type: cross Abstract: We study whether latent sequence compression of prompt prefixes preserves the capabilities that large language models rely on during inference. We introduce SemanticFold, a compression scheme that folds prefix hidden states at lea…
arXiv cs.CL
TIER_1English(EN)·Maverick Morales, Tom\'a\v{s} Dominik, Vermut Gao, Katrina Shirey, Paulius Rimkevi\v{c}ius, Aaron Schurger, Uri Maoz·
arXiv:2610.10405v1 Announce Type: cross Abstract: Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models. However, semantic chain-of-thought monitoring depends on…
arXiv:2505.19912v3 Announce Type: replace Abstract: We present Adjacent Possible Exploration (APE), a selective fine-tuning method for adapting large language models that systematically explores parameter modifications while maintaining model stability. Inspired by evolutionary o…
arXiv:2610.05366v2 Announce Type: replace Abstract: \`{I}r\`{o}y\`{i}nSpeech is a 42-hour, 80-speaker Yor\`ub\'a read-speech corpus whose audio has been distributed by ELRA since 2024. This paper describes the release of its text component: 24,905 unique, hand-verified, tone-mark…
arXiv:2610.10261v1 Announce Type: cross Abstract: Context: Once defined a taxonomy of stages structuring Machine Learning (ML) pipelines (e.g. Data Preprocessing, Modeling...), extracting these stages from source code is key for better understanding ML practices. However, the div…
arXiv:2610.06387v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved substantial performance gains through increases in model size, training data, and computational resources. However, traditional scaling approaches produce diminishing returns, rising fi…
arXiv cs.CL
TIER_1English(EN)·James Ravi Kirkpatrick, Alexandru Radulescu, Rachel Katharine Sterken·
arXiv:2610.09064v1 Announce Type: cross Abstract: When we interact with large language models (LLMs), are we having a conversation? They are designed to invite us to treat them as intelligent interlocutors who remember, act, and make commitments. But appearances deceive. We intro…
arXiv:2610.10049v1 Announce Type: new Abstract: Reasoning models allocate extra computation at inference time and present their answers as the product of deliberate thought. If this deliberation works the way dual-process accounts of human cognition suggest, longer thinking shoul…
arXiv:2610.09346v1 Announce Type: new Abstract: Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated…
arXiv cs.CL
TIER_1English(EN)·Marek \v{S}uppa, Ivan Vykopal, Andrej Ridzik, Kristi\'an Sopkovi\v{c}, Nat\'alia K\v{n}a\v{z}ekov\'a, Jaroslav Kop\v{c}an, Miroslav Bl\v{s}t\'ak, Vikt\'oria Ondrejov\'a, Daniel Hl\'adek, Michal Gregor, Martin Tamajka, Mari\'an \v{S}imko·
arXiv:2610.09152v1 Announce Type: new Abstract: Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored…
arXiv cs.CL
TIER_1English(EN)·Tobias Braun, Nils Loose, Alexander Herzog, Virginia Ceccatelli, Marcus Rohrbach, Thomas Eisenbarth, Lorenzo Cavallaro·
arXiv:2610.09087v1 Announce Type: new Abstract: Large language models are informing decisions with ever-higher stakes. As the consequences of their errors grow, a central question becomes harder to ignore: how much can we trust an individual answer? Yet recognizing when to defer …
arXiv:2610.09033v1 Announce Type: new Abstract: Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and …
arXiv cs.AI
TIER_1English(EN)·Xingru Zhou, Luis Sentis, Aarti Choudhary·
arXiv:2610.07276v1 Announce Type: cross Abstract: Deployment-time safety of language models is commonly implemented through runtime guardrails such as input moderation, routing, retrieval verification, and output filtering. Existing deployment frameworks provide increasingly capa…
arXiv:2510.22954v2 Announce Type: replace Abstract: Language models (LMs) often struggle to generate diverse, human-like creative content, raising concerns about the long-term homogenization of human thought through repeated exposure to similar outputs. Yet scalable methods for e…
arXiv:2610.08164v1 Announce Type: cross Abstract: Low-rank quantization error compensation (LQEC) recovers the accuracy lost under aggressive weight quantization by attaching a closed-form rank-$r$ adapter beside each frozen quantized weight, without any training. We show that ex…
arXiv:2610.07853v1 Announce Type: cross Abstract: Ternary language models such as BitNet b1.58, Falcon-E and BitCPM are fine-tuned with higher-precision latent weights and deployed as ternary codes produced by an export step that, in the labs' documented pipelines, first casts th…
arXiv:2610.08630v1 Announce Type: new Abstract: Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, user preferences, documents, and interaction experience. In-context learning (ICL)…
arXiv:2610.07587v1 Announce Type: new Abstract: LLM orchestration investigates how an orchestrator coordinates a group of autonomous agents to achieve common goals or maximize collective welfare. The agents are typically heterogeneous, each holding a private preference that it pu…
arXiv:2610.06940v1 Announce Type: new Abstract: Continual adaptation of language models can change their output distribution on prompts learned earlier, while retaining every old prompt-answer pair may be undesirable or impossible. We study condition-anchored generative distillat…
arXiv:2601.16390v2 Announce Type: replace-cross Abstract: Large language models exhibit strong multilingual capabilities, yet significant performance gaps persist between dominant and non-dominant languages. Prior work attributes this gap to imbalances between shared and language…
arXiv:2508.10020v2 Announce Type: replace-cross Abstract: Enhancing LLM reasoning in federated settings is nontrivial due to stringent computational, communication, and privacy constraints, especially in healthcare, where clinically consequential decisions require not only accura…
arXiv:2503.22764v3 Announce Type: replace-cross Abstract: The large language model (LLM) is typically integrated into the mainstream optimization protocol. However, it remains underexplored whether maintaining the model integrity is \textit{indispensable} for promising performanc…
arXiv:2511.20471v3 Announce Type: replace Abstract: Recent advances in Large Language Model (LLM) reasoning have improved conventional problem solving, but creative reasoning remains comparatively underexplored. Inspired by cognitive science, we formalize combinational, explorato…
arXiv:2610.07940v1 Announce Type: cross Abstract: Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T. The larger cache limits how many sequences a GPU can deco…
arXiv:2610.07848v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning (PEFT) has become a standard approach for adapting large language models to downstream tasks. However, most existing PEFT methods rely on uniform and static adaptations, without accounting for the s…
arXiv cs.AI
TIER_1English(EN)·Byeonghu Na, Donghyeok Shin, Yeongmin Kim, Mina Kang, Il-Chul Moon·
arXiv:2610.07706v1 Announce Type: cross Abstract: Autoregressive large language models (LLMs) have rapidly advanced in capability, but their increasing scale comes with substantial computational and memory costs at inference time. Knowledge distillation (KD) offers a practical so…
arXiv:2610.07553v1 Announce Type: cross Abstract: LoRA fine-tuning adapts small language models (SLMs) to heterogeneous instruction data within a low-rank update subspace, making it vulnerable to three structural problems: conflicting gradients that cancel, static data selection …
arXiv:2610.08258v1 Announce Type: new Abstract: Auditing the claimed outcomes of large language model (LLM) training is challenging when model weights and training data are private, while cryptographically proving the full training process is prohibitively expensive at Transforme…
arXiv cs.AI
TIER_1English(EN)·Minjia Mao, Shi Chen, Bowen Yin, Xiao Fang·
arXiv:2610.07661v1 Announce Type: new Abstract: Massive activations, a phenomenon in which a small number of hidden channels exhibit exceptionally large magnitudes, are pervasive in large language models (LLMs). However, the mechanism by which a token develops massive activations…
arXiv:2610.07606v1 Announce Type: new Abstract: This paper establishes a theoretical framework for vertical adaptive layer skipping, proving three foundational results: (i) an Expected FLOPs formula (theorem 2) giving a closed-form expression for the computational cost of arbitra…
arXiv:2610.08341v1 Announce Type: cross Abstract: Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic de…
arXiv:2610.07247v1 Announce Type: new Abstract: Large language models have shown strong reasoning capabilities, but their high inference costs make knowledge distillation an important approach for transferring such capabilities to compact models in resource-constrained scenarios.…
Low-rank quantization error compensation (LQEC) recovers the accuracy lost under aggressive weight quantization by attaching a closed-form rank-$r$ adapter beside each frozen quantized weight, without any training. We show that existing compensators are limited by two shared simp…
This paper establishes a theoretical framework for vertical adaptive layer skipping, proving three foundational results: (i) an Expected FLOPs formula (theorem 2) giving a closed-form expression for the computational cost of arbitrary per-sample skip schedules as a function of la…
Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, user preferences, documents, and interaction experience. In-context learning (ICL) and ICL-based agent harness remain flexible, bu…
arXiv:2610.02705v1 Announce Type: cross Abstract: The Muon optimizer derives its update rule for hidden linear layers by solving a local linearization of the loss penalized by the spectral norm, motivated by an RMS-stability argument for dense linear layers. Standard Muon impleme…
arXiv cs.AI
TIER_1English(EN)·Mohsen Larni (Department of Computer Science, University of Nevada, Las Vegas), Sobhan Ebrahimi Azar (Department of Computer Science, University of Nevada, Las Vegas), Pouyan Nahed (Department of Computer Science, University of Nevada, Las Vegas), Kazem …·
arXiv:2610.03329v1 Announce Type: cross Abstract: Large language models are increasingly used where small syntactic errors matter, yet character-level reasoning is still evaluated mostly through isolated probes and aggregate accuracy. We introduce SyntaxBench, a diagnostic benchm…
arXiv cs.AI
TIER_1English(EN)·Shuyang Cao, Karthik Radhakrishnan, David Rosenberg, Steven Lu, Pengxiang Cheng, Lu Wang, Shiyue Zhang·
arXiv:2505.21870v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) generally enhances large language models' (LLMs) ability to solve knowledge-intensive tasks. But RAG could also lead to performance degradation due to imperfect retrieval and the model'…
arXiv:2610.02766v1 Announce Type: new Abstract: Retaining language-model prefix states trades recomputation against storage time. Optimizing each cached block independently can overcount savings: a resident block is usable only when the required preceding prefix is also available…
arXiv cs.LG
TIER_1English(EN)·Nicolas Martorell, Wendy Brau, Gonzalo A. Heredia, Tom\'as Pablo Korenblit, Gaspar Labasti\'e, Tom\'as Gimenez Molina·
arXiv:2610.02303v1 Announce Type: new Abstract: Language models increasingly assist people with power-related requests, so systematic differences in whom they help could shift the distribution of power at scale, or be exploited by users who learn which identities are refused less…
arXiv:2610.03199v1 Announce Type: cross Abstract: Large language models fine-tuned from a shared base can be merged by averaging their task vectors, but some merges collapse far below the base model, and common merge operators give no warning before evaluation. We show that one s…
arXiv:2610.02911v1 Announce Type: cross Abstract: Methods for training language models on stale samples are judged by comparisons against importance-corrected baselines. We show that details of the experimental harness can reverse the observed ranking of methods, and we introduce…
arXiv:2610.02632v1 Announce Type: new Abstract: As large language models are increasingly deployed for multi-step reasoning, verifying the correctness of their outputs has become essential for maintaining reliability at scale. Verifying the correctness of large language model out…
arXiv:2610.02432v1 Announce Type: cross Abstract: Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM r…
arXiv:2610.02986v1 Announce Type: new Abstract: Membership inference on large language models (LLMs) aims to determine whether a given text sample was included in an LLM's training data, without access to its training corpus. Despite recent progress, existing benchmarks suffer fr…
arXiv:2610.02856v1 Announce Type: new Abstract: Multi-task post-training of large language models (LLMs) aims to improve performance across tasks with unequal amounts of training data. Existing methods focus primarily on balancing task contributions during single-model training. …
arXiv cs.LG
TIER_1English(EN)·Zhenghao Zhao, Gaowen Liu, Zhiling Lan, Yan Yan·
arXiv:2610.02593v1 Announce Type: new Abstract: Data selection is already a central bottleneck in large-language-model training, where web-scale corpora are noisy and token budgets are finite. In continual pre-training (CPT), it becomes a forgetting-control problem: a poorly chos…
arXiv cs.AI
TIER_1English(EN)·Han Wang, Ishwar B Balappanawar, Huan Zhang·
arXiv:2610.02549v1 Announce Type: new Abstract: Time and event expression extraction are fundamental temporal reasoning tasks, but the problem remains difficult due to annotation ambiguity, domain sensitivity, and unstable model behavior. Existing evaluations focus on in-domain p…
arXiv cs.AI
TIER_1English(EN)·Zizhuo Zhang, Xiong Peng, Jingwei Sun, Rong Yao, Shixiong Kai, Mingxuan Yuan, Bo Han·
arXiv:2610.02695v1 Announce Type: cross Abstract: Reliable large language models (LLMs) must not only produce accurate answers but also express confidence that faithfully reflects their probability of being correct. Such calibration is essential for identifying uncertain predicti…
arXiv:2610.01560v1 Announce Type: cross Abstract: Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but g…
arXiv:2610.00265v1 Announce Type: new Abstract: Large Language Bayes (LLB) answers an informal modelling question by sampling candidate probabilistic programs from a language model, running approximate inference on each, and averaging them with weights proportional to an exponent…
arXiv cs.LG
TIER_1English(EN)·Bangji Yang, Jingyuan Li, Jiajun Fan, Yi Evie Zhang, Ruihan Guo, Hongba Ma, Neil He, Chumeng Liang, Qinglong Zheng, Zhanghan Ni, Ge Liu·
arXiv:2610.01110v1 Announce Type: new Abstract: How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge s…
arXiv:2610.01165v1 Announce Type: new Abstract: Layer interventions are widely used to probe the internal organization of language models, yet most analyses examine a single training checkpoint even though model representations and computations evolve throughout pretraining. This…
arXiv:2610.01592v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly applied to process streaming data, with practitioners relying on benchmarks to select the best model even though these signals only approximate real performance. While oracle annotatio…
arXiv:2508.13408v3 Announce Type: replace Abstract: Chemical Language Models (CLMs) are increasingly used in de novo drug design, driven by recent growth in model scale, compute, and dataset size. However, the relationship between design choices, training dynamics, and downstream…
arXiv cs.LG
TIER_1English(EN)·Chayne Thrash, Kevin Chen, Soheil Kolouri·
arXiv:2609.35579v2 Announce Type: replace Abstract: Residual stream pruning methods reduce inference cost by shrinking the model's hidden dimension, but existing approaches typically choose these dimensions by minimizing activation reconstruction error. This criterion implicitly …
Numerical Linear Algebra (NLA) has consistently played a vital role in advancing science by providing tools to solve fundamental problems encountered in scientific and engineering applications. Over the decades, it has continually evolved to meet the demands driven by successive …
A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which …
arXiv:2610.00694v1 Announce Type: cross Abstract: Compression reports summarize how far a compressed language model moved from the dense one, usually by a KL divergence; a deployment that relies on the dense model's outputs needs to know how many of its decisions changed. We show…
arXiv:2610.00673v1 Announce Type: cross Abstract: Looped language models increase effective depth by repeatedly applying a shared block of layers, but existing large-scale recipes require multi-stage training over trillions of tokens, while the benefits of recurrence remain diffi…
arXiv:2610.00568v1 Announce Type: cross Abstract: Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tens…
arXiv cs.AI
TIER_1English(EN)·Timoth\'ee Lesort, Alejandra L\'opez de Aberasturi G\'omez, Tristan Karch, Tom Veniat, Philippe Modard, Karl Tuyls, Ludovic Denoyer·
arXiv:2610.00416v1 Announce Type: new Abstract: Evaluating large language models becomes increasingly challenging as their capabilities advance: benchmarks can saturate, public test sets risk contamination, and assessing harder tasks can require expensive grading or execution inf…
arXiv cs.AI
TIER_1English(EN)·Oguzhan Baser, Elahe Sadeghi, Eric Wang, Nico Vergauwen, Sam Kazemian, Hong Kang, Sandeep P. Chinchali, Sriram Vishwanath·
arXiv:2602.12630v2 Announce Type: replace-cross Abstract: Most large language models (LLMs) run on external clouds: users send a prompt, pay for inference, and must trust that the remote GPU executes the LLM without any adversarial tampering. We critically ask how to achieve veri…
arXiv cs.AI
TIER_1English(EN)·Jiangfeng Chen, Xinyu Wang, Tianshuo Yan, Hanwei Wu, Xiao-Wen Chang, Yang Zhang, Lei Ding·
arXiv:2610.00717v1 Announce Type: cross Abstract: Post-training compression of LLM attention is often formulated as independent matrix approximation, ignoring both the shared structure among attention projections and the representation shift introduced by earlier compression. We …
arXiv:2606.22676v2 Announce Type: replace Abstract: Refusal on a safety benchmark does not reveal how stable that behavior will remain after model updates. Benign downstream fine-tuning can weaken refusal, yet behavioral evaluations typically expose this fragility only after an i…
A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which …
A trainable input embedding table assigns each vocabulary item an independently adjustable vector. We investigate whether this token-specific parameterization is required for substantial language-modeling capability, or whether a shared Transformer can learn from fixed token iden…
arXiv:2511.00763v3 Announce Type: replace Abstract: We investigate the performance of large language models (LLMs) on repetitive deterministic prediction tasks and study how the sequence accuracy rate (SAR) scales with output length. Each such task involves the repetition of the …
arXiv:2609.39869v1 Announce Type: new Abstract: Grammar-constrained generation guarantees syntactic validity, but can substantially degrade semantic quality when the model's preferred outputs are poorly aligned with the imposed grammar. This trade-off is particularly severe when …
arXiv:2609.38261v1 Announce Type: cross Abstract: In this paper we propose NinaXander, a series of composed language models obtained by connecting layers of frozen language models from different architecture families with a single trained shared-latent adapter. A composed model r…
arXiv:2609.39321v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) has emerged as a memory-efficient reinforcement fine-tuning (RFT) technique for reasoning-intensive tasks. How- ever, GRPO training dynamics on small language models (SLMs) remain poorly u…
arXiv cs.AI
TIER_1English(EN)·Zhen Liang, Hai Huang, Wentao Chen·
arXiv:2609.39902v1 Announce Type: cross Abstract: Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety gen…
arXiv cs.AI
TIER_1English(EN)·Hoa Quynh Nhung Nguyen, Jacopo Staiano, Michael Sullivan·
arXiv:2609.40121v1 Announce Type: cross Abstract: While Abstract Meaning Representation (AMR) has historically improved performance on a range of NLP tasks, the benefit---or lack thereof---of AMR augmentation for modern LLMs is thus far unclear. In this paper, we attempt to repro…
arXiv:2603.26164v2 Announce Type: replace-cross Abstract: Data-centric training has emerged as a promising direction for improving large language models (LLMs) by optimizing not only model parameters but also the selection, composition, and weighting of training data during optim…
arXiv:2511.10661v2 Announce Type: replace Abstract: It is increasingly important to evaluate the characteristics of systems based on large language models (LLMs). Evaluations in this context often rely on a curated benchmark set of input prompts provided to the LLM, where the out…
arXiv cs.CL
TIER_1English(EN)·Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov·
arXiv:2410.02343v2 Announce Type: replace Abstract: Large language models (LLMs) routinely fail to output the correct option in multiple-choice question answering (MCQA) while encoding the answer internally. We expose this latent knowledge via the Query--Key (QK) score, defined f…
arXiv:2609.40185v1 Announce Type: new Abstract: Constrained generation aims to sample from language models (LMs) conditioned on hard constraints. Existing constrained-generation techniques for nondeterministic finite automaton (NFA) constraints either distort the distribution or …
arXiv cs.CL
TIER_1English(EN)·Jhen-Ke Lin, Chung Chun Wang·
arXiv:2609.38612v1 Announce Type: new Abstract: As natural language drives more applications, language models increasingly run inside programs as decision components: the program sends them the current state and acts on the returned decision until a newer one arrives. When eviden…
arXiv cs.CL
TIER_1English(EN)·Parisa Salmani, Peter R. Lewis·
arXiv:2609.38357v1 Announce Type: new Abstract: Conversational safety evaluations often test language models with a single harmful prompt, even though real-world systems interact with users through long, adaptive conversations. This study examines whether models continue to respo…
arXiv cs.CL
TIER_1(CA)·Juan M Zambrano Chaves, Peniel Argaw, Risa Ueno, Carlo Bifulco, Kristina Young, Rom Leidner, Tristan Naumann, Hoifung Poon·
arXiv:2609.38181v1 Announce Type: new Abstract: Survival analysis estimates time-to-event outcomes from patient covariates and is widely used for medical risk assessment. Patients seeking prognostic information after a diagnosis may turn to large language models (LLMs), now readi…
arXiv cs.AI
TIER_1English(EN)·Kenan Alkiek, David Jurgens, Vinod Vydiswaran·
arXiv:2510.13935v3 Announce Type: replace-cross Abstract: The facts a language model stores are tied to its parameter count, so small models that fit on edge devices fail on expert problems, which need specialized knowledge and follow multi-step procedures. Fine-tuning for a spec…
arXiv cs.AI
TIER_1English(EN)·Mansi Sakarvadia, Aswathy Ajith, Arham Khan, Nathaniel Hudson, Caleb Geniesse, Kyle Chard, Yaoqing Yang, Ian Foster, Michael W. Mahoney·
arXiv:2410.02159v3 Announce Type: replace-cross Abstract: Language models (LMs) can "memorize" information, i.e., encode training data in their weights in such a way that inference-time queries can lead to verbatim regurgitation of that data. This ability to extract training data…
arXiv cs.MA (Multiagent)
TIER_1English(EN)·Roberto Di Pietro·
Large Language Model (LLM) agents are increasingly deployed as populations of interacting entities, in which consensus --agreement on a shared answer-- emerges as a collective, unengineered behaviour. Prior work on LLM consensus shows that agents can cross-verify their answers an…
arXiv cs.CL
TIER_1English(EN)·Elsayed Eshra, Ali Al-Lawati, Dongwon Lee, Suhang Wang·
arXiv:2609.33638v2 Announce Type: replace-cross Abstract: We introduce RareTrap, a framework for estimating the probability of severe behaviors in black box large language models (LLMs). A key challenge for probability estimation is defining a tractable distribution over the inpu…
arXiv:2609.37076v1 Announce Type: new Abstract: Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this issue, existing unlearning methods typically rely on training-based parameter updat…
arXiv cs.CL
TIER_1English(EN)·Prosper Arineitwe Asiimwe, Francois Meyer, Jan Buys·
arXiv:2609.37543v1 Announce Type: new Abstract: Cross-lingual zero-shot transfer and multilingual fine-tuning are promising approaches for NLP tasks such as Named Entity Recognition (NER) in low-resource languages, but in the absence of target language benchmarks, it is unclear w…
arXiv:2602.13540v2 Announce Type: replace-cross Abstract: Accurate confidence estimation is critical for reliable use of large language models (LLMs). Prior work on LLM calibration largely focuses on response-level confidence, which estimates the correctness of a single generated…
arXiv:2406.00083v3 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by retrieving relevant information from external knowledge bases to provide more accurate, contextually informed, and up-to-date responses. However…
arXiv cs.AI
TIER_1English(EN)·Rebecca Ramnauth, Brian Scassellati·
arXiv:2609.37371v1 Announce Type: cross Abstract: Model adaptation is typically governed by a fixed recipe, even though different update programs can produce substantially different behavioral outcomes. We introduce adaptation compilation, which reframes where, how, and to what e…
arXiv cs.AI
TIER_1English(EN)·Kenan Alkiek, Moontae Lee, David Jurgens, V. G. Vinod Vydiswaran·
arXiv:2609.38006v1 Announce Type: new Abstract: Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud mo…
arXiv cs.AI
TIER_1Deutsch(DE)·Yuansen Liu, Yixuan Tang, Anthony Kum Hoe Tung·
arXiv:2609.37700v1 Announce Type: new Abstract: Language models expose internal signals that predict whether an answer is correct, readable from a single forward pass of a frozen model without additional generations. Yet existing probes often commit to one signal family or layer …
Language models expose internal signals that predict whether an answer is correct, readable from a single forward pass of a frozen model without additional generations. Yet existing probes often commit to one signal family or layer and can be brittle under distribution shift; in …
arXiv:2501.09223v3 Announce Type: replace-cross Abstract: This is a book about large language models. As indicated by the title, it primarily focuses on foundational concepts rather than comprehensive coverage of all cutting-edge technologies. The book is structured into six main…
arXiv:2609.29504v1 Announce Type: cross Abstract: Aggregate factuality scores hide where a language model succeeds, which relations it confuses, and whether an answer survives innocuous changes to the question or decoder. We introduce PROOF, a profile-oriented benchmark for factu…
We study a practical question: can a small correction module fix errors in a frozen language model's outputs without degrading its base capabilities? We propose CRN v2, a lightweight logit-level correction module (~34M trainable parameters, 0.73% of the 4.65B text module) that si…
arXiv:2505.06592v2 Announce Type: replace Abstract: In this paper, we propose batch augmentation with unimodal fine-tuning for multimodal learning. We start with pre-trained unimodal models. We fine-tune the unimodal models with the application data. After that, we form a Multi-L…
arXiv cs.CV
TIER_1English(EN)·Pengcheng Zheng, Chaoning Zhang, Jiaxin Yan, Sihan Cao, Jianwei Zhang, Xudong Wang, Jiaquan Zhang, Jewon Lee, Tae-Ho Kim, Yang Yang, Heng Tao Shen·
arXiv:2610.08830v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities across vision and language tasks. However, their massive computational and memory demands hinder real-world deployment. While recent effort…
Latent Space (podcast video)
TIER_1English(EN)·Latent Space·
From GPU kernels and KernelBench to Recursive Language Models, agent harnesses, and massive multi-agent swarms, Alex Zhang is exploring how much capability we’re leaving on the table by wrapping increasingly powerful models in primitive systems. In this episode, the MIT researche…
<h1> Introduction </h1> <p>Deploying large language models (LLMs) in a production environment presents a different set of challenges than running them in a notebook. Engineers need to balance latency, cost, and reliability while keeping the model up to date with the latest data. …
<p>Há um tempo estudei Engenharia de IA, usando como base o livro Engenharia de IA da Chip Huyen e fui escrevendo alguns artigos ..<br /> Agora que a série está completa, achei que valia a pena compartilhar por aqui.</p> <p>Se você está estudando IA, LLMs ou só quer entender melh…
<!-- SC_OFF --><div class="md"><p>Learning from data as we observe it is easy for humans, but most machine learning models have limited ability to learn from new data that they have not seen during training. Prior-fitted networks (the idea behind TabPFN) showed that a model train…
<p>German AI company Aleph Alpha released Kolibri 1 on October 3, 2026, an open-weight language model built for German and English. It has 78.1 billion parameters, but only 3.46 billion do work on each token, which keeps it fast for its size. The weights are free to download and …