English(EN)Understanding Alignment in Multimodal LLMs: A Comprehensive Study
新的大语言模型研究涵盖多模态对齐、推理审计和能源使用 · 已追踪 10 个来源
作者PulseAugur 编辑部·[441 个来源]·
近期研究探讨了大语言模型(LLM)能力和局限性的各个方面。一项研究调查了多模态大语言模型的对齐问题,提出了一种新的数据生成方法来提高图像-文本一致性。另一篇论文介绍了一种用于评估大语言模型推理的协议级审计,突出了基准分数与实际行为属性之间的差异。进一步的研究利用围棋死活题探究了大语言模型的推理效率,揭示了当前模型在搜索组织方面存在困难。此外,还有研究考察了大语言模型的能耗、默认设置和代币定价对推理服务的影响,以及低资源编程语言基准的开发。
AI
Large language models (LLMs) exhibit a wide range of human-like behaviors, from expressing thoughts and emotions, to engaging in relationship-building with users, to refusing requests and maintaining boundaries. Despite their prevalence, researchers and practitioners lack methods…
Apple Machine Learning Research
TIER_1English(EN)·
Preference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains comparatively underexplored. Similar to language models, MLLMs for image understanding tasks encoun…
arXiv cs.AI
TIER_1English(EN)·Stephen Mell, Botong Zhang, David Mell, Shuo Li, Ramya Ramalingam, Nathan Yu, Stephan Zdancewic, Osbert Bastani·
arXiv:2506.12202v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often call external tools to solve tasks. One effective strategy is for LLMs to write code, enabling them to use complex control flow such as conditionals and loops. Such code actions are typic…
arXiv:2508.15813v2 Announce Type: replace-cross Abstract: A big issue in modern LLM applications is they tend to feed long context to LLM, which results in high inference cost and latency, and may exceed the context limit. Prompt compression addresses this issue by reducing the l…
arXiv cs.AI
TIER_1English(EN)·Erik Thureck, Robert K\"uhnen, Tim Jacobowitz·
arXiv:2608.21074v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in research workflows and software development pipelines, yet their output remains sensitive to input prompt variations. This paper presents $\unicode{x00AB}$PromptResponse$\unico…
arXiv:2608.19207v1 Announce Type: new Abstract: Production deployments of Multimodal Large Language Models (MLLMs) increasingly rely on system messages to govern model behavior. Yet existing benchmarks either evaluate constraints in text only or embed them into the user turn, lea…
arXiv:2608.19677v1 Announce Type: cross Abstract: Prefix caching avoids prefill only when a repeated request returns to a server that still holds the prefix KV. Cache-blind balancing disperses that reuse; fixed affinity preserves it but can overload a server. CacheRoute resolves …
arXiv:2608.19758v1 Announce Type: new Abstract: Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. Our previous work, FlashPr…
arXiv:2608.19662v1 Announce Type: new Abstract: Agentic language models repeatedly encode tool and skill schemas that recur across requests in different combinations and orders, preventing standard prefix caching from reusing their key--value (KV) states. We introduce \textbf{ReC…
arXiv:2608.19395v1 Announce Type: cross Abstract: Hybrid Transformer-Mamba large language models (LLMs) enhance long-context efficiency, but their heterogeneous computation and communication patterns complicate efficient hardware acceleration. Chiplet-based architectures offer a …
arXiv:2608.20202v1 Announce Type: new Abstract: Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, s…
arXiv:2608.19889v1 Announce Type: new Abstract: The entire ecosystem of open-source language models effectively relies on a single platform. What if this platform was forced to shut down tomorrow? Implementing and maintaining efficient model definitions and translating them betwe…
Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking …
arXiv:2608.18149v1 Announce Type: cross Abstract: Energy-aware LLM serving requires comparing configurations under realistic request shapes, yet exhaustive target-GPU profiling is costly and a cheap predictor can be dangerously confident outside its measured scope. We present Tok…
arXiv:2608.18795v1 Announce Type: cross Abstract: Majority voting over multiple LLM samples is widely used to raise answer accuracy, yet its gain varies erratically: on hard questions it can even backfire. This paper gives a quantitative account of this failure. A pluralistic agr…
arXiv cs.AI
TIER_1English(EN)·Pattaraphon Kenny Wongchamcharoen, Kris Gulati, Min Min Fong, Abhishek Nagaraj·
arXiv:2608.18554v1 Announce Type: cross Abstract: Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model…
arXiv:2608.18539v1 Announce Type: cross Abstract: The remarkable capabilities of large language models (LLMs) are often undermined by their instability. Even subtle and semantically irrelevant changes in prompts can cause dramatic fluctuations in performance, a phenomenon known a…
arXiv:2608.18158v1 Announce Type: cross Abstract: LLMs have been increasingly used to catch data quality issues automatically, but we know very little about how consistent these judgments actually are. This study tests an LLM on two e-commerce data quality tasks, entity matching …
Retrieved memories can induce reasoning errors and belief distortions in large language models, and an inference-time strategy helps avoid these cognitive traps while maintaining benchmark performance.
Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most im…
arXiv:2608.17515v1 Announce Type: cross Abstract: Background: Large Language Models (LLMs) are increasingly being applied to Software Engineering (SE) tasks, achieving high accuracy across problems such as clone detection, vulnerability prediction, and code summarization. However…
arXiv:2608.18027v1 Announce Type: new Abstract: Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative e…
arXiv:2605.09041v2 Announce Type: replace Abstract: LLM bias scores can depend on audit design. We introduce BiAxisBias, a prespecified audit varying task, role, perspective, sentiment, and wording over 200 stereotype statements while retaining forced Selection and Rationale as s…
arXiv cs.CL
TIER_1English(EN)·Marianne Menglin Liu, Leonid Boytsov, Daniel W. Peterson, Pramuditha Perera, Rongguang Wang, Sai Ashish Somayajula, Syed Hamza Rafique, Rohit Saini, Shubham Pathak, Sujeeth Bharadwaj, Tao Sheng, Graham Horwood, Fahad Shah, Ankan Bansal, Sujith Ravi, Dan …·
arXiv:2608.15931v1 Announce Type: new Abstract: We present PLSQLBench, to our knowledge the first benchmark for evaluating whether LLMs can write executable PL/SQL programs, with correctness measured through execution-based tests. Existing LLM evaluations largely target general-p…
arXiv:2608.15964v1 Announce Type: cross Abstract: Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the language of the prompt. In other words, a language model outputs different (and potentially incorrect) responses …
arXiv:2608.14999v1 Announce Type: cross Abstract: Constructing special graphs is an important task within graph theory and computer science. Many popular graph constructions are the result of a comprehensive exploration of relevant graphs and human ingenuity. Given the rise of ge…
arXiv:2608.14626v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual settings than in high-resource languages. In this paper,…
arXiv:2608.16438v1 Announce Type: new Abstract: In a world where valuable artifacts are increasingly created, completed, or processed by LLMs, the central economic question is not only what the LLM can produce, but what \emph{value} remains in the inputs (i.e., the prompts) we pr…
arXiv:2608.15592v1 Announce Type: new Abstract: Efficient LLM serving is often bottlenecked by the need to pad sequences to a fixed maximum length, and this wastes compute and degrades throughput. Predicting output lengths in advance makes it possible to adopt length-aware schedu…
arXiv:2608.15451v1 Announce Type: new Abstract: Large language models process large amounts of information but usually lack an explicit mechanism for maintaining compact and evolving conceptual representations. We introduce Mental Model Management (3M), a framework in which knowl…
arXiv cs.AI
TIER_1English(EN)·Zahra Fazel, Sunanda Gamage, Shayan Shirahmad Gale Bagi, Amir H. Ashouri, Tomasz S. Czajkowski, Bryan Chan, Reza Azimi, Yaoqing Gao·
arXiv:2608.14953v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have opened opportunities to apply high-level code transformations to the field of code optimization, and it has since emerged as one of the most fundamental tasks for LLMs to perform;…
Large language models improve continuously through iterative test-time feedback loops called Chain-of-Experience, outperforming zero-shot baselines with lower cost and higher token efficiency.
arXiv:2608.13612v1 Announce Type: new Abstract: Natural-language interfaces to enterprise data must translate underspecified requests into governed, executable behavior while controlling invalid queries, policy failures, cost, and nondeterminism. SemPlan Benchmark evaluates this …
arXiv cs.AI
TIER_1English(EN)·William Nixon, Jon Durbin, Florian Standhartinger, Haryadi S. Gunawi, Juncheng Yang·
arXiv:2608.13573v1 Announce Type: new Abstract: Large Language Model (LLM) serving has become a critical cloud workload, and realistic traces are essential for motivating and benchmarking serving systems. However, existing LLM serving workload studies remain limited in scale and …
arXiv:2608.14425v1 Announce Type: new Abstract: LLM evaluations often use fixed sampling budgets, testing every item the same number of times even after estimates are precise. We introduce optstop, a precision-based adaptive stopping framework that treats evaluation as a sequenti…
arXiv cs.AI
TIER_1English(EN)·Yiderigun Borjigin, Alexander Hermann, Christian Cyron, Roland Aydin·
arXiv:2608.14320v1 Announce Type: new Abstract: The anchoring effect is a cognitive bias in which an initial reference value shifts a later judgment toward itself. This effect is well established in human judgment and decision-making, and recent work suggests that large language …
Large language models process large amounts of information but usually lack an explicit mechanism for maintaining compact and evolving conceptual representations. We introduce Mental Model Management (3M), a framework in which knowledge is represented as mental models consisting …
HOTFIXR is a data generation framework that targets multilingual reasoning weaknesses to improve cross-lingual performance without sacrificing overall capability.
arXiv cs.AI
TIER_1English(EN)·Diego Manya, Ethan I. Thorpe, Ji Zhang, Myranda Shirk, Jiamian He, Angel Hsu, Michael P. Vandenbergh·
arXiv:2608.12350v1 Announce Type: cross Abstract: The energy demand growth and environmental impacts of artificial intelligence (AI) have generated substantial interest in supplying sufficient low-cost electricity for AI-driven data center development. Research on the ability of …
arXiv cs.AI
TIER_1English(EN)·Conor F. Hayes, Elliot Meyerson, Kajetan Schweighofer, Roberto Dailey, Babak Hodjat, Risto Miikkulainen, Xin Qiu·
arXiv:2608.12679v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science. The usual approach is to present the problem to the model and use its answer as the proposed solution. However, beyond this best g…
arXiv:2608.13221v1 Announce Type: new Abstract: The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and allocate reasoning resources--that is, how they organiz…
arXiv:2401.11641v5 Announce Type: replace Abstract: In recent years, Large Language Models (LLMs) like ChatGPT have seen considerable advancements and have been applied in diverse fields. Built on the Transformer architecture, these models are trained on extensive datasets, enabl…
arXiv cs.CL
TIER_1English(EN)·Junhao Luo (School of Statistics,Data Science, Southwestern University of Finance,Economics), Ning Huang (School of Statistics,Data Science, Southwestern University of Finance,Economics), Ziqi Sha (School of Statistics,Data Science, Southwestern Universi…·
arXiv:2608.13326v1 Announce Type: new Abstract: LLM benchmark scores can be precise even when the observation protocol does not identify the behavioral property they are intended to measure. In a controlled, solver-grounded setting, we formalize a protocol-level identifiability a…
arXiv:2608.13304v1 Announce Type: new Abstract: Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refused. We propose Wrapper-Based Intent-Form Augmentatio…
arXiv:2603.14501v2 Announce Type: replace-cross Abstract: Large Language Models excel in high-resource programming languages but struggle with low-resource ones. Existing research related to low-resource programming languages primarily focuses on Domain-Specific Languages (DSLs),…
arXiv cs.AI
TIER_1English(EN)·Dananjay Srinivas, Saksham Khatwani, Maria Pacheco·
arXiv:2608.13484v1 Announce Type: cross Abstract: When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rather than backing off to safer, more general claims. We frame this failure through a Gricean lens: a cooperative spe…
arXiv:2608.13315v1 Announce Type: cross Abstract: We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the allocation, or exit. Larger allocations can i…
We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the allocation, or exit. Larger allocations can improve accuracy but increase token cost and latenc…
arXiv cs.AI
TIER_1English(EN)·Vu Duc Anh, Nhat M. Hoang, Do Xuan Long, Cong-Duy Nguyen, Ponhvoan Srey, Luu Anh Tuan·
arXiv:2608.11573v1 Announce Type: cross Abstract: Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we propose Self-Fix Step-DPO (SFS-DPO), a reinforcement learn…
arXiv:2608.11705v1 Announce Type: new Abstract: Aligned large language models (LLMs) are expected to exhibit safety behavior based on the content of the user request: they should refuse unsafe requests and comply with safe ones. However, we show that the same request can elicit s…
arXiv cs.AI
TIER_1English(EN)·Vincent Siu, Nicholas Crispino, David Park, Nathan W. Henry, Zhun Wang, Yang Liu, Dawn Song, Chenguang Wang·
arXiv:2509.13450v3 Announce Type: replace Abstract: We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets. While prior work highlights the general capabilities of representation steering, we fo…
arXiv cs.AI
TIER_1English(EN)·Shailja Thakur, Sungeun An, Chad DeLuca, Hima Patel·
arXiv:2608.11694v1 Announce Type: cross Abstract: A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked, but it does not. We show that rephrasing a problem while k…
arXiv cs.AI
TIER_1English(EN)·Huiqi Miao, Xinbao Sun, Bo Wang, Fanyu Meng, Lijun Mei, Na Wu, Di Jin, Chao Deng, Junlan Feng·
arXiv:2608.11584v1 Announce Type: new Abstract: Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.8% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap. Existing benchmarks …
arXiv:2608.11408v1 Announce Type: new Abstract: Prior white-box studies show that large language models can retain latent traces of target knowledge after unlearning, even when the knowledge is no longer expressed in their outputs. However, existing audits remain limited to one-o…
arXiv:2604.17244v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents for sequential decision-making struggle to produce diverse outputs. This leads to insufficient exploration, suboptimal solutions, and repeated actions. Actions are generated at the sequenc…
arXiv cs.AI
TIER_1English(EN)·Duo Lu, Andrew Crotty, U\u{g}ur \c{C}etintemel·
arXiv:2604.14401v2 Announce Type: replace Abstract: Agentic AI systems are becoming commonplace in domains that require long-lived, stateful decision-making in continuously evolving conditions. As such, correctness depends not only on the output of individual model calls, but als…
Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science. The usual approach is to present the problem to the model and use its answer as the proposed solution. However, beyond this best guess, discovery can be enhanced by increasing te…
Natural-language interfaces to enterprise data must translate underspecified requests into governed, executable behavior while controlling invalid queries, policy failures, cost, and nondeterminism. SemPlan Benchmark evaluates this architectural design space with a deterministic …
arXiv cs.AI
TIER_1English(EN)·William Lugoloobi, Thomas Foster, William Bankes, Chris Russell·
arXiv:2602.09924v4 Announce Type: replace-cross Abstract: Running LLMs with extended reasoning on every problem is expensive, but determining which inputs actually require additional compute remains challenging. We investigate whether their own likelihood of success is recoverabl…
arXiv:2608.10042v1 Announce Type: cross Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization. We introduce UserToolBench , a benchmark …
arXiv:2608.11047v1 Announce Type: new Abstract: While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains comparatively less explored. We introduce V-FiLLM, a framework that generates financia…
arXiv cs.AI
TIER_1English(EN)·Haoze Liu, Run Liu, Haiying Xu, Jiahui Han, Siyuan Fang, Siyu Yan, Huiqi Deng, Guanchu Wang, Na Zou·
arXiv:2608.10703v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM personality studies largely rely on self-report quest…
arXiv:2608.10567v1 Announce Type: new Abstract: Analytic dashboards combine coordinated views and interactions for data exploration and decision-making. Recent models can generate them from data and natural-language goals, but evaluating their usefulness remains difficult. Dashbo…
arXiv:2608.10528v1 Announce Type: cross Abstract: Anchor-based pointwise LLM reranking scores each candidate against a shared reference passage to recover cross-document context at pointwise cost. We study when this actually helps, using GCCP/PAGC as a representative method. Our …
Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we propose Self-Fix Step-DPO (SFS-DPO), a reinforcement learning based, two-stage framework for step-level self…
Anchor-based pointwise LLM reranking scores each candidate against a shared reference passage to recover cross-document context at pointwise cost. We study when this actually helps, using GCCP/PAGC as a representative method. Our study is reproduction-first. We use reproduction a…
Anchor-based pointwise LLM reranking scores each candidate against a shared reference passage to recover cross-document context at pointwise cost. We study when this actually helps, using GCCP/PAGC as a representative method. Our study is reproduction-first. We use reproduction a…
Anchor-based pointwise LLM reranking scores each candidate against a shared reference passage to recover cross-document context at pointwise cost. We study when this actually helps, using GCCP/PAGC as a representative method. Our study is reproduction-first. We use reproduction a…
arXiv cs.LG
TIER_1English(EN)·Calvin Higgins, Marco Alvarez·
arXiv:2608.07894v1 Announce Type: new Abstract: Recent advances have highlighted the potential of machine learning, particularly Large Language Models (LLMs), for analyzing and optimizing programs. We present the first application of program embeddings from LLMCompiler---an LLM m…
arXiv:2608.09106v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. In legal practice, time is a critical concept that governs the validity of statutes, the progression of legal cases, and the enfor…
arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages. To…
arXiv:2604.01029v2 Announce Type: replace-cross Abstract: Multi-LLM revision pipelines, in which a second model reviews and improves a draft produced by a first, are widely assumed to derive their gains from genuine error correction. We question this assumption with a controlled …
arXiv cs.AI
TIER_1English(EN)·Delip Rao, Chris Callison-Burch·
arXiv:2603.00077v3 Announce Type: replace-cross Abstract: Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduced to exact programmatic checks. Yet the underlying judges remain vulnerable to …
arXiv cs.AI
TIER_1English(EN)·Adrian Marius Dumitran, Theodor-Pierre Moroianu, Mihnea-Vicentiu Buca·
arXiv:2507.03162v2 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) has transformed various domains, particularly computer science (CS) education. These models exhibit remarkable capabilities in code-related tasks and problem-solving, r…
arXiv cs.AI
TIER_1English(EN)·Yupeng Cao, Haohang Li, Weijin Liu, Wenbo Cao, Anke Xu, Lingfei Qian, Xueqing Peng, Minxue Tang, Zhiyuan Yao, Jimin Huang, K. P. Subbalakshmi, Zining Zhu, Jordan W. Suchow, Yangyang Yu·
arXiv:2604.10015v3 Announce Type: replace Abstract: Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for long-horizon financial tasks. While existing benchmarks have begun evaluating financial tool…
arXiv:2608.08467v1 Announce Type: new Abstract: The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs). A common server design embeds frequently used reference data, such as identifier lookup tables, directly in the server …
arXiv:2608.08254v1 Announce Type: new Abstract: Structured output, where an LLM populates a predefined JSON schema, has become a default mechanism for data labeling and information extraction, but it also introduces a second instruction channel through schema descriptions. We tes…
arXiv cs.AI
TIER_1English(EN)·Ananya Sahu, Mohit Bansal, Elias Stengel-Eskin·
arXiv:2608.07460v1 Announce Type: cross Abstract: While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively impacting tasks that explicitly require creativity (e.g., story generation) as we…
arXiv cs.AI
TIER_1English(EN)·Rodrigo Ferreira Rodrigues, Karim Radouane, Jose G Moreno, Lynda Tamine·
arXiv:2608.07411v1 Announce Type: new Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comp…
The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs). A common server design embeds frequently used reference data, such as identifier lookup tables, directly in the server instructions: the system-prompt text a server ha…
In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comprehensive benchmark for probing LLMs on geo-rela…
arXiv:2608.06301v1 Announce Type: new Abstract: As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes…
arXiv:2608.05420v1 Announce Type: cross Abstract: Large language models (LLMs) can generate text that resembles a mathematical proof, but resemblance does not establish correctness. A formal proof checker verifies whether each proof step follows established logical rules. Coq bas…
arXiv cs.LG
TIER_1English(EN)·Louis Mandel, Guillaume Baudart, Mandana Vaziri, Martin Hirzel·
arXiv:2608.05234v1 Announce Type: new Abstract: Building reliable applications that leverage large language models (LLMs) remains a significant challenge. While LLMs offer impressive capabilities across diverse tasks, their outputs often lack accuracy and provide no clear measure…
arXiv cs.CL
TIER_1English(EN)·Lukas Twist, Twm Stone, Helen Yannakoudakis, Jie M. Zhang·
arXiv:2608.06041v1 Announce Type: cross Abstract: Large language models (LLMs) have been shown to exhibit strong Python preferences when generating project-level code, but there is currently no systematic way to measure this behaviour across new models. To bridge this gap, we int…
arXiv:2608.06312v1 Announce Type: new Abstract: Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer…
arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing …
arXiv cs.AI
TIER_1Deutsch(DE)·Jewon Yeom, Jaewon Sok, Seonghyeon Park, Jeongjae Park, Hwiyeong Lee, Taesup Kim·
arXiv:2608.05687v1 Announce Type: cross Abstract: Masked diffusion language models (dLLMs) can commit tokens in any order -- a freedom marketed as their core advantage over autoregressive decoding. We show that on reasoning tasks this freedom is instead the axis of failure. Loggi…
arXiv cs.AI
TIER_1English(EN)·Itamar Pres, Belinda Z. Li, Laura Ruis, Zifan Carl Guo, Keya Hu, Mehul Damani, Isha Puri, Ekdeep Singh Lubana, Jacob Andreas·
arXiv:2608.05188v1 Announce Type: cross Abstract: Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models overcondition on user framing ("sycophancy"), exhibit incomplete logical generalization, and p…
Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer a representative testbed: they are lengthy, hig…
arXiv:2608.04463v1 Announce Type: new Abstract: Prior work on LLM conformity largely measures discrete answer flips under verifiable labels. Open-ended revisions require a different measurement strategy because answer quality is graded, latent, and judged imperfectly. We introduc…
arXiv cs.LG
TIER_1English(EN)·Usha Shrestha, Dmitry Ignatov, Radu Timofte·
arXiv:2601.03808v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have achieved notable performance in code synthesis; however, data-aware augmentation remains a limiting factor, handled via heuristic design or brute-force approaches. We introduce a performan…
arXiv:2608.04552v1 Announce Type: new Abstract: Black-box language-model reliability is commonly pursued by sampling, prompting, voting, verifying, or iteratively revising individual answers. We ask a prior question: \emph{what determines whether a collection of black-box respons…
arXiv cs.AI
TIER_1English(EN)·Shahed Masoudian, Passant Shafaei, Monorama Swain, Markus Schedl·
arXiv:2608.04714v1 Announce Type: cross Abstract: Benchmark scores are reported as properties of a model, yet the inference framework used to produce them, such as HuggingFace, vLLM, or Ollama, are considered non-influential and their names and versions are almost never disclosed…
arXiv cs.AI
TIER_1English(EN)·Enrico Mensa, Lorenzo Zane, Calogero Jerik Scozzaro, Matteo Delsanto, Tommaso Milani, Daniele Paolo Radicioni·
arXiv:2608.04670v1 Announce Type: cross Abstract: Large Language Models (LLMs) have transformed computational linguistics and achieved remarkable performance across numerous natural language processing tasks, yet significant gaps persist in understanding how these systems process…
As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automated harness optimization -- the iterative…
Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon: SFT suffers from severe task conflicts under multi-stage t…
arXiv cs.AI
TIER_1English(EN)·Farhan Ahmed, Yuya Jeremy Ong, Chad DeLuca·
arXiv:2603.24929v2 Announce Type: replace Abstract: Understanding and quantifying uncertainty in large language model (LLM) outputs is critical for reliable deployment. However, traditional evaluation approaches provide limited insight into model confidence at individual token po…
arXiv cs.AI
TIER_1English(EN)·Mobina Kashaniyan, Ali Jannesari·
arXiv:2608.03961v1 Announce Type: new Abstract: Test-time scaling improves LLM reasoning by generating and aggregating multiple candidate answers, yet many pipelines use fixed per-query budgets that spend the same compute on easy and difficult prompts. These fixed budgets are als…
arXiv cs.AI
TIER_1English(EN)·Zhinan Liu, Jie Li, Mingyu Kang, Jiayi Ji·
arXiv:2608.03838v1 Announce Type: new Abstract: Reasoning-based guard models improve LLM safeguards, but decoding explicit rationales for every interaction makes them costly to deploy. Although latent-reasoning methods reduce token generation by moving reasoning into continuous s…
arXiv:2608.03501v1 Announce Type: new Abstract: AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research process, prior work has focused primarily on code implementation and execution, o…
arXiv:2608.03219v1 Announce Type: new Abstract: Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior. A model may reach new answers, or produce answers that were already within reach. Aggregate …
arXiv:2608.03071v1 Announce Type: new Abstract: Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting the right tool and orchestrating the order of calls. However, correctly filling the parameters…
arXiv:2608.03887v1 Announce Type: new Abstract: Fine-tuning a large language model on new data degrades what it previously learned. We present Omega-S, a drop-in penalty computed from the weight matrix alone: it needs no previous-task data, no Fisher matrix and no stored copy of …
arXiv cs.LG
TIER_1English(EN)·Anton Rasmussen, Hong Qin·
arXiv:2608.03854v1 Announce Type: new Abstract: When decoder language models are used as classifiers, predicted class probabilities depend on implementation choices, including the prompt template, verbalizer (label-to-token mapping), and scoring rule, that are rarely treated as e…
arXiv cs.LG
TIER_1English(EN)·Ameen Patel, Max Zhang, Nathan Hu·
arXiv:2608.02632v1 Announce Type: new Abstract: Circuit tracing is an exciting technique for revealing the internal computation of language models, but it requires a time-intensive manual step of grouping individual features or MLP neurons into supernodes. We present a simple pip…
arXiv:2608.04010v1 Announce Type: cross Abstract: Existing scaling strategies for Multimodal Large Language Models (MLLMs) typically expand either model parameters or sequential inference computation, incurring substantial memory or latency overhead. More importantly, most existi…
arXiv cs.CL
TIER_1English(EN)·Kejian Zhu, Zhuoran Jin, Shangqing Tu, Hongbang Yuan, Yushi Bai, Kang Liu, Juanzi Li, Jun Zhao·
arXiv:2608.03573v1 Announce Type: new Abstract: Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon: SFT suffer…
arXiv:2608.03063v1 Announce Type: new Abstract: Merchant risk control at large payment platforms screens tens of millions of merchants daily, where false positives harm legitimate merchants and false negatives leave harmful activity undetected. The hardest cases require jointly u…
arXiv cs.CL
TIER_1English(EN)·Xiao Fei, Yang Zhang, Sarah Almeida Carneiro, Michalis Vazirgiannis·
arXiv:2608.02966v1 Announce Type: new Abstract: Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scoring treats all incorrect responses alike, even though an LLM's preferences among i…
arXiv:2608.02612v1 Announce Type: new Abstract: Formulating an optimization problem strongly affects the quality of the final solution, yet good formulations usually require substantial expertise. Recent studies have therefore examined how to automatically derive optimization pro…
arXiv cs.AI
TIER_1English(EN)·Bakbergen Ryskulov, Iker Garc\'ia-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Rom\'an Or\'us·
arXiv:2608.03796v1 Announce Type: cross Abstract: Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained from scratch: a compressed model is usually recovered through knowledge distillatio…
arXiv:2608.03794v1 Announce Type: cross Abstract: Large Language Models (LLMs) are transforming database interaction paradigms, evolving from simple query translators to autonomous database administrators (DBAs). However, current evaluation benchmarks remain disproportionately fi…
arXiv cs.AI
TIER_1English(EN)·Forough Majidi, Mohammad Mehdi Morovati, Foutse Khomh, Heng Li·
arXiv:2608.03036v1 Announce Type: cross Abstract: Large Language Models (LLMs) are integrated into software systems and AI services, making efficient LLM serving a concern for software engineering. Serving LLMs is challenging because inference requires computation, memory, GPU re…
arXiv cs.AI
TIER_1English(EN)·Yongwan Jo, Jinyoung Park, Euihyun Lee, Dokyung Song·
arXiv:2608.02995v1 Announce Type: cross Abstract: Modern large language models (LLMs) exhibit activation sparsity, wherein only a subset of their neurons is activated for given input tokens. Researchers have leveraged this property to optimize LLM serving systems by omitting weig…
arXiv cs.AI
TIER_1English(EN)·Soumadeep Saha, Krish Sharma, Akshay Chaturvedi, Nicholas Asher·
arXiv:2608.02867v1 Announce Type: cross Abstract: Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large language models (LLMs) across a variety of reasoning tasks, there is significant debate as to whether RLVR expands the reasoning …
<p>I released <a href="https://llm.datasette.io/en/stable/changelog.html#v0-32">LLM 0.32</a> this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider tools, r…
Existing scaling strategies for Multimodal Large Language Models (MLLMs) typically expand either model parameters or sequential inference computation, incurring substantial memory or latency overhead. More importantly, most existing methods fail to alter the rigid, fixed computat…
Fine-tuning a large language model on new data degrades what it previously learned. We present Omega-S, a drop-in penalty computed from the weight matrix alone: it needs no previous-task data, no Fisher matrix and no stored copy of the old weights. It is three lines in an existin…
AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research process, prior work has focused primarily on code implementation and execution, overlooking the importance of this stage, and no …
Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior. A model may reach new answers, or produce answers that were already within reach. Aggregate scores do not distinguish these changes question…
arXiv:2510.27313v3 Announce Type: replace-cross Abstract: Generation novelty is a key indicator of an LLM's ability to generalize, yet measuring it against full pretraining corpora is computationally challenging. Existing evaluations often rely on lexical overlap, failing to dete…
arXiv cs.CL
TIER_1English(EN)·Yutong Ke, Ming Yin, Chongwen Zhao, Kaizhu Huang·
arXiv:2608.00218v1 Announce Type: new Abstract: Agentic LLMs exhibit three consequential tool-use failures: invalid arguments (validity), unnecessary calls (over-calling), and omitted calls when tools are needed (missing). We find that a small, failure-specific set of MLP neurons…
arXiv:2608.01810v1 Announce Type: new Abstract: Rubric-based LLM-as-judge pipelines often assume that evaluation criteria provide independent signals. In practice, however, criteria can be behaviorally coupled: improving one criterion may systematically change scores on another, …
arXiv:2608.02415v1 Announce Type: new Abstract: Intent classification in Large Language Models (LLMs) involves categorizing user prompts into predefined classes. For instance, given a user prompt, the system must determine whether it primarily concerns mathematics, coding, or gen…
arXiv:2608.02515v1 Announce Type: new Abstract: Long-running assistants and agents consume interaction streams that eventually outgrow the context. Existing context retention, summarization, and retrieval preserve access to selected history, but do not provide a persistent state …
arXiv cs.CL
TIER_1English(EN)·Shrenil Shaun Sharma, Avi Sharma·
arXiv:2608.00991v1 Announce Type: cross Abstract: This paper introduces SCHEDBench, a natural-language benchmark for evaluating combinatorial scheduling constraint faithfulness under surface-form variation. Grounded in canonical scheduling instances and solver-derived feasibility…
arXiv:2608.01050v1 Announce Type: cross Abstract: Production LLM agents that select from large skill libraries face a limitation that semantic relevance alone cannot resolve: a skill may match a user's topic yet be impossible to execute in the current account state. We present a …
arXiv:2608.01975v1 Announce Type: cross Abstract: Large language model (LLM) inference has evolved from an offline workload into a continuously operated software service, yet root-cause analysis remains difficult because a single request spans the inference engine, Python/C++ bac…
arXiv:2602.01132v2 Announce Type: replace Abstract: Tasks such as solving arithmetic equations, evaluating truth tables, and completing syllogisms are handled well by large language models (LLMs) in their standard form, but they often fail when the same problems are posed in logi…
arXiv cs.CL
TIER_1English(EN)·Xiangming Gu, Soham De, Michalis Titsias, Larisa Markeeva, Petar Veli\v{c}kovi\'c, Razvan Pascanu·
arXiv:2604.06543v2 Announce Type: replace Abstract: In this work, we demonstrate that reliable stochastic sampling is a fundamental yet unfulfilled requirement for Large Language Models (LLMs) operating as agents. Agentic systems are frequently required to sample from distributio…
arXiv cs.LG
TIER_1English(EN)·Zhichao Xu, Xueguang Ma, Shengyao Zhuang, Luyu Gao, Wenqian Ye, Yu Wang, Jamie Callan, Jimmy Lin·
arXiv:2608.00916v1 Announce Type: cross Abstract: Modern reranking recipes---billion-scale cross-encoders, mixture-of-experts (MoE) backbones, and distillation against strong teachers---have outpaced the training infrastructure available to most academic groups. Existing Tevatron…
arXiv:2608.02031v1 Announce Type: cross Abstract: This paper investigates collaborative mobile edge computing (MEC) servers for large language model (LLM) inference under soft deadline constraints. In this system, to improve the quality of service, computations are expected to be…
arXiv:2511.00382v2 Announce Type: replace-cross Abstract: Organizations increasingly adapt Large Language Models (LLMs) from public repositories such as HuggingFace to downstream tasks. Prior work shows that even fine-tuning on benign datasets can weaken safety alignment, raising…
Large Language Models (LLMs) are integrated into software systems and AI services, making efficient LLM serving a concern for software engineering. Serving LLMs is challenging because inference requires computation, memory, GPU resources, and execution while maintaining latency a…
Omega-S is a lightweight, data-free regularization penalty for low-rank fine-tuning that improves retention of original model capabilities by penalizing variance in weight-matrix node degrees.
arXiv:2607.28966v1 Announce Type: new Abstract: Large language models often improve task performance by generating long reasoning traces, but the resulting computation is frequently wasted on redundant verification and revision. Existing probe-based early-exit approaches mainly i…
arXiv:2607.29433v1 Announce Type: new Abstract: As large language model (LLM) agents evolve into personalized companions, memory has emerged as a core capability. However, LLMs face a knowledge utilization problem: they may fail to act on relevant user preferences even when they …
arXiv:2607.28657v1 Announce Type: new Abstract: Large Language Models (LLMs) often require carefully crafted prompts to unlock their full potential, which can be a barrier for non-expert users. This work addresses the challenge by introducing a Task-Aware Prompt Rewriter (TAPR), …
arXiv:2607.29431v1 Announce Type: new Abstract: Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-succes…
arXiv cs.AI
TIER_1English(EN)·Philipp D. Siedler, Jordan Sassoon·
arXiv:2607.28801v1 Announce Type: cross Abstract: Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial variation in the demands of individual samples. We introduce a dataset-centric …
arXiv cs.AI
TIER_1English(EN)·Martin Lukk (University of Toronto)·
arXiv:2607.28934v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent …
arXiv cs.LG
TIER_1English(EN)·Jiaxuan Chen, Jianshu She, Ye Yuan, Rajat Ghosh, Karan Gupta, Qirong Ho, Xue Liu, Oana Balmau·
arXiv:2607.28848v1 Announce Type: cross Abstract: LLM serving systems are provisioned for peak load to meet strict latency targets, leaving substantial GPU compute idle whenever traffic falls below peak. We present DeltaServe, a host-agnostic co-serving design that converts this …
arXiv cs.LG
TIER_1English(EN)·Jiajia Tang, Sizhe Yuen, Francisco Gomez Medina, Yali Du, Adam Sobey·
arXiv:2607.29601v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) commonly adapts large language models using a single shared Low-Rank Adapter (LoRA). This shared optimization space often suffers from interference when adapting heterogeneous task sequences, l…
arXiv:2607.28640v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy in model predictions under such cross-modal variation…
arXiv:2603.09161v2 Announce Type: replace-cross Abstract: Learning effective netlist representations is fundamentally constrained by the scarcity of labeled datasets, as real designs are protected by Intellectual Property (IP) and costly to annotate. Existing work therefore focus…
Modern reranking recipes---billion-scale cross-encoders, mixture-of-experts (MoE) backbones, and distillation against strong teachers---have outpaced the training infrastructure available to most academic groups. Existing Tevatron reranker training relies on the Hugging Face Trai…
arXiv:2508.02092v3 Announce Type: replace-cross Abstract: Large language models represent significant investments in computation, data, and engineering expertise, making them extraordinarily valuable intellectual assets. Nevertheless, these AI assets remain vulnerable to unauthor…
arXiv:2607.25018v2 Announce Type: replace Abstract: Large language model (LLM) cascades reduce inference cost by routing easy queries to a small model and deferring hard queries to a larger one. Production cascades govern this deferral through a confidence threshold, but LLM conf…
arXiv cs.LG
TIER_1English(EN)·Dikshant Kukreja, Kritarth Prasad, Avinash Anand, Zhengkui Wang, Erik Cambria, Timothy Liu, Aik Beng Ng, Simon See, Bapi Chatterjee·
arXiv:2606.22932v2 Announce Type: replace Abstract: Reverse-mode differentiation computes every weight gradient, writes it to memory, and only then lets the optimizer read it back. This two-phase schedule sets the memory ceiling of modern training: at the seam between the phases,…
arXiv cs.LG
TIER_1English(EN)·Leandro Giusti Mugnaini, Bruno Lopes Yamamoto, Lucas Lauton de Alcantara, Victor Zacarias, Edson Bollis, Lucas Pellicer, Anna Helena Reali Costa, Artur Jordao·
arXiv:2504.21174v2 Announce Type: replace Abstract: Deep learning drives a new wave in computing systems and triggers the automation of increasingly complex problems. In particular, Large Language Models (LLMs) have significantly advanced cognitive tasks, often matching or even s…
arXiv cs.LG
TIER_1English(EN)·Jinliang Gao, Ning Yang, Hai Wang, Baili Xiao, Pin Lyu·
arXiv:2607.27273v1 Announce Type: new Abstract: Post-training of large language models is expensive, and existing efficiency improvements mainly focus on selecting informative samples or designing training schedules. However, data organization itself is usually treated as a stati…
arXiv cs.CL
TIER_1English(EN)·Lukas Hinterleitner, Loris Schoenegger, Benjamin Roth·
arXiv:2601.16651v3 Announce Type: replace Abstract: Gradient-based methods for instance-based explanation for large language models (LLMs) are hindered by the immense dimensionality of model gradients. In practice, influence estimation is restricted to a subset of model parameter…
arXiv:2607.28418v1 Announce Type: cross Abstract: Pruning is a promising approach for improving the efficiency of LLMs. Existing static structured pruning methods are hardware-friendly and can deliver practical throughput gains, but their input-agnostic computation allocation oft…
arXiv cs.CL
TIER_1English(EN)·Nayera Hasan, Jack Greff, Alvin Grissom II·
arXiv:2607.27405v1 Announce Type: new Abstract: Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes doma…
arXiv:2607.27379v1 Announce Type: new Abstract: High-quality, diverse data are vital for large language models (LLMs) but remain scarce and costly. Data synthesis is a viable alternative and succeeds on closed tasks, yet the humanities and social sciences (HSS) are overlooked, an…
arXiv:2507.02259v2 Announce Type: replace Abstract: Despite improvements by length extrapolation, efficient attention and memory modules, handling infinitely long documents with linear complexity without performance degradation during extrapolation remains the ultimate challenge …
arXiv cs.CL
TIER_1English(EN)·Ekaterina Trofimova, Zosia Shamina, Maria Selifanova, Artem Zaitsev, Remi Savchuk, Maxim Minets, Daria Ozerova, Emil Sataev, Denis Zuenko, Andrey E. Ustyuzhanin·
arXiv:2509.22768v3 Announce Type: replace Abstract: We introduce ML2B, the first benchmark for evaluating cross-lingual task comprehension in end-to-end ML pipeline generation by large language models. Despite growing global AI adoption, no systematic evaluation exists for ML pip…
arXiv:2607.26457v1 Announce Type: new Abstract: Reinforcement learning is a natural post-training paradigm for code-oriented large language models because generated programs can be evaluated through parsing, execution, unit tests, and structural analysis.However, existing methods…
arXiv:2607.26828v1 Announce Type: new Abstract: Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive discovery controllers assign credit based only on score progress, even though…
Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive discovery controllers assign credit based only on score progress, even though prompt length, retries, and guidance calls caus…
arXiv:2607.24780v1 Announce Type: new Abstract: Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specific failure modes -- while human preference is subjec…
arXiv:2607.25554v1 Announce Type: new Abstract: Future event prediction carries broad social impact yet remains challenging. SOTA approaches augment LLMs with external agent frameworks whose predictive capability vanishes once the harness is removed. While recent Tool-Integrated …
arXiv cs.AI
TIER_1English(EN)·Chandan Kumar Sah, Xiaoli Lian, Li Zhang·
arXiv:2607.25880v1 Announce Type: cross Abstract: LLM provenance testing asks whether a suspect LLM belongs to the same lineage as a source. Existing black-box methods largely infer this relationship from response-level characteristics, but these characteristics may shift under a…
arXiv:2607.24769v1 Announce Type: new Abstract: With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-context scheming -- the covert pursuit of misaligned ob…
Retrieval-augmented generation improves knowledge-intensive question answering, but indiscriminate retrieval can introduce irrelevant evidence and unnecessary computation. We investigate whether verbalized confidence from black-box language models can serve as an actionable signa…
Retrieval-augmented generation improves knowledge-intensive question answering, but indiscriminate retrieval can introduce irrelevant evidence and unnecessary computation. We investigate whether verbalized confidence from black-box language models can serve as an actionable signa…
arXiv:2607.24539v1 Announce Type: new Abstract: Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish that task-appropriate evidence was used. This letter proposes a general framewo…
arXiv cs.AI
TIER_1English(EN)·Yijia Dai, Zhaolin Gao, Yahya Sattar, Jennifer J. Sun, Sarah Dean·
arXiv:2607.22646v1 Announce Type: new Abstract: Large language models (LLMs) display a striking ability to predict next observations from Hidden Markov Models (HMMs) via in-context learning (ICL), but the algorithm underlying this capability remains undetermined: prior work has p…
arXiv:2607.22625v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge, but suffers from knowledge conflicts: when retrieved information contradicts parametric memory, the shared self-attention pathway pr…
arXiv cs.AI
TIER_1English(EN)·Samyak Jhaveri, Erel Kaplan, Tom Yotam, Le Chen, Tomer Bitan, Niranjan Hasabnis, Gal Oren·
arXiv:2607.22588v1 Announce Type: new Abstract: Modern compute-intensive software must migrate across a changing ecosystem of accelerators, programming APIs, compiler stacks, and portability layers, including CUDA, OpenMP, OpenCL, and OpenMP target offload. Large language models …
arXiv:2607.22583v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved widespread adoption because of their strong reasoning and query-response capabilities. However, deploying them in embedded and edge computing environments remains challenging because of str…
arXiv:2607.22578v1 Announce Type: new Abstract: The proliferation of Large Language Models (LLMs) has shifted serving systems from processing isolated requests to orchestrating high-concurrency, multi-tenant agentic workflows. However, existing solutions typically prioritize intr…
arXiv:2607.22554v1 Announce Type: new Abstract: Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study how m…
arXiv cs.CL
TIER_1English(EN)·Mengqi Wang (UNSW Sydney), Jianwei Wang (UNSW Sydney), Qing Liu (Data61, CSIRO), Xiwei Xu (Data61, CSIRO), Zhenchang Xing (Data61, CSIRO), Liming Zhu (Data61, CSIRO), Michael Bain (UNSW Sydney), Wenjie Zhang (UNSW Sydney)·
arXiv:2512.07246v2 Announce Type: replace Abstract: Error detection (ED), which aims to identify incorrect or inconsistent cell values in tabular data, is important for ensuring data quality. Recent state-of-the-art ED methods leverage the pre-trained knowledge and semantic capab…
arXiv:2602.05547v2 Announce Type: replace-cross Abstract: RL-based post-training with GRPO is widely used to improve large language models on individual reasoning tasks. However, real-world deployment requires reliable performance across diverse tasks. A straightforward multi-tas…
arXiv:2607.22898v1 Announce Type: cross Abstract: Large language models (LLMs) generate code from natural-language prompts, yet real-world prompts rarely provide complete specifications. When prompts leave input formats, error handling, or design decisions unspecified, LLMs fill …
arXiv:2607.22759v1 Announce Type: cross Abstract: Large language models (LLMs) show promise in code generation, but their capabilities to produce correct, synthesizable hardware description language (HDL) code still remain to be properly benchmarked. Existing evaluations are prim…
arXiv:2607.24260v1 Announce Type: new Abstract: Modern LLM systems increasingly rely on knowledge-selection processes that produce high-value structured priors, such as ranked evidence, graph topology, multimodal alignment, and confidence signals. Yet LLM serving remains fundamen…
arXiv:2607.21927v1 Announce Type: new Abstract: Full self-attention in large language models scales as O(N^2), which limits long-context document analysis to 65,536 tokens and requires costly GPU clusters. The Reduced Interaction Sampling (RIS) inference engine addresses this con…
arXiv:2601.03190v4 Announce Type: replace Abstract: Machine unlearning aims to forget sensitive knowledge from Large Language Models (LLMs) while maintaining general utility. However, existing approaches typically treat all tokens in a response indiscriminately and enforce uncert…
arXiv cs.LG
TIER_1English(EN)·Jinhyeok Kim, Yejoon Lee, Jaeyoung Do·
arXiv:2607.21985v1 Announce Type: cross Abstract: The increasing deployment of large language models (LLMs) has magnified the computational and memory bottlenecks of autoregressive decoding, where low compute intensity and bandwidth-bound kernels dominate inference cost. Weight p…
arXiv cs.AI
TIER_1English(EN)·Zishan Shao, Lixun Zhang, Kangning Cui, Yixiao Wang, Ting Jiang, Hancheng Ye, Qinsi Wang, Zhixu Du, Yuzhe Fu, Fan Yang, Danyang Zhuo, Yiran Chen, Hai Helen Li·
arXiv:2607.20469v1 Announce Type: new Abstract: Large language models (LLMs) handle many tasks with one set of parameters, but under KV-cached inference it is unclear what task-general structure, if any, is used at decode time rather than during prefill. We propose DecodeShare, a…
arXiv:2607.20515v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is critical for aligning Large Language Models (LLMs) with human preferences. However, its efficacy is often compromised by the inherent inconsistency and subjectivity of human annot…
arXiv cs.AI
TIER_1English(EN)·Yichuan Liu, Daniel Cummings, Nick Vadlamudi·
arXiv:2607.20638v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated strong capabilities in code generation and reasoning, yet their ability to perform temporal reasoning over digital waveform data remains largely unexplored. Although reasoning over digi…
arXiv cs.AI
TIER_1English(EN)·Yangjun Lu, Hongyi Zhou, Fabian Spill, Kai Ye, Chengchun Shi, Jin Zhu·
arXiv:2607.21458v1 Announce Type: new Abstract: The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents. Existing methods for detecting LLM-generate…
arXiv cs.AI
TIER_1English(EN)·Mikail Khona, Aditya Vavre, Boxiang Wang, Deyu Fu, Hao Wu, Mike Chrzanowski, Bryan Catanzaro, Dheevatsa Mudigere, Jeff Pool, Michael Lightstone, Mohammad Shoeybi, Mostofa Patwary, Nima Tajbakhsh, Tijmen Blankevoort·
arXiv:2607.20548v1 Announce Type: cross Abstract: Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges have limited adoption at scale. In this work, we adapt and enhance preconditioned g…
arXiv cs.AI
TIER_1English(EN)·Emma Kondrup, Zachary Yang, Anne Imouza, Reihaneh Rabbany·
arXiv:2607.20558v1 Announce Type: cross Abstract: AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government services. Yet their real-world behavior remains poorly understood due to strong context dependence. Current evaluation protocols fol…
arXiv cs.AI
TIER_1English(EN)·Minh Ngoc Ta, My Anh Tran Nguyen, Duong D. Nguyen, Yuxia Wang, Preslav Nakov·
arXiv:2607.21143v1 Announce Type: cross Abstract: Ambiguous user requests make clarification a sequential decision problem for conversational LLM assistants: they must decide whether to ask, what to ask, when to stop, and when to answer. We introduce RegretBench, a multi-turn ben…
arXiv cs.AI
TIER_1English(EN)·Lorenzo Orsingher, Thomas De Min, Massimiliano Mancini, Davide Talon, Elisa Ricci·
arXiv:2607.21300v1 Announce Type: cross Abstract: Machine unlearning has emerged as a tool for removing personal data from trained models to comply with recent AI regulations. To evaluate unlearning effectiveness in multimodal large language models (MLLMs), prior works fine-tune …
arXiv cs.AI
TIER_1English(EN)·Heejun Kim, Seungpil Lee, Jewon Yeom, Jaewon Sok, Seonghyeon Park, Jeongjae Park, Taesup Kim, Sundong Kim·
arXiv:2605.11936v2 Announce Type: replace Abstract: Recent soft prompt research has tried to improve reasoning by inserting trained vectors into LLM inputs, yet whether the gain comes from the learned content or from the act of injection itself has not been carefully separated. W…
arXiv cs.CL
TIER_1English(EN)·Juho Leinonen, Paul Denny·
arXiv:2607.20446v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly used by students to generate natural language responses and program code, there is growing interest in whether LLMs themselves can be used to distinguish AI-generated work from human-…
arXiv:2607.21063v1 Announce Type: new Abstract: Almost every large language model that reaches a broad audience is quantized: trained in full precision, then compressed for efficiency. This step is assumed harmless and its safety is rarely re-checked. We find its principal side e…
arXiv:2607.20465v1 Announce Type: cross Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows actually prepare training data en…
arXiv:2607.20864v1 Announce Type: cross Abstract: Position bias in multiple-choice LLM evaluation is widely cited as a confound in capability comparisons, but published measurements rely on single answer-order shuffles whose results confound the bias signal with content-level noi…
arXiv cs.LG
TIER_1English(EN)·Amir Sarfi, Benjamin Th\'erien, Joel Lidin, Eugene Belilovsky·
arXiv:2508.15706v3 Announce Type: replace Abstract: Communication-efficient distributed training algorithms (e.g., DiLoCo) have received considerable interest due to their benefits for training large language models (LLMs) in bandwidth-constrained settings, such as across datacen…
arXiv:2607.20470v1 Announce Type: new Abstract: Enhancing the task-specific capabilities of Large Language Models (LLMs) primarily requires substantial instruction-tuning datasets. However, the sheer volume of such data imposes a considerable annotation cost, and a lack of optimi…
The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents. Existing methods for detecting LLM-generated text mainly focus on document-level classifica…
Almost every large language model that reaches a broad audience is quantized: trained in full precision, then compressed for efficiency. This step is assumed harmless and its safety is rarely re-checked. We find its principal side effect is increased bias that standard safety eva…
arXiv:2604.03447v2 Announce Type: replace-cross Abstract: LLM-based software engineering assistants often reason over multiple artifacts, including code, documentation, signatures, and tests, even when those artifacts are incomplete or mutually inconsistent. Existing evaluations …
arXiv cs.AI
TIER_1English(EN)·Beining Wang, Weihang Su, Hongtao Tian, Hao Kong, Tao Yang, Ting Yao, Qingyi Pan, Yueyue Wu, Qingyao Ai, Min Zhang, Yiqun Liu·
arXiv:2607.20083v1 Announce Type: cross Abstract: Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models. As policies improve, these sampled responses become close in quality. These close candidates create a…
arXiv:2607.19824v1 Announce Type: new Abstract: LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has become a major post-training approach for this goal, but existing proxy rewards are often outcome…
arXiv:2607.19349v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as always-on online services, making efficient LLM serving a critical systems challenge. Achieving low latency and high throughput under volatile demand requires deep understand…
arXiv cs.CL
TIER_1English(EN)·Andrei Chetvergov, Stepan Ukolov, Timofei Sivoraksha, Alexander Evseev, Mikhail Solovev, Valeriia Kuschenko, Maria Chistyakova, Sergey Bolovtsov·
arXiv:2607.20270v1 Announce Type: new Abstract: Large language models are increasingly evaluated through the values they endorse, but such evaluations presuppose that models can identify the value expressed in a concrete situation. We study this prerequisite as controlled top-1 r…
Large language models are increasingly evaluated through the values they endorse, but such evaluations presuppose that models can identify the value expressed in a concrete situation. We study this prerequisite as controlled top-1 recognition over Schwartz's ten basic values. Our…
arXiv:2607.18259v1 Announce Type: new Abstract: Steering vectors (SVs), an inference-time intervention technique for large language models (LLMs), guide the generation process by adding a concept-specific direction vector to intermediate activations during inference. However, exi…
arXiv:2607.18261v1 Announce Type: new Abstract: LLM agents are increasingly used as transaction compilers: a user states an intent in natural language, and the model emits a structured object that an API can execute. JSON Schema and provider-level structured-output modes are usef…
arXiv:2607.18639v1 Announce Type: cross Abstract: Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e.g., unlearning, filtering) and suppressing it at the output layer (e.g., refusal training); both pay a tax in adjacent-domain comp…
arXiv:2607.18915v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), modern systems not only possess strong foundational capabilities and extensive knowledge, but can also solve complex problems via long, multi-step reasoning. However, as re…
arXiv cs.AI
TIER_1English(EN)·Shvetank Prakash, Andrew Cheng, Mark Mazumder, Arya Tschand, Varun Gohil, Jeffrey Ma, Jason Yik, Zishen Wan, Jessica Quaye, Elisavet Lydia Alvanaki, Avinash Kumar, Chandrashis Mazumdar, Tuhin Khare, Alexander Ingare, Ikechukwu Uchendu, Radhika Ghosal, Ab…·
arXiv:2510.22087v2 Announce Type: replace-cross Abstract: The field of computer architecture, which bridges high-level software abstractions and low-level hardware implementations, remains absent from current large language model (LLM) evaluations. To this end, we present QuArch …
arXiv:2607.18496v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM performance on security tasks is an active area of measurement and research, often with a fo…
arXiv:2607.18360v1 Announce Type: cross Abstract: Large language models (LLMs) now routinely draft literature reviews and assist with academic writing, which means a higher risk of fabricated references: GPTZero found 53 papers with hallucinated citations among NeurIPS 2025's acc…
arXiv cs.AI
TIER_1English(EN)·Meena Jagadeesan, Tatsunori Hashimoto, Jon Kleinberg·
arXiv:2607.19300v1 Announce Type: new Abstract: As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example through LLM detection tools and through heuristics based on language patterns. Detectors operate as an intervention…
arXiv:2607.18098v1 Announce Type: new Abstract: Large language models are increasingly used in practical systems, making efficient model selection important for reducing deployment cost. LLM routing has emerged as a practical solution for allocating each input query to an appropr…
arXiv cs.LG
TIER_1English(EN)·Qi Huang, Furong Ye, Ananta Shahane, Thomas B\"ack, Niki van Stein·
arXiv:2603.02792v2 Announce Type: replace Abstract: Large Language Models (LLMs) have already been widely adopted for automated algorithm design, demonstrating strong abilities in generating and evolving algorithms across various fields. Existing work has largely focused on exami…
arXiv cs.CL
TIER_1English(EN)·Roberto Pietrantuono, Antonio Guerriero, Pouya Sattari·
arXiv:2607.16808v1 Announce Type: new Abstract: Event Argument Extraction (EAE) converts documents into structured event records by identifying argument spans and assigning them schema-defined roles. Document-level EAE is challenging due to long-range dependencies between trigger…
arXiv cs.AI
TIER_1English(EN)·Hippolyte Gisserot-Boukhlef, Nicolas Boizard, Emmanuel Malherbe, C\'eline Hudelot, Pierre Colombo·
arXiv:2604.09497v2 Announce Type: replace-cross Abstract: Accurate evaluation is central to the large language model (LLM) ecosystem, guiding model selection and downstream adoption across diverse use cases. In practice, however, evaluating generative outputs typically relies on …
arXiv:2505.20161v2 Announce Type: replace-cross Abstract: Effective generalization in language models depends critically on the diversity of their training data. Yet existing diversity metrics often fall short of this goal, relying on surface-level heuristics that are decoupled f…
arXiv cs.AI
TIER_1English(EN)·Yongkang Du, Xiaohan Zou, Minhao Cheng, Lu Lin·
arXiv:2603.27958v2 Announce Type: replace Abstract: Analogical reasoning tests a fundamental aspect of human cognition: mapping the relation from one pair of objects to another. Existing evaluations of this ability in multimodal large language models (MLLMs) overlook the ability …
arXiv:2607.17047v1 Announce Type: cross Abstract: LLM constraint reasoners are often evaluated near the random-SAT phase transition, confounding density and solver hardness. We test instance-level transfer while near-matching clause density. At aligned size bins, with near-matche…
arXiv cs.AI
TIER_1English(EN)·Jiahe Fan, Yinghao Hou, Si Chen, Aiyuan Zhang, Hong Xie, Defu Lian·
arXiv:2607.18026v1 Announce Type: new Abstract: Can large language models with substantially different parameter spaces be merged by direct weighted averaging, without training or semantic alignment? Existing heterogeneous fusion methods typically introduce distillation, adapters…
arXiv:2607.17063v1 Announce Type: new Abstract: The rapid advancement of large language models (LLMs) has led practitioners to increasingly rely on them for answering questions about hardware description languages (HDLs). Because HDL is ultimately synthesized into physical hardwa…
Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e.g., unlearning, filtering) and suppressing it at the output layer (e.g., refusal training); both pay a tax in adjacent-domain competence or over-refusal. We argue that the right op…
arXiv:2607.16122v1 Announce Type: new Abstract: Evaluations should do more than measure a models current performance. They should tell us what to fix for the next model iteration and provide a way to generate targeted post training data. Most evaluation pipelines identify weak ex…
arXiv:2607.15957v1 Announce Type: new Abstract: Large Language Models (LLMs) can generate natural language explanations that rationalize their own decisions, a phenomenon commonly referred to as self-explanations.Such explanations have emerged as a promising direction for explain…
Ahead of AI (Sebastian Raschka)
TIER_1English(EN)·Sebastian Raschka, PhD·
How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes
arXiv cs.AI
TIER_1English(EN)·Jimmy T. H. Smith, Tarek Dakhran, Alberto Cabrera, Simon S. Lee, Paul Pak, Aditya Tadimeti, Tim Seyde, Maxime Labonne, Alexander Amini, Mathias Lechner·
arXiv:2607.15232v1 Announce Type: cross Abstract: A tokenizer fixed at the start of pre-training allocates vocabulary in proportion to the pre-training corpus, reflecting the deployment priorities at that time. When those priorities shift, languages added later are split into man…
arXiv:2602.12429v2 Announce Type: replace Abstract: Foundation models have achieved remarkable success, yet their growing parameter counts pose significant computational and memory challenges. Low-rank factorization offers a promising route to reduce training and inference costs,…
arXiv:2607.14618v1 Announce Type: new Abstract: CPUs are the most universal target for on-device LLM inference, but existing low-bit quantization methods offer either coarse operating points or fine-grained mixed precision that is difficult to execute efficiently on CPUs. We pres…
arXiv:2607.14306v1 Announce Type: new Abstract: In this paper, we study the connection between an LLM's output distribution and the data used to train it. Specifically, we study the degree to which an LLM's next-token distribution agrees with the empirical next-token distribution…
arXiv:2607.14118v1 Announce Type: new Abstract: Large language models (LLMs) can generate research ideas that appear novel to expert reviewers, but recent work also shows that such ideas often lack diversity, are difficult for LLMs to evaluate reliably, and may fail to translate …
arXiv cs.AI
TIER_1English(EN)·Jiayi Zhang, Simon Yu, Derek Chong, Anthony Sicilia, Michael R. Tomz, Christopher D. Manning, Weiyan Shi·
arXiv:2510.01171v4 Announce Type: replace-cross Abstract: Post-training alignment often reduces LLM diversity, leading to a phenomenon known as mode collapse. Unlike prior work that attributes this effect to algorithmic limitations, we identify a fundamental, pervasive data-level…
arXiv:2607.14109v1 Announce Type: cross Abstract: Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central challenges in natural language understanding. Furthermore, the rapid proliferation…
arXiv:2607.14111v1 Announce Type: cross Abstract: Can small language models detect and report on perturbations their own internal activations? We investigate this question through the lens of activation steering: injecting concept vectors into a model's residual stream and measur…
A tokenizer fixed at the start of pre-training allocates vocabulary in proportion to the pre-training corpus, reflecting the deployment priorities at that time. When those priorities shift, languages added later are split into many more tokens per word, which can raise latency, c…
CPUs are the most universal target for on-device LLM inference, but existing low-bit quantization methods offer either coarse operating points or fine-grained mixed precision that is difficult to execute efficiently on CPUs. We present PolyQ, a CPU-oriented compiler/quantization …
arXiv:2607.13425v1 Announce Type: cross Abstract: Learning effectively from limited data is critical in domains like security where labeled examples are scarce. Large language models (LLMs) have demonstrated some capabilities for data-efficient learning, especially through parame…
arXiv:2607.13069v1 Announce Type: new Abstract: Large language models produce chain-of-thought (CoT) reasoning that appears logically sound yet may not genuinely depend on its stated premises. We introduce interventional grounding audits, a black-box, step-level test of premise d…
arXiv:2607.13205v1 Announce Type: cross Abstract: Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking tokens on accumulated attention mass, treated here as signal energy, and keeping the heaviest. …
arXiv:2601.20379v2 Announce Type: replace Abstract: Large language models (LLMs) struggle with complex, long-horizon reasoning due to instability caused by their frozen policy assumption. Current test-time scaling methods treat execution feedback merely as an external signal for …
arXiv:2510.26707v2 Announce Type: replace-cross Abstract: As LLMs occupy an increasingly important role in society, they are more and more confronted with questions that require them not only to draw on their general knowledge but also to align with certain human value systems. T…
Learning effectively from limited data is critical in domains like security where labeled examples are scarce. Large language models (LLMs) have demonstrated some capabilities for data-efficient learning, especially through parameter-efficient adaptation methods, but continue to …
Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking tokens on accumulated attention mass, treated here as signal energy, and keeping the heaviest. On schema-dense input streams such as nested JSON,…
As LLM technology advances, the space of model families, compute hardware, quantization schemes, parallelization strategies, and specialized optimization kernels continues to expand, sharply increasing the code complexity and maintenance cost of general-purpose inference framewor…
arXiv:2607.11207v1 Announce Type: cross Abstract: Table-based reasoning with large language models (LLMs), which requires reasoning based on natural language questions and structured tabular data, has gained widespread attention. However, a series of issues still constrain the ap…
arXiv:2607.09693v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become the dominant paradigm for improving the reasoning capabilities of large language models, but it requires expensive training, curated data, and reward signals. Recent work shows that sampling …
arXiv:2607.10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and tr…
arXiv:2607.11505v1 Announce Type: cross Abstract: Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution align…
arXiv:2607.11614v1 Announce Type: cross Abstract: Extending the context length of large language models (LLMs) is critical for many real-world applications, yet standard transformers remain constrained by quadratic compute and linear memory scaling. In this work, we investigate t…
arXiv cs.AI
TIER_1English(EN)·Arjun Ashok, Andrew Robert Williams, Vincent Zhihao Zheng, Irina Rish, Nicolas Chapados, \'Etienne Marcotte, Valentina Zantedeschi, Alexandre Drouin·
arXiv:2508.09904v3 Announce Type: replace-cross Abstract: Real-world forecasting requires models to integrate not only historical data but also relevant contextual information provided in textual form. While large language models (LLMs) show promise for context-aided forecasting,…
arXiv cs.CL
TIER_1English(EN)·Renuka Oladri, Mohan Vamsi Varadaraju Priya, Jerry Wu·
arXiv:2607.09999v1 Announce Type: new Abstract: We show that post-training quantization can silently alter how large language models reason even when task accuracy is preserved. Using a six-category failure taxonomy validated by two independent human annotators (Cohen's $\kappa$ …
arXiv cs.LG
TIER_1English(EN)·Liangqi Yuan, Dong-Jun Han, Shiqiang Wang, Christopher G. Brinton·
arXiv:2502.11007v5 Announce Type: replace Abstract: Compared to traditional machine learning models, recent large language models (LLMs) can exhibit multi-task-solving capabilities through multi-modal data sources and multi-turn conversations. These unique characteristics of LLMs…
arXiv cs.LG
TIER_1English(EN)·Yuchen Zhu, Wei Guo, Jaemoo Choi, Petr Molodyk, Bo Yuan, Molei Tao, Yongxin Chen·
arXiv:2510.08233v3 Announce Type: replace Abstract: Diffusion large language models (dLLMs) are promising alternatives to autoregressive large language models (AR-LLMs), as they potentially allow higher inference throughput. Reinforcement learning (RL) is crucial to enabling dLLM…
Extending the context length of large language models (LLMs) is critical for many real-world applications, yet standard transformers remain constrained by quadratic compute and linear memory scaling. In this work, we investigate the Associative Recurrent Memory Transformer (ARMT)…
Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive exploration d…
Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive exploration d…
Table-based reasoning with large language models (LLMs), which requires reasoning based on natural language questions and structured tabular data, has gained widespread attention. However, a series of issues still constrain the application of this task. The previous approaches su…
Table-based reasoning with large language models (LLMs), which requires reasoning based on natural language questions and structured tabular data, has gained widespread attention. However, a series of issues still constrain the application of this task. The previous approaches su…
Large language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate otherwise plausible responses. We study inference-time feed-forward network (FFN) intervention for improving structured outputs without retraining…
arXiv:2604.00130v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting has significantly improved the reasoning capabilities of large language models (LLMs). However, conventional CoT often relies on unstructured, flat reasoning chains that suffer from redundancy an…
arXiv cs.AI
TIER_1English(EN)·Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal·
arXiv:2507.18043v2 Announce Type: replace-cross Abstract: Inference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLMs) by modifying internal activations at test time without updating model weights. H…
arXiv:2607.08774v1 Announce Type: new Abstract: Reliability in large language model (LLM) systems is typically framed as a function of model capability. We challenge this by demonstrating that reliability is significantly influenced by \emph{inference-time control} -- the computa…
Large language models (LLMs) increasingly improve their reasoning at test time via additional computation, yet most existing works treat each problem in isolation. When problems arrive sequentially, accumulating reusable experience across them can further improve performance. Exi…
Recent LLM-based agent systems continuously accumulate context across multi-turn interactions, tool invocations, and cross-session workflows. Replaying the full history for every request quickly becomes impractical: long contexts increase prefill cost, may exceed context limits, …
arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models to abstain when uncertain improves reliability but introduces a coverage accurac…
arXiv stat.ML
TIER_1English(EN)·Zeyu Zhang, Bradly C. Stadie·
arXiv:2608.02985v1 Announce Type: cross Abstract: The standard check for contamination in LLM backtests is simple: compare scores before and after the training cutoff. We show this check is uninformative. Four flagship models fail it on questions they cannot have memorized: every…
arXiv:2607.22951v1 Announce Type: new Abstract: Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified operational profile. Conventional benchmark-based evaluation, often summarized by agg…
arXiv:2607.20357v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance across vision-language tasks. However, their high inference cost, arising from both the large number of input visual tokens and the heavy computa…
arXiv:2607.16259v1 Announce Type: cross Abstract: Pretrained models are typically ranked on multi-task leaderboards to assess their effectiveness across diverse tasks. Rank confidence intervals were recently introduced as a method to quantify the uncertainty in these rankings by …
<p><span>The classic "</span><a href="https://transformer-circuits.pub/2025/attribution-graphs/biology.html#dives-tracing" rel="noreferrer"><span>Dallas</span></a><span>" example from Anthropic focuses on an </span><b><span>internal circuit</span></b><span> in an LLM.</span></p><…
arXiv:2511.23231v2 Announce Type: replace Abstract: Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) demonstrate strong reasoning capabilities, yet their performance in English significantly outperforms that in low-resource languages, raising fairness concern…
GPU utilization can read healthy while your queue backs up, and a new replica takes minutes to warm. Here's how to pick autoscaling metrics, tune scale-up/down windows, and budget for cold starts on dedicated inference.
Learn how to stream LLM responses in a voice pipeline with tool calling and structured outputs using AssemblyAI's LLM Gateway for sub-second response times.
<p>Implement an end-to-end fine-tuning pipeline for tool-calling language models. This tutorial covers parsing trajectories, structured tool-call extraction, Qwen-compatible ChatML rendering, and efficient LoRA adaptation using PyTorch.</p> <p>The post <a href="https://www.markte…
<p>The Generative AI landscape is moving at breakneck speed. While API access to frontier models has democratized artificial intelligence, building reliable, production-grade applications requires much more than sending a single prompt to an endpoint. It requires orchestration.</…
<h2> How we cut repo-wide symbol indexing for LLM agents from 30s to 98ms </h2> <p>If your coding agent has ever stalled for tens of seconds on "what's in this repo?" — or burned hundreds of tokens re-reading a file after a failed edit — this is the story of why that happens and …
<h1>Locking Down the AI Kernel: A Network Isolation Blueprint for Self-Hosted LLMs</h1> <p>Learn to secure self-hosted AI models by binding them to localhost and using an nginx reverse proxy for TLS termination and authentication. This guide details robust network isolation patte…
Medium — fine-tuning tag
TIER_1English(EN)·Let’s Automate ️·
<h4>Why expanding context windows degrade transformer attention — and how to build temporal precision gates and cross-encoder reranking layers for production vector retrieval.</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*vxDaACE5fD6FTn1ABn_bxg.png" /></…
Towards AI
TIER_1English(EN)·Satsawat Natakarnkitkul (Net)·
<h4>KV cache, prefill, quantization, FlashInfer — all downstream of one fact: to write three quarters of a word, the model reads all seventy gigabytes of itself.</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*4CnE_tHScHIQ358S-tRSYQ.jpeg" /><figcaption><em…
<h4><em>One pip install. No API key is needed. Your chat history, stripped down to what actually matters.</em></h4><p>By Taha Azizi — AI Engineer | Data Scientist | Tech Writer</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/769/1*AyD6rtHwTFtIxQnX-lFlMQ.png" /><fi…
Medium — Claude tag
TIER_1English(EN)·zeeshan khan·
Big new release of my LLM CLI tool and Python library for talking to hundreds of different LLMs - reasoning traces, OpenAI Responses support, server-side tools, smarter logging and a whole lot more
simonwillison.net/2026/Aug/4/n...
<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*1T9lccXvmhg31LCFHwoM4A.jpeg" /><figcaption>Long-Output LLM Architecture</figcaption></figure><p>Big context windows get the attention. Big outputs create the production mess. If your model can write tens of thous…
<h1>The LLM Waterfall Pattern: Your Shield Against Rate Limits and Downtime</h1> <p>Master the LLM waterfall pattern to eliminate workflow disruptions from API rate limits. This guide diagrams the failover cascade from primary providers to local models, ensuring zero downtime for…
Medium — MLOps tag
TIER_1English(EN)·Tedi Ikonomi·
<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*p9J-Z9rIE98LJqcqhI8CmA.jpeg" /><figcaption>LLM Provider Quirks</figcaption></figure><p>Every multi-model AI app eventually reaches the same uncomfortable moment: the abstraction layer works in the demo, then brea…
<p>A distributed trace is one of the few places where a system tells you the truth about itself. It records what actually called what, in what order, and how long each hop took. It is also, for a human being at 2am, a wall of a thousand spans.</p> <p>That combination makes tracin…
<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*qMxjyqguq6oSh1Zmd4nllw.jpeg" /><figcaption>LLM Reasoning Budget</figcaption></figure><p>A reasoning model can feel brilliant on one task and painfully slow on the next. The model did not suddenly get worse. You p…
<p>I kept running into the same problem: after a few weeks away from a project, I'd forgotten why I made certain architecture decisions. Switching between Claude, ChatGPT, and Gemini meant re-explaining everything from scratch. And in long chats, the platform's compaction slowly …
<h1> The LLM Waterfall Pattern: Never Let a Rate Limit Kill Your </h1> <p>Workflow</p> <p>## The Problem</p> <p>You're in the zone. Your AI agent is cranking out code,<br /> reviewing PRs, and generating documentation. Then it happens:<br /> </p> <div class="highlight js-code-hig…
<h1>The LLM Waterfall Pattern: Never Let a Rate Limit Kill Your Workflow</h1> <p>Implementing a provider failover strategy is critical for production AI applications. Learn why the LLM waterfall pattern outperforms simple retries and circuit breakers for zero downtime AI inferenc…
Medium — MLOps tag
TIER_1English(EN)·Hitarthdesai·
<p><strong>Everyone talks about LLMs writing essays and code. Far more useful, and far less discussed, is the LLM as a <em>translator</em> — a layer that sits between messy human language and rigid, structured machine commands, and converts one into the other. That unglamorous ro…
<p>Free model quotas have a hidden enemy: synchronous calls. Every request blocks on the network. Timeouts get wasted. Retries pile up.</p> <p>Core conclusion: an asynchronous queue turns a free model quota from a demo tool into a batch engine. A SQLite-backed queue, a few Python…
dev.to — LLM tag
TIER_1Español(ES)·Silviu Technology·
<p>Cuando un flujo con LLMs empieza a usar herramientas de verdad, el primer impulso suele ser mejorar prompts, pulir mensajes o meter otro retry. Yo suelo ir por otro lado. Si el run no deja un contrato corto y congelado antes de ejecutar, el sistema parece flexible pero en real…
<p>Running a language model locally used to be a hobbyist experiment. In 2026, it's a viable engineering decision for a growing slice of real workloads — and the tooling has finally caught up. StorageReview just published a comprehensive roundup of what actually works, with one u…
dev.to — LLM tag
TIER_1English(EN)·Rahul singh Shekhawat·
<h2> Why Single-LLM Evaluators Produce 80%+ False Alarms on Generated Code — and How a Hybrid Engine Fixes It </h2> <p>Over the past two years, the AI industry converged on a single standard for evaluating generative outputs: <strong>LLM-as-a-Judge</strong>.</p> <p>The pitch was …
dev.to — LLM tag
TIER_1English(EN)·The Unmeshed Team·
<p>There is a scene in <em>Moneyball</em> where Brad Pitt sits across from a room of old scouts who want to spend big on famous players. He tells them they are asking the wrong question. The point was never to buy the best players. <em>It was to buy wins, and wins were hiding in …
dev.to — LLM tag
TIER_1English(EN)·Mobisoft Infotech·
<p>"AI po prostu wie, co odpowiedzieć" — to najczęstsze uproszczenie, jakie słyszę o dużych modelach językowych. W praktyce za każdą odpowiedzią stoi konkretny mechanizm: tokenizacja, wektory, warstwy transformera i statystyka prawdopodobieństwa. Ten artykuł rozkłada to na czynni…
<p><em>Part 4 of a series. Previously: <a href="https://dev.to/USERNAME/PART-3-SLUG">Part 3 — A Multi-Agent Setup with Hermes</a>.</em></p> <p>Part 3 ended on a specific failure: Gemma 4 26B, the best of the local models as a <em>worker</em>, could not orchestrate. It handled one…
<div class="ltag__link--embedded"> <div class="crayons-story "> <a class="crayons-story__hidden-navigation-link" href="https://dev.to/mikeross27/ai-agents-dont-need-more-context-they-need-memory-470o">AI Agents Don’t Need More Context. They Need Memory.</a> <div class="crayons-st…
<p>This LLM life cycle is an easy reading for an Infra admin comparing to Deployment life cycle of an OS. </p> <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to…
dev.to — LLM tag
TIER_1Polski(PL)·Andrzej Klusiewicz·
<p>Model LLM działający lokalnie, bez wysyłania danych do chmury i bez miesięcznego rachunku za API — Ollama robi to prościej, niż się wydaje. Przechodzę całość: instalację, pierwsze zapytanie, REST API i polskiego Bielika gadającego na żywo.</p> <p>Blog JSystems · AI dla program…
<p>Part one of this series ended on a specific boundary:LangChain's chain model, even with LCEL's clean pipe syntax, is built for pipelines that flow forward, not ones that loop, branch based on a runtime decision, or need to pause and resume days later. That boundary is exactly …
<p>Here's the state of building with LLMs in 2026: DeepSeek has one API, Moonshot has another, Zhipu GLM has a third, Alibaba's Qwen a fourth. Each with its own key, its own SDK quirks, its own billing console, its own free-tier rules.</p> <p>We built AIBridge to collapse all of …
dev.to — LLM tag
TIER_1English(EN)·SullivanReed1247·
<p>Short answer: put each US or EU extraction worker behind a region-local, rate-aware queue; validate every structured finding before completion; and use provider batch processing only for review work whose deadline can absorb asynchronous turnaround. A 429 is an admission-contr…
<p>You stood up vLLM on a box with a GPU. It works. You share the URL with a couple of teams and go back to your actual job.</p> <p>Here is what you just deployed:</p> <p>No authentication. vLLM doesn't ship with any by default. Anyone who learns the address can use it — includin…
dev.to — LLM tag
TIER_1English(EN)·Priyanka-Chettri·
<p>Before jumping into APIs, RAG, agents, and AI applications, I wanted to understand what actually happens inside an LLM.</p> <p>I kept coming across terms like <strong>neural networks, deep learning, Transformers, attention, tokens, embeddings, BERT, GPT, and causal language mo…
<p>A few weeks ago, I started seriously learning how Large Language Models actually work.</p> <p>Not how to <em>use</em> ChatGPT.</p> <p>Not how to write better prompts.</p> <p>Not how to call an LLM API from Python.</p> <p>I mean what actually happens <strong>after I press Enter…
<p>The logging rewrite is the second half of a provider migration that nobody schedules, because the log line was written by copying whatever the SDK response object happened to contain.</p> <h2> Why the log is coupled to the provider </h2> <p>The usual call log is a serialised r…
dev.to — LLM tag
TIER_1English(EN)·Divyakush Punjabi·
<p><strong>A large language model is frozen in time and confidently wrong about anything it wasn't trained on. Retrieval-Augmented Generation is how you fix both problems at once — without retraining the model.</strong></p> <p>RAG has become one of the most important patterns in …
dev.to — LLM tag
TIER_1English(EN)·Ravindranath Guptha K·
<blockquote> <p>Large Language Models have quickly become part of everyday software development.</p> </blockquote> <p>We ask them to explain code, debug errors, generate tests, write Python scripts, summarize documentation, or help us understand an unfamiliar codebase.</p> <p>Wit…
<p><strong>TL;DR</strong></p> <p>Traditional caching checks if you've seen the exact same input before. Semantic caching checks if you've seen a meaningfully similar input before — using embeddings instead of string matching. For LLM-backed apps, this can cut API costs by 30–70% …
<p>📦 Project: <a href="https://github.com/VampiricCyborg/sluice" rel="noopener noreferrer">https://github.com/VampiricCyborg/sluice</a></p> <h2> 1. The Problem: When Capacity Becomes the Bottleneck </h2> <p>A self-hosted vLLM deployment runs on a GPU pool of fixed size. That's th…
<p>Large language models can be adapted to enterprise applications in two common ways: <strong>Retrieval-Augmented Generation (RAG)</strong> and <strong>fine-tuning</strong>.</p> <p>They solve different problems.</p> <p>RAG gives an LLM access to external knowledge at query time,…
dev.to — LLM tag
TIER_1English(EN)·daxharrington5274·
<p>Short answer: submit historical posts and comments as batch work, persist the returned job identity beside the tenant, and poll that job instead of firing one LLM request per record.</p> <p>For an edtech knowledge base, that decision is less about raw model price than about re…
<p>Hypothesis generates inputs and checks that a property holds for all of them. Applied to a language model, the model is the system under test and the property is whatever you can assert about its output without knowing what the output will say. That second half is the part wor…
dev.to — LLM tag
TIER_1English(EN)·Adolfo Pedernera·
<p><strong>A comprehensive framework for deciding between local LLMs and cloud APIs. Covers cost, privacy, latency, control, and the hybrid approach.</strong></p> <h2> The Cathedral and the Bazaar, Revisited </h2> <p>In 1997, Eric Raymond published an essay that framed a fundamen…
<p>Use one bulk classification job over the whole archive, and make schema-valid JSON the acceptance criterion you check first — before model quality, before cost per row. If you need to moderate existing posts and comments after a policy change, the batch endpoint of whatever LL…
<p>There's a pattern I see every release cycle: a new budget-friendly model ships, the discourse explodes with hot takes, and within 48 hours half my feed has declared it a drop-in replacement for everything. The claim might even be true. But here's what nobody posting those take…
<p>Thirty seconds is about how long a nurse will wait for an answer from an internal-policy assistant before giving up and opening the PDF. That deadline, not the price list, is what decides this comparison. Pick a gateway on two numbers at once — what one answered question costs…
<h2> Evolusi Pemrosesan Data: Dari Batch ke Real-Time </h2> <p>Transformasi infrastruktur data mendorong transisi dari pemrosesan batch statis ke arsitektur streaming. Integrasi LLM kini menjadi komponen inti sistem data terdistribusi, di mana kecepatan pemrosesan informasi menen…
<p>Every time a new open model trends, the same ritual plays out: impressive demos on Monday, horror stories by Wednesday, and by Friday nobody remembers what they concluded. I used to ride that wave. What pulled me off it wasn't a better opinion — it was realizing I already solv…
<h2> The decision nobody writes down </h2> <p>Ask a team why they run model X in production and the honest answer is usually "someone tried it in a chat window and it looked good." That works right up until the inputs stop looking like the demo: an edge case arrives, the model bl…
<p>Every week there's a new model, a new benchmark chart, and a new thread arguing about which one is "smartest." Meanwhile, the question that actually matters for your project — <em>does this model handle my prompts, my edge cases, my output format?</em> — goes untested.</p> <p>…
dev.to — LLM tag
TIER_1English(EN)·JamesAnderson121·
<p>Short answer: move summarization, tagging, and extraction to batch LLM jobs when nobody needs the answer immediately, but keep realtime calls for interactive work; the useful savings come from making latency flexible and forecasting tokens before dispatch, not from assuming ev…
dev.to — LLM tag
TIER_1English(EN)·AidenSterling3417·
<p>The answer is conditional: move summarization, tagging, and extraction into a batch lane only when the work can wait, inputs can be replayed safely, and your measured all-in cost is lower than the realtime path. Keep interactive requests realtime. A provider's advertised disco…
<p>I had eight models sitting on my laptop and no idea which one I should actually be using.</p> <p>I could get tokens per second out of <code>llama-bench</code>. That told me <code>llama3.2</code> was fast. It told me nothing about whether <code>llama3.2</code> was <em>good</em>…
dev.to — LLM tag
TIER_1English(EN)·Emylton Leunufna·
<p>For years, most Large Language Models (LLMs) have started from the same assumption: <strong>Language is first broken into tokens, and computation happens on those tokens.</strong></p> <p>Whether it's BPE, SentencePiece, or WordPiece, the <em>tokenizer</em> remains one of the m…
<h1> How to Tell If an LLM Was Really Trained From Scratch: A Reproducible Fingerprinting Method </h1> <p><em>Detect whether an LLM was trained from scratch or derived from Qwen, Llama, or DeepSeek — by fingerprinting architecture, tokenizer, and weight provenance from public Hug…
<p>In this video, we perform a deep dive into evaluating Retrieval-Augmented Generation (RAG) performance using <em>Amazon Bedrock Knowledge Bases</em> and the <em>LLM-as-a-Judge</em> framework.</p> <p> </p>
<p>In December 2024, Jeremy Howard, the creator of fast.ai, proposed a simple standard: an llms.txt file placed in a domain's root directory, telling AI systems about the structure and content of the site. The standard is supported by Anthropic (Claude), Perplexity and a dozen or…
<p>The technical accuracy of content generated by LLMs is an undeniable requirement, especially in critical applications. When manual checks are insufficient to ensure the reliability of an LLM output, setting up an automated fact-check pipeline is the primary way to prevent the …
<p>The reflexive answer is “use a model” and the contrarian answer is “classical methods are underrated”. Both are slogans. The decision is a calculation with four inputs, and once you write it down it usually answers itself.</p> <h2> The four variables </h2> <ul> <li> <strong>Vo…
dev.to — LLM tag
TIER_1English(EN)·Lightning Developer·
<h2> Introduction to Modern Local LLM Workflows </h2> <p>Historically, fine-tuning an 8B parameter Large Language Model (LLM) required access to expensive enterprise hardware like the NVIDIA A100. Developers often faced the anxiety of whether their training run would complete bef…
<p>Our ML Services need LLMs to process the large documents and data. We started using Ollama since 2024 and the learnings below:</p> <p><strong>Ollama</strong>, an open-source tool that packages LLMs into a simple CLI and REST server, has emerged as a game-changer for developers…
<p>LLM (Large Language Model) API calls can become a significant expense due to token-based costs as applications scale. Prompt caching aims to directly reduce this cost by returning cached responses for previously answered or similar queries instead of re-querying the LLM. This …
<!-- SC_OFF --><div class="md"><p>We are investigating whether recurring LLM workloads can be replaced, where appropriate, by automatically constructed pipelines of regexes, deterministic parsers, traditional ML and NLP models. </p> <p>As an example, suppose an application repeat…
dev.to — LLM tag
TIER_1English(EN)·Dinesh_gowtham·
<p>Large Language Models can be finicky, but with the right prompts, they can produce astonishing results. However, crafting effective prompts is more art than science. What if you could systematically improve your prompts to get better outputs? </p> <h2> Introduction to Prompt E…
<h2> Introduction </h2> <p>LLM (Large Language Model) structured output plays a critical role in natural language processing and artificial intelligence applications. Using tools like JSON Schema, we can make LLM output more reliable and understandable. In this article, we will e…
dev.to — LLM tag
TIER_1English(EN)·jamesanderson3589·
<p><strong>Short answer:</strong> Backfill existing posts and comments with a resumable batch job, not a loop that sends one LLM classification request at a time; persist the source identity, poll job state, then export or fetch results before applying moderation flags in an idem…
<p>I've built BiasCheck: Critical Query Template. This is a tool designed to help users get more nuanced and reliable outputs from large language models, especially when dealing with sensitive subjects.</p> <p>The template works by structuring prompts to encourage the LLM to cons…
<blockquote> <p><strong>Before reading this</strong>: This is the final post in the series. I'd recommend reading the previous five — Part 1 covers the <code>.meph</code> format, Part 2 walks through a hands-on tutorial, Parts 3 and 4 explain how the parser outputs a <code>domain…
<p>Every week there's a new "best free model" thread, and every week I watch developers pick one based on vibes, a screenshot, or someone else's benchmark that has nothing to do with their workload. Then three days later they're rewriting prompts because the model falls apart on …
<p>Scrolling DEV this week, roughly every third post is about AI coding tools, agents, or model comparisons. Most of them share the same weakness: the benchmarks happened on someone else's machine, on someone else's code, with prompts I can't rerun. When the result disagrees with…
<p>Most developers pick a model the same way: they paste one prompt into a chat UI, like the answer, and start building on it. A week later the integration falls apart on edge cases nobody tested, and the 'evaluation' turns out to have been a vibe.</p> <p>I got tired of doing thi…
<blockquote> <p>Originally published on <a href="https://freedevkit.com/blog/rag-architecture-for-developers-enhancing-llm-accuracy-relevance/" rel="noopener noreferrer">FreeDevKit</a>.</p> </blockquote> <p>Retrieval Augmented Generation (RAG) architecture is a critical paradigm …
<p>If you've been using <a href="https://github.com/ray-project/llmperf" rel="noopener noreferrer">ray-project/llmperf</a>, you may have noticed it's now in archive mode. No new updates, no fixes, no responses to issues. If you're evaluating it for the first time, that's worth kn…
<p>This guide walks through a complete chatbot you can run locally in about two minutes and read in an afternoon: a FastAPI server, one static page, no database, no build step. It covers the design of each subsystem, the order things happen in on every turn, and the failure modes…
dev.to — LLM tag
TIER_1English(EN)·Prabhakar Chaudhary·
<p>Most business prediction problems do not arrive as prose. They arrive as rows: account attributes, transactions, sensor readings, test results, and a target column. For this kind of data, gradient-boosted trees and other conventional methods remain hard to displace.</p> <p>A n…
dev.to — LLM tag
TIER_1English(EN)·Mohammad Jawad (Kasir) Barati·
<p>Large Language Models (LLMs) like GPT are incredibly powerful, but they work in ways that are often counterintuitive. Here I wanna break down 3 core concepts that explain what's really happening under the hood.</p> <h2> tl;dr </h2> <ul> <li>LLMs are stateless.</li> <li>Some mo…
<p>Part 1 established the hardware, the runner, and the primary model. This entry covers what governs inference on that hardware — the two phases of inference, the cost of long context, and the cost of loading a model from disk — and compares the local machines against a free-tie…
<blockquote> <p><strong>TL;DR</strong> — A glossary to actually understand the terms you hit when reading about LLMs: token, embedding, attention, KV cache, GQA, MoE, quantization and the rest. But not alphabetical — <strong>in dependency order</strong>: every entry uses only con…
<blockquote> <p><strong>📖 Cross-posted.</strong> The original has interactive, step-through diagrams for every stage (attention, the feed-forward step, sampling, and more) that can't run here. For the full experience, <a href="https://codemug.github.io/blog/sre-guide-deploying-ll…
<p>Your GPU's TFLOPS rating does not decide how fast it generates tokens. Memory bandwidth does. Here's the spec that actually separates A100-era inference from H200-era inference.</p> <p><strong>Key takeaways</strong></p> <ul> <li><p>HBM3e reaches up to 9.6 Gb/s per pin versus H…
dev.to — LLM tag
TIER_1English(EN)·kapil Maheshwari·
<h2> Key takeaways </h2> <ul> <li>Batching can reduce costs by 30-50% in non-urgent tasks.</li> <li>Streaming minimizes latency but may lead to higher operational costs.</li> <li>Choosing the right strategy can increase reliability and user satisfaction.</li> <li>Understanding yo…
<table> <tr><td> <a href="https://www.reddit.com/r/MachineLearning/comments/1vdsgcj/context_degradation_in_llms_what_the_papers/"> <img alt="Context degradation in LLMs: what the papers actually show, and the habits I built for long analysis sessions [R]" src="https://external-pr…
<p>Context engineering is one of those terms that you understand when you actually build with LLMs in production. Then it becomes clear that this is the real work of deciding what information the model should see, how that information should be arranged, what should be remembered…
dev.to — LLM tag
TIER_1English(EN)·kapil Maheshwari·
<h2> Key takeaways </h2> <ul> <li>Implement fallback responses to maintain user engagement.</li> <li>Use caching mechanisms to reduce LLM calls during outages.</li> <li>Establish alerting systems for LLM performance degradation.</li> <li>Balance the trade-offs between user experi…
<p>Digital strategist and researcher Inna Udalaya (publishing under the pseudonym Inna Story) has introduced a formalized methodological framework designed to stabilize digital identity in algorithmic environments. Her research addresses a fundamental structural issue in modern A…
<p>This is the first in a series of build-log posts documenting a local LLM project, in which models are run on owned consumer hardware rather than through a cloud API. The present entry covers the hardware, the software stack, and the benchmarks by which a primary model was sele…
<p>Recently I read aboyt this article:<br /> <a href="https://arxiv.org/abs/2605.19537" rel="noopener noreferrer">The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility</a></p> <p>Here is what I learnt:<br /> LLM inference engines provide a…
dev.to — LLM tag
TIER_1English(EN)·Lightning Developer·
<p>The performance gap between proprietary coding models like Claude and GPT and open-weight alternatives has become remarkably small. As of August 2026, self-hosting is no longer about compromising on quality. It is about running production-ready coding assistants that keep sens…
<p>Most CI pipelines assume a function called with the same input twice returns the same output. That assumption breaks the moment an LLM call enters your test suite. Ask GPT-4 or Claude the same question twice and you can get two different (both correct) answers. Teams shipping …
<p>Every LLM evaluation framework today invents its own test case format, its own grader definitions, and its own results schema. DeepEval, Promptfoo, Inspect AI, and lm-evaluation-harness all solve the same core problem (checking whether a model's output is correct) but none of …
<p>Единый прокси решает одну проблему и тут же создаёт другую. Он убирает из приложений десятки ключей и base_url, но забирает себе все решения о маршрутизации, которые раньше были размазаны по коду. Пока прокси работает, это выглядит как чистая экономия сложности. В момент отказ…
dev.to — LLM tag
TIER_1English(EN)·Alina Trofimova·
<p>Everyone tells you to just drop a semantic cache in front of OpenAI or Anthropic to cut costs. I tried it. tbh, it's mostly an illusion.</p> <p>Here is the reality of AI agents: they don't ask the same questions twice. They inject dynamic context, timestamps, user-specific IDs…
<h2> Key takeaways </h2> <ul> <li>Semantic caching can reduce LLM costs by up to 70%.</li> <li>Accuracy risks arise when cached responses are stale or misaligned.</li> <li>Implementing effective cache invalidation strategies is crucial.</li> <li>Evaluate use cases carefully to ba…
<h2> LLM Context Management, PyTorch Attention Profiling & Open-Source LLM Code Review </h2> <h3> Today's Highlights </h3> <p>This week's highlights include a practical CLI for managing LLM context windows, a deep dive into profiling PyTorch attention for performance, and an …
<p>An exact-match cache misses "how do I reverse a list in Python" when it has already answered "python list reverse". A semantic cache doesn't: it embeds the prompt, finds the closest one it has seen, and replays that answer instead of calling the model. Fewer API calls, lower l…
<h2> Part 3: Speculative Decoding via Sampling-Mode Accept/Reject </h2> <p><em>Part 3 of a 4-part series on system-level LLM inference internals. Part 1 tracked entropy during decode. Part 2 measured attention sinks during prefill. This one implements sampling-mode speculative de…
<p>A validation failure is not an exception to hide. It is a record your system does not yet know how to trust.</p> <p>That distinction matters in LLM extraction pipelines. A malformed invoice, an unexpected OCR layout, a model response that violates the schema, and a semanticall…
<p>Having an optimal Batch size can decrease your models cost per token at the time of inference.</p> <p>This will be an explanation on how Batch size affects the cost at the inference, We will be going deep and building the framework from the ground up. This will be an informal …
<p><em>Why a solar production report generator treats grounding as a hard constraint, not a nice-to-have — and what actually enforcing that looks like in a system prompt.</em></p> <h2> The problem </h2> <p>Most PV monitoring dashboards are good at showing data and bad at explaini…
<p><strong>Agent Determinism Illusions (Part 6)</strong></p> <blockquote> <p><strong>Where this fits:</strong> <a href="https://dev.to/zxpmail/six-experiments-on-adversarial-verification-and-the-75-wall-that-didnt-move-2d1m">Part 5</a> closed the experimental arc with an honest a…
<p>Financial institutions have moved beyond experimenting with large language models, they’re learning how to run them in production now. Building a deployment that actually works is not simply about picking the latest model but integrating enterprise data, meeting regulatory nee…
<h2> How LLM-as-Judge Works </h2> <p>LLM-as-Judge delegates the quality evaluation of one LLM's output to another LLM (or the same one).</p> <p>Two basic forms:</p> <p><strong>Pointwise scoring (single answer)</strong><br /> </p> <div class="highlight js-code-highlight"> <pre cla…
dev.to — LLM tag
TIER_1English(EN)·Rahul Vijayvergiya·
<p><strong>Before we begin:</strong> The goal of this article is to help you understand how LLMs work in simple language. I've intentionally avoided deep technical details.</p> <p>If you've used ChatGPT, Claude, Gemini, or another AI chatbot, you've probably wondered:</p> <p><str…
<p>An LLM Wiki fails when old facts remain plausible, contradictions become polished, and generated summaries drift from their sources.</p> <p>Maintenance is the real product of any compiled knowledge system. Creating wiki pages is straightforward compared with keeping them trust…
dev.to — LLM tag
TIER_1English(EN)·Sreeraj Sreenivasan·
<p><em>Nine tools, three layers, one decision framework. Everything you need to run open-source models in 2026.</em></p> <h2> Why This Guide Exists </h2> <p>The local LLM inference ecosystem has quietly matured into one of the most consequential layers of the open-source AI stack…
<p>You've got 200k tokens. So why do you keep running out of room halfway through your API call?</p> <p>Most developers treat context like a gas tank—fill it up and hope you don't run empty. That's the wrong mental model. Context is inventory. You need to <em>manage</em> it.</p> …
<h2> What Changed </h2> <p>The emergence of Chain-of-Thought (CoT) reasoning has significantly advanced the capability of large language models (LLMs) to handle complex, multi-step tasks. However, a persistent challenge in human-AI interaction with these models has been the ineff…
<p>From data preparation and tokenizer selection to pretraining, LoRA, RLHF, evaluation, and production monitoring, this guide covers the major stages involved in training an AI model.</p> <p>Training an artificial intelligence model is not simply a matter of loading a dataset on…
<h2> Part 2: The Attention Sink Detector </h2> <p><em>Part 2 of a 4-part series on system-level LLM inference internals. Part 1 built the entropy tracker; this one looks one step earlier in the pipeline — at prefill, before a single token is generated.</em></p> <h2> Where This Fi…
<h2> What Changed </h2> <p>Prism ML has introduced Bonsai-27B, a 27B-class language model that leverages binary transformer weights, achieving a deployed footprint of approximately 3.9 GB. This represents a significant reduction in size, roughly 14.2 times smaller than its FP16 c…
<h2> What Changed </h2> <p>GnLOLot has introduced the MiniCPM5-1B-Claude-Opus-Fable5-Thinking model, a specialized 1-billion parameter language model designed to enhance coding and instruction-following performance. This new model is a fine-tuned version of the <code>openbmb/Mini…
<h2> Part 1: The Entropy Tracker </h2> <p><em>Part 1 of a 4-part series on system-level LLM inference internals.</em></p> <h2> What This Series Builds </h2> <p>Most LLM tooling treats inference as a black box. Hosted APIs make this worse; they strip away logits, attention weights…