English(EN)A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
新研究探讨LLM的忠实性、推理和评估问题
作者PulseAugur 编辑部·[95 个来源]·
近期研究探索了提高大型语言模型(LLM)可靠性和评估方法。几篇论文介绍了评估LLM对证据的忠实性、检测模型规范中的不一致性以及评估推理能力的框架和基准。一项研究提出了VERITYGATE,用于根据结构化证据检查LLM的叙述,揭示了GPT-4o-mini和Claude Sonnet 4.6等当前模型存在显著的失败率。另一篇论文介绍了VeriSpec,它使用LLM作为验证器来查找模型规范中的不一致性,并成功识别了OpenAI Model Spec中的问题。其他研究则专注于为科学AI创建本体论基础的基准,通过稳定性而非仅准确性来评估泛化能力,以及理解LLM如何基于推理方法而非主题来迁移数学知识。
AI
arXiv:2610.08840v1 Announce Type: new Abstract: Large language models (LLMs) often abandon a correct answer, or endorse a user's position, once the user pushes back. This behavior, called sycophancy, is usually reported as a single rate per model, which says little about when it …
arXiv cs.LG
TIER_1English(EN)·Harry Mayne, Justin Singh Kang, Dewi Gould, Kannan Ramchandran, Adam Mahdi, Noah Y. Siegel·
arXiv:2602.02639v2 Announce Type: replace-cross Abstract: LLM self-explanations are often presented as a promising tool for AI oversight, yet their faithfulness to the model's true reasoning process is poorly understood. Existing faithfulness metrics have critical limitations, ty…
arXiv:2610.09229v1 Announce Type: new Abstract: LLM-as-judge is now a standard tool for scalable evaluation, but judge performance is still often summarized by a single accuracy number. This aggregate view hides the deployment conditions under which a judge succeeds or fails. We …
arXiv cs.CL
TIER_1English(EN)·Xiaoshu Chen, Sihang Zhou, Ke Liang, Xinwang Liu·
arXiv:2609.33181v2 Announce Type: replace-cross Abstract: Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or…
arXiv:2609.35308v2 Announce Type: replace Abstract: Large language models treat conversation history as unverified context, so false premises injected into prior turns can be adopted as fact, a failure mode we term session-level contamination. We introduce five contamination prot…
arXiv cs.CL
TIER_1English(EN)·Hongru Cai, Ran Wei, Wenjie Wang, Chengfa Wu, Ning Song, Yongqi Li, Wenjie Li·
arXiv:2610.10533v1 Announce Type: new Abstract: Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this archite…
arXiv cs.AI
TIER_1English(EN)·Maoqi Liu, Quan Fang, Yufei He·
arXiv:2610.08312v1 Announce Type: new Abstract: Continual learning (CL) is essential for Large Language Models (LLMs) to sequentially adapt to evolving tasks. To mitigate catastrophic forgetting, recent advances implement low-rank adaptation with orthogonal projections (e.g., O-L…
arXiv:2610.08778v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, pre…
arXiv:2610.08563v1 Announce Type: new Abstract: Sequence-level power sampling has recently emerged as a training-free approach to reasoning by sampling from a sharpened output distribution of a base large language model (LLM). Nevertheless, existing methods typically sharpen the …
arXiv cs.AI
TIER_1English(EN)·Bach Nguyen, Zhaonan Li, Mau Son Nguyen, Sanika Chavan, Nilay Kumar, Hong Anh Nguyen, Khoa Vo, Ben Zhou·
arXiv:2610.07751v1 Announce Type: new Abstract: Real-world reasoning rarely reduces to static question answering: agents must actively gather information from tools and sensors that are often noisy and unreliable. Yet most existing active reasoning benchmarks assume that environm…
arXiv:2610.08129v1 Announce Type: new Abstract: Cooperation with unfamiliar partners requires adapting to communication conventions that are not known in advance. We study this problem in a controlled Hanabi-derived environment with scripted hint generation, LLM-controlled receiv…
arXiv cs.LG
TIER_1English(EN)·Chen Chen, Dongjie Wang, Mei Liu, Zijun Yao·
arXiv:2610.07739v1 Announce Type: new Abstract: Post-discharge risk prediction from electronic health records (EHRs) is difficult because many dependencies that link discharge-time observations to downstream complications, such as comorbidity cascades and drug-disease interaction…
arXiv cs.CL
TIER_1English(EN)·Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas-Hinostroza, Mattia Nee, Eliot Krzystof Jones, Ir\`ene Girard, David Mach, Anastasia Stasenko, Ivan P. Yamshchikov·
arXiv:2506.01732v4 Announce Type: replace Abstract: Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions of copyrighted or proprietary content, which rai…
arXiv cs.CL
TIER_1English(EN)·Sofia Torres, Gabriel Almeida, Carter Adams, Camila Rocha·
arXiv:2610.06861v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become the dominant paradigm for eliciting multi-step reasoning in large language models, and a recent wave of methods (LUFFY, ExPO, PAPO, TAPO) further augments RL with \e…
arXiv:2610.08018v1 Announce Type: new Abstract: Reliable tool use requires more than triggering a mechanism or matching a query to an API description. Before selecting a specific tool, an agent must first infer the capability requirements implied by the user query. In this paper,…
arXiv cs.CL
TIER_1English(EN)·Zizhuo Zhang, Xiong Peng, Jingwei Sun, Rong Yao, Borui Jiang, Bo Han·
arXiv:2610.07894v1 Announce Type: new Abstract: Large language models (LLMs) are expected to answer questions faithfully based on the provided context, abstaining when the context information is insufficient to answer the questions. Existing faithfulness evaluations typically ass…
arXiv cs.CL
TIER_1English(EN)·Yiqi Liu, Joseph James, Yang Wang, Kun Zhao, Chenghao Xiao, Chenghua Lin·
arXiv:2610.07109v1 Announce Type: new Abstract: When an LLM judge scores an output, its score distribution retains uncertainty and disagreement information that is lost after scalar compression. We introduce JudgeMoE, a lightweight aggregator that assigns example-specific weights…
arXiv cs.AI
TIER_1Français(FR)·Stephanie Buttigieg, Maeve Madigan, Parameswaran Kamalaruban, Stuart Burrell·
arXiv:2610.08559v1 Announce Type: cross Abstract: Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models. However, prior work reports that steering vectors generalise poorly, with unintended effects on model perform…
arXiv cs.AI
TIER_1English(EN)·Hans Schabert, Christoph Peters·
arXiv:2610.07817v1 Announce Type: cross Abstract: Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step. When an agent is the executor that predictabil…
arXiv:2610.07535v1 Announce Type: cross Abstract: Multi-agent simulations with large language models (LLMs) often operate networks of agents with a single base model. This overlooks the inter-model effects which may dominate engagement dynamics in real-world deployments. To show …
arXiv cs.AI
TIER_1English(EN)·Wonjun Lee, Kyungsik Yang, Gaeun Ji, Vaidehi Patil, Haon Park, Bumsub Ham, Mohit Bansal, Suhyun Kim·
arXiv:2610.07532v1 Announce Type: cross Abstract: LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks including defenses at decoding stage that leverage models' hidden states. However, e…
arXiv cs.AI
TIER_1English(EN)·Hashmath Shaik, Gnaneswar Villuri, Alex Doboli·
arXiv:2610.07115v1 Announce Type: cross Abstract: The order in which candidate responses are presented can change an LLM judge's verdict. Detecting such a position flip ordinarily requires judging each pair in both orders, which doubles the number of judgments. We investigate whe…
Large language models can extract useful signals from heterogeneous enterprise data, but high-recall extraction often produces outputs that are duplicated, uneven in granularity, semantically overlapping, or too numerous for downstream systems and human reviewers to use effective…
Multi-agent simulations with large language models (LLMs) often operate networks of agents with a single base model. This overlooks the inter-model effects which may dominate engagement dynamics in real-world deployments. To show this, we simulate a heterogeneous social network p…
Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria…
arXiv cs.AI
TIER_1English(EN)·Seoyeon Ye, Gayoung Kim, Jiyoung Hong, Sookyung Kim, Hyunsoo Cho·
arXiv:2610.02902v1 Announce Type: new Abstract: Current analyses of LLMs' parametric knowledge are largely output-centric, drawing conclusions about what a model knows without verifying what it was actually trained on. This leaves fundamental questions, such as whether a correct …
arXiv:2610.02793v1 Announce Type: new Abstract: Research on LLMs continually uncovers model limitations, their causes, and potential solutions. Yet these human discoveries remain largely disconnected from model evolution: an LLM does not automatically learn from new research abou…
arXiv:2602.10273v3 Announce Type: replace-cross Abstract: Reasoning ability in large language models is often attributed to \emph{distribution sharpening}: concentrating output probability on high-likelihood sequences. Recent works show that this sharpening effect can be obtained…
arXiv:2605.10247v2 Announce Type: replace Abstract: Applying Large Language Models (LLMs) to graph-structured data usually involves multi-step pipelines in which textual node attributes are compressed into single tokens and further processed by GNNs, discarding most of their sema…
arXiv:2603.13725v2 Announce Type: replace Abstract: Memristor-based analog compute-in-memory (CIM) architectures provide a promising substrate for the efficient deployment of Large Language Models (LLMs), owing to superior energy efficiency and computational density. However, the…
arXiv:2610.03080v1 Announce Type: cross Abstract: Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader descr…
arXiv:2604.14128v3 Announce Type: replace-cross Abstract: Rhetorical questions are asked not to seek information but to persuade or signal stance. How large language models internally represent them remains unclear. We analyze rhetorical questions in LLM representations using lin…
arXiv:2603.29112v2 Announce Type: replace Abstract: We introduce GISTBench, a benchmark for evaluating Large Language Models' (LLMs) ability to understand users from their interaction histories in recommendation systems. Unlike traditional RecSys benchmarks that focus on item pre…
arXiv cs.AI
TIER_1English(EN)·Katherine Van Koevering, Jon Kleinberg·
arXiv:2406.00092v2 Announce Type: replace Abstract: Randomness is central to human cognition and to many applications in which large language models are deployed, yet probabilistic token generation does not imply that LLMs can produce unbiased random sequences. We study how conte…
arXiv:2610.02571v1 Announce Type: cross Abstract: As AI-assisted programming becomes increasingly mainstream, the environmental impact of AI-generated software has emerged as an important consideration. This motivates evaluating LLM-generated code beyond functional correctness by…
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this…
arXiv:2610.00833v1 Announce Type: cross Abstract: Fluent LLM explanations may not follow the evidence from a structured system. We present VERITYGATE, a four-gate checker for declared evidence IDs, entities, numbers, and claim types. It checks a fixed schema; it does not verify e…
arXiv:2610.01847v1 Announce Type: cross Abstract: Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet these specifications may themselves contain defects: two individually reasonable …
arXiv:2602.03006v3 Announce Type: replace Abstract: Deploying Large Language Models (LLMs) for discriminative workloads is often limited by inference latency, compute, and API costs at scale. Active distillation reduces these costs by querying an LLM oracle to train small discrim…
arXiv:2608.28421v2 Announce Type: replace Abstract: Post training a language model to reason means updating its weights. Supervised finetuning and reinforcement learning both place the acquired capability inside the model where it cannot be inspected cannot be checked step by ste…
arXiv cs.AI
TIER_1English(EN)·Hongbo Chen, Guohua Lu, Ting Dang, Hong Jia·
arXiv:2609.33149v2 Announce Type: replace Abstract: A common principle of effective learning is to practice material that is neither already mastered nor too difficult to permit progress. We ask how to apply this principle to structured reasoning tasks such as Sudoku and maze sol…
arXiv:2609.39801v1 Announce Type: new Abstract: Post-training quantization (PTQ) has become a widely adopted technique for reducing the memory footprint and inference cost of large language models (LLMs). However, recent studies reveal that when applied to reasoning models, PTQ n…
arXiv:2610.01471v1 Announce Type: cross Abstract: Large language models now generate code, documentation, and analyses, and are increasingly used to review such output. We ask when a second review by a different model helps. Building on the author's earlier preprints, which varie…
arXiv cs.AI
TIER_1English(EN)·Nagham Omar, Mahmoud Jabarin, Maya Rozenshtein, Rom Himelstein, Avi Mendelson, Amit LeVi·
arXiv:2610.01428v1 Announce Type: cross Abstract: Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggre…
arXiv:2610.00562v1 Announce Type: cross Abstract: Longitudinal clinical reasoning requires large language models (LLMs) to identify and integrate relevant evidence distributed across extended patient histories. Although long-context models can process increasingly large amounts o…
arXiv cs.AI
TIER_1English(EN)·Matthieu Zimmer, Xiaotong Ji, Tu Nguyen, Haitham Bou-Ammar·
arXiv:2610.00332v1 Announce Type: cross Abstract: Distilling the reasoning capabilities of large language models (LLMs) into smaller students is a central challenge for efficient deployment. Current approaches face a fundamental tension: optimizing purely for verifiable task rewa…
arXiv:2610.00296v1 Announce Type: cross Abstract: Token-level certainty is widely used as a proxy for correctness in LLM training and inference. However, the performance of certainty-based methods depends both on the information in certainty scores and on how those scores are use…
arXiv cs.AI
TIER_1English(EN)·Nishtha N. Vaidya, Stephan Grimm, Thomas Hubauer, Thomas A. Runkler·
arXiv:2610.00682v1 Announce Type: new Abstract: Large language models (LLMs) increasingly underpin scientific AI applications that reason over structured knowledge, from biomedical question answering to materials informatics. However, their logical reasoning often falls short, pr…
arXiv:2610.00331v1 Announce Type: new Abstract: When selecting mathematical training data for LLMs, a natural organizing principle is topic: probability examples for probability targets. An alternative is reasoning approach: worked solutions that share a solution method with the …
arXiv cs.AI
TIER_1English(EN)·Khawaja Murad ul Hassan, Mehran Ebrahimi·
arXiv:2610.00047v1 Announce Type: new Abstract: Diversity collapse in parallel chain-of-thought has motivated inference-time interventions built on a natural design: when a process reward model (PRM) prunes a chain, its high-PRM prefix is extracted and grafted verbatim as an in-c…
arXiv:2609.39572v1 Announce Type: new Abstract: We analyze how lexical ambiguity and underspecification affect language model training. We create artificial homonyms and artificial hypernyms as pseudowords and analyze the generative performance of language models as they are trai…
arXiv:2609.38260v1 Announce Type: cross Abstract: Values such as honesty, autonomy, and confidentiality are often regarded as general principles underpinning AI alignment. However, what it means to act in accordance with these values can depend on the context in which a decision …
arXiv:2609.38672v1 Announce Type: cross Abstract: Beam-search-based test-time methods provide an effective way to improve large language model (LLM) performance on long-horizon generation by pruning invalid reasoning paths early, leading to significantly improved reasoning effici…
arXiv:2609.38829v1 Announce Type: new Abstract: Multi-path reasoning methods such as self-consistency (SC) sample $K$ reasoning paths and choose the most frequent answer. However, their gains quickly plateau as $K$ increases, and existing methods do not predict when this saturati…
arXiv:2609.38972v1 Announce Type: cross Abstract: Chain-of-thought (CoT) traces often serve as a proxy for how Large Language Models (LLMs) arrive at their answers. However, growing evidence shows that models' CoT often fails to reflect their internal computations and can be chan…
arXiv:2609.39967v1 Announce Type: cross Abstract: Recursive reasoning models apply a small shared Transformer block many times to refine a latent state. This gives them large effective depth with few parameters and makes them strong on algorithmic tasks. Such compact solvers are …
arXiv:2609.39346v1 Announce Type: new Abstract: Large language models (LLMs) offer strong reasoning capabilities but are often costly to access through commercial APIs, while small language models (SLMs) are easier to deploy locally yet remain weaker in reasoning. This capability…
Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or …
Recursive reasoning models apply a small shared Transformer block many times to refine a latent state. This gives them large effective depth with few parameters and makes them strong on algorithmic tasks. Such compact solvers are natural candidates for tools that an LLM can call …
arXiv:2609.37915v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher …
arXiv cs.LG
TIER_1English(EN)·Qihao Wen, Jiahao Wang, Yang Nan, Pengfei He, Ravi Tandon, Han Xu·
arXiv:2602.02427v3 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved significant breakthroughs across various domains, but they can still produce unreliable or misleading outputs. For responsible LLM applications, uncertainty quantification techniques ar…
arXiv cs.LG
TIER_1English(EN)·Shenghong Dai, Shiva Kumar Pentyala, Yingchi Liu, Shubham Mehrotra, Suman Banerjee, James Zhu, Bin Bi, Sitaram Asur, Phil Mui·
arXiv:2609.36115v1 Announce Type: new Abstract: Industry applications often demand low-latency classification, yet current large language model (LLM) approaches remain poorly suited for latency-critical applications. Existing prompting and constrained decoding produce verbose, mu…
arXiv:2609.36159v1 Announce Type: new Abstract: Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final deco…
arXiv:2609.37054v1 Announce Type: new Abstract: Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe r…
arXiv:2609.38070v1 Announce Type: new Abstract: As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current meth…
arXiv cs.AI
TIER_1English(EN)·Gurbir Arora, Toni J. B. Liu, Jiajun Bao, Rapha\"el Sarfati, Christopher J. Earls·
arXiv:2609.35970v1 Announce Type: cross Abstract: Large language models are able to solve tasks whose answers depend on not only individual input tokens, but also on relations among them. How is such relational information represented and processed across transformer layers? We s…
arXiv cs.AI
TIER_1English(EN)·Ali Mohammadi Esfahani, Nafiseh Kahani, Samuel A. Ajila·
arXiv:2609.37405v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for software engineering tasks that require understanding existing source code, including behavior prediction, function explanation, debugging, and code review. However, aggregate…
arXiv:2602.11328v2 Announce Type: replace Abstract: As people turn to LLMs for social advice, understanding their behavior in such contexts becomes essential. In this work, we focus on behavioral dispositions: the underlying tendencies that shape responses in social contexts. We …
arXiv cs.CL
TIER_1English(EN)·Thomas Reiter, Christoph Kern, Fedor Miasnikov, Sofiia Nikolenko, Rob Chew, Stephanie Eckman, Frauke Kreuter·
arXiv:2609.35824v1 Announce Type: new Abstract: Large language models (LLMs) can give reliable labels under one setup yet change those labels when researchers make other reasonable design choices. We tested seven LLMs, 12 task designs, three independent runs, and 3,000 tweets lab…
arXiv cs.CL
TIER_1English(EN)·Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois, Pavel Chizhov, Carlos Rosas-Hinostroza, Neil Si Smail, Benjamin Burtin, Hanna Shcharbakova, Ivan Yamshchikov, Anastasia Stasenko·
arXiv:2609.37891v1 Announce Type: new Abstract: Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--for instance, they contain little explicit reasoning. Thus, many frontier labs have…
arXiv:2609.38137v1 Announce Type: new Abstract: Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated acc…
Chain-of-thought (CoT) traces often serve as a proxy for how Large Language Models (LLMs) arrive at their answers. However, growing evidence shows that models' CoT often fails to reflect their internal computations and can be changed without affecting their final answers. In this…
Large language models (LLMs) can now improve themselves by revising the instructions they follow, and LLM agents are increasingly orchestrated to work together on complex problems. However, self-improvement methods typically optimize one system at a time, and multi-agent framewor…
As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confidence of large langu…
On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solut…
We study reasoning in Rule-Governed Decision Tasks (RGDTs), where models apply external rules to case facts and justify decisions, as required in policy, contract, and compliance settings. Beyond the deductive capability emphasized by standard mathematical and logical reasoning t…
Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. T…
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-ins…
arXiv:2610.00809v1 Announce Type: new Abstract: Vision-Language Models (VLMs) enable video annotation at scale, but costs accumulate quickly: processing a typical 60-second short-form video at one frame per second requires millions of tokens. To reduce costs, researchers rely on …
Medium — fine-tuning tag
TIER_1English(EN)·Joya Parveen·
<p>Most AI-generated strategies suffer from the same terminal flaw: they are forward-only. If you ask an LLM to propose a market entry plan or a scaling architecture, it will present a linear progression of successful steps. It describes the ascent, but it never maps the cliffs.<…
dev.to — MCP tag
TIER_1English(EN)·Renato Marinho·
<p>Large Language Models are notoriously bad at arithmetic. When you ask a model to interpret complex financial oscillators, it isn't performing calculus; it is predicting the most likely next token based on training data. In technical analysis—where a decimal error in a volatili…
Medium — fine-tuning tag
TIER_1English(EN)·SINAPSA Infocomplex·
<p>Large Language Models (LLMs) have changed how we interact with computers.</p> <p>Tools such as ChatGPT, Claude, Gemini, and many open-source models can write code, explain complex topics, summarize documents, translate languages, and generate natural-sounding conversations.</p…
<h1> LLM Evaluation: How a Benchmark Turns Raw Answers Into Comparable Numbers </h1> <p><strong>TL;DR</strong> — LLM evaluation only produces comparable numbers when every model faces the same prompts, the same fixed judge, and the same per-axis rubrics, with a public verbatim tr…
<h2> <strong>Mistral Large 4 (“Le Chonk”) Takes the Stage: What the Benchmarks Really Reveal</strong> </h2> <h3> A Sharper Lead </h3> <p>On <strong>October 3 2026</strong>, Mistral released a <strong>1‑trillion‑parameter</strong> model that anyone can download under an open‑weigh…
dev.to — LLM tag
TIER_1English(EN)·Denis Lavrentyev·
<h2> Introduction: The Rise of LLMs and Their Impact </h2> <p>The integration of <strong>Large Language Models (LLMs)</strong> into professional workflows has sparked a critical debate: <em>Is advanced knowledge of local LLMs essential for leveraging AI tools effectively?</em> Th…
dev.to — LLM tag
TIER_1English(EN)·Filipe Martins·
<h1> How LLMs Work: A Journey Through Tokens, Attention, and Transformers </h1> <p>“My favourite rock band is…”</p> <p>How does a large language model take that unfinished sentence and decide what comes next?</p> <p>You’ve probably heard that LLMs “predict the next token”. But th…
<h2> TL;DR </h2> <p>LLM evaluation only becomes useful when every model faces the same prompts, the same fixed judge, per-axis rubrics, and a public verbatim trail. The LFORLA Reverse Engineering benchmark does exactly that: it restores C source from stripped binaries, scores wit…
dev.to — LLM tag
TIER_1English(EN)·Lightning Developer·
<p>Running large language models (LLMs) on microcontrollers has long been considered a "what if" scenario reserved for theoretical discussions. However, the recent emergence of the <a href="https://github.com/Low-Zi-Hong/ESP32s3-LLM-Cluster" rel="noopener noreferrer">ESP32s3-LLM-…
<p>Fine-tuning a large language model on your own data used to require serious GPU budgets and weeks of infrastructure work. LoRA (Low-Rank Adaptation) changed the math: you can adapt a 7B-parameter model to a specific domain in a few hours on a single consumer GPU, without touch…
A mechanism explainer on LLM evaluation using the Team Recruitment Oracle benchmark: same prompts, a fixed judge, per-axis rubrics, and a public verbatim trail. Real scores for Nemotron 3 Ultra, HY3, and our own GLM 5.2. # llm # evaluation # benchmark # ai # software # coding # d…
<!-- SC_OFF --><div class="md"><p>Source: <a href="https://github.com/ying-hui-he/Hi-ToM_dataset">https://github.com/ying-hui-he/Hi-ToM_dataset</a></p> <p>I asked several LLMs to solve this Hi-ToM puzzle:</p> <p>The following story happens in chronological order. You will be give…