PulseAugur
实时 20:10:44
English(EN) Making LLMs more accurate by using all of their layers

Google Research 评估大语言模型对齐并提高事实准确性

Google Research 开发了一个新的框架来评估大语言模型与人类社会倾向的行为对齐情况。该方法将已建立的心理学问卷改编成大规模情境判断测试,从而能够量化模型在现实场景中的倾向。研究发现了模型行为偏离人类共识或未能捕捉人类意见范围的差距,旨在改善大语言模型在社会动态中的导航能力。另外,Google Research 还推出了 SLED,这是一种新颖的解码策略,通过利用模型的所有层而不是仅最后一层来提高大语言模型的准确性,且无需外部数据或微调。 AI

影响 评估大语言模型对齐和提高事实准确性的新方法可能带来更值得信赖和更善于社交的AI系统。

排序理由 该集群包含两篇来自 Google Research 的研究论文,详细介绍了评估大语言模型对齐和提高大语言模型事实准确性的新方法。

在 Google AI / Research 阅读 →

AI 生成摘要 · Google Gemini · 来自 598 个来源。 我们如何撰写摘要 →

Google Research 评估大语言模型对齐并提高事实准确性

报道来源 [598]

  1. Google AI / Research TIER_1 English(EN) ·

    评估大型语言模型行为倾向的对齐性

    Generative AI

  2. Google AI / Research TIER_1 English(EN) ·

    通过利用其所有层来提高 LLM 的准确性

    Algorithms & Theory

  3. Hugging Face Blog TIER_1 English(EN) ·

    Consilium:当多个大型语言模型协同工作时

  4. Hugging Face Blog TIER_1 English(EN) ·

    使用 KVPress 掌握 LLM 的长上下文

  5. Hugging Face Blog TIER_1 English(EN) ·

    Judge Arena:将大型语言模型作为评估者进行基准测试

  6. Hugging Face Blog TIER_1 English(EN) ·

    专家支持案例研究:使用 LLM-as-a-Judge 加强 RAG 应用

  7. Hugging Face Blog TIER_1 English(EN) ·

    CodeGemma - 谷歌官方发布的代码大语言模型

  8. Hugging Face Blog TIER_1 English(EN) ·

    推出 Open Ko-LLM 排行榜:引领韩国 LLM 评估生态系统

  9. Hugging Face Blog TIER_1 English(EN) ·

    开源LLM作为LangChain代理

  10. Hugging Face Blog TIER_1 English(EN) ·

    隆重推出 Agents.js:为您的 LLM 赋予 JavaScript 工具

  11. arXiv cs.LG TIER_1 English(EN) · Advik Raj Basani, Anshuman Chhabra ·

    揭示大型语言模型知识编辑中“擦除幻觉”的真相

    arXiv:2606.23276v2 Announce Type: replace Abstract: Knowledge Editing (KE) has emerged as a frontier for updating specific facts in LLMs without costly retraining, but its reliability and underlying mechanisms remain poorly understood. In this work, we examine KE from an adversar…

  12. arXiv cs.AI TIER_1 English(EN) · Anshuman Chhabra ·

    揭示大型语言模型知识编辑中“擦除幻觉”的真相

    Knowledge Editing (KE) has emerged as a frontier for updating specific facts in LLMs without costly retraining, but its reliability and underlying mechanisms remain poorly understood. In this work, we examine KE from an adversarial elicitation perspective, revealing that edited k…

  13. arXiv cs.CL TIER_1 English(EN) · Haw-Shiuan Chang, Jeffrey Gomez, Mehul Patwari, Aryan Sajith, Hamed Zamani ·

    你的鼠标和眼睛会秘密泄露你的偏好:利用用户隐式反馈进行LLM对齐

    arXiv:2606.20482v1 Announce Type: new Abstract: To align a Large Language Model (LLM), most existing methods collect explicit human feedback and train a reward model to predict the human preference based on the response text. These existing methods have two key limitations. First…

  14. arXiv cs.AI TIER_1 English(EN) · Zunchen Huang, Songgaojun Deng ·

    分析LLM求解器循环中的叙事差距

    arXiv:2606.19588v1 Announce Type: new Abstract: Formal tools such as SAT and SMT solvers are increasingly embedded in language model reasoning pipelines when a safety or security critical question can be formulated in logic. Unlike chain of thought whose steps are sampled from th…

  15. arXiv cs.CL TIER_1 English(EN) · Hamed Zamani ·

    你的鼠标和眼睛会秘密泄露你的偏好:利用用户隐式反馈进行LLM对齐

    To align a Large Language Model (LLM), most existing methods collect explicit human feedback and train a reward model to predict the human preference based on the response text. These existing methods have two key limitations. First, the users rarely provide explicit feedback for…

  16. arXiv cs.MA (Multiagent) TIER_1 Nederlands(NL) · Sankalp Nayak ·

    异构大模型在对抗性同伴下的争论:诚实的收益、替代成本和弹性

    Heterogeneous LLM debate is motivated by the promise that diverse peers correct one another, but the same exchange that carries correction also carries adversarial influence. We measure which dominates by tracking how a heterogeneous peer changes the honest agents' revision behav…

  17. arXiv cs.CL TIER_1 English(EN) · Naihao Deng, Yiming Feng, Chimaobi Okite, Kaijian Zou, Lu Wang, Rada Mihalcea, Yulong Chen ·

    错误的正确:量化和定位大型语言模型中失准的对齐问题

    arXiv:2606.18656v1 Announce Type: new Abstract: Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment. Instead, this paper high…

  18. arXiv cs.AI TIER_1 English(EN) · Sunnie S. Y. Kim, Margit Bowler, Leon A Gatys ·

    探究大型语言模型中的类人行为:模型行为、用户因素和系统提示的多维度分析

    arXiv:2606.18258v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit a wide range of human-like behaviors, from expressing thoughts and emotions, to engaging in relationship-building with users, to refusing requests and maintaining boundaries. Despite their prev…

  19. arXiv cs.LG TIER_1 English(EN) · Zilong Zhang, Yi-Ting Hung, Lei Ding, Chi-Kuang Yeh ·

    通过正负样本学习量化和审计大语言模型评估

    arXiv:2606.19057v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used as judges for scalable evaluation, yet such LLM--as--a--Judge systems exhibit systematic biases that are decoupled from semantic quality, most notably verbosity bias. Meanwhile, h…

  20. arXiv cs.AI TIER_1 English(EN) · Xi Fang, Weijie Xu, Yuchong Zhang, Stephanie Eckman, Scott Nickleach, Chandan K. Reddy ·

    个性化陷阱:用户记忆如何改变大型语言模型的情感推理

    arXiv:2510.09905v2 Announce Type: replace Abstract: When an AI assistant remembers that Sarah is a single mother working two jobs, does it interpret her stress differently than if she were a wealthy executive? As personalized AI systems increasingly incorporate long-term user mem…

  21. arXiv cs.AI TIER_1 English(EN) · Mika M\"antyl\"a, Patricia Matsubara, Katia Romero Felizardo, Miikka Kuutila, Marco Gerosa, Savio de Sousa Sampaio, Tayana Conte, Igor Steinmacher ·

    理解LLMs在标题-摘要筛选中的应用:从分歧到建议

    arXiv:2606.17588v1 Announce Type: cross Abstract: Several studies have examined the use of large language models (LLMs) for title-abstract screening in systematic reviews (SRs), reporting mixed accuracy. However, questions of reliability remain largely unaddressed. In this study,…

  22. arXiv cs.CL TIER_1 English(EN) · Hyungwon Kim, Kandarp Joshi, Lillian Zhou, Pavel Golik, Petar Aleksic ·

    你说我的语言吗?关于多模态大语言模型中的口语依从性

    arXiv:2606.17281v1 Announce Type: new Abstract: While Large Language Model (LLM) based Automatic Speech Recognition (ASR) enables seamless multilingual use, models often misidentify the output language, compromising transcription fidelity and downstream application quality. To pr…

  23. arXiv cs.CL TIER_1 English(EN) · Ali Marashian, Alexis Palmer, Katharina von der Wense ·

    能言善辩,自我评估:论大型语言模型在机器翻译中的口头自信

    arXiv:2606.17234v1 Announce Type: new Abstract: The rapid rise in popularity of large language models (LLMs) for translation calls for a thorough study of the reliability of their confidence in their own outputs. Unlike many generation tasks, translation errors and confidence lev…

  24. arXiv cs.LG TIER_1 English(EN) · SongEun Kim, Seungyoo Lee, Edwin Fong, Hyungi Lee, Juho Lee ·

    从漂移到连贯:稳定大型语言模型中的信念

    arXiv:2606.17832v1 Announce Type: new Abstract: Large language models (LLMs) are often hypothesized to perform implicit Bayesian inference, yet a key coherence condition, the martingale property of predictive beliefs, has been shown to fail in controlled synthetic in-context lear…

  25. arXiv cs.CL TIER_1 English(EN) · Omar Sharif, Eftekhar Hossain, Nikhil Singh, Patrick Ng ·

    通过奖励设计解耦多模态大模型中的感知与推理

    arXiv:2601.00215v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards has driven major gains in LLM reasoning, and it is intuitive to assume this recipe will transfer well to multimodal models. However, multimodal models do two things: first, pe…

  26. arXiv cs.CL TIER_1 English(EN) · Rui Wen, Lu Sun, Jiayang Liu, Zesheng Xu, Tianshuo Cong, Zheng Li ·

    基准测试的幻觉:剪枝的LLM可以通过多项选择题但无法回答问题

    arXiv:2606.17609v1 Announce Type: new Abstract: Compressing large language models reduces memory use and inference cost, but it can also create failures that standard benchmarks miss. A pruned model may still perform well on multiple-choice evaluations, yet fail to answer the sam…

  27. arXiv cs.CL TIER_1 English(EN) · Ramaravind Kommiya Mothilal, Terry Jingchen Zhang, Raiyan Ahmed, Zhijing Jin, Shion Guha, Syed Ishtiaque Ahmed ·

    通过认知权利评估LLM的二阶偏差

    arXiv:2606.17506v1 Announce Type: new Abstract: Evaluations of social bias in LLMs largely focus on whether models generate or imply biased content. However, as LLMs are increasingly used as judges of bias, they may exhibit social biases in subtler ways in how they evaluate biase…

  28. arXiv cs.CL TIER_1 English(EN) · Yulong Chen ·

    错误的好:量化和定位 LLM 中失准的对齐

    Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment. Instead, this paper highlights the need for principled approaches to mor…

  29. arXiv cs.LG TIER_1 English(EN) · Juho Lee ·

    从漂移到连贯:稳定大型语言模型中的信念

    Large language models (LLMs) are often hypothesized to perform implicit Bayesian inference, yet a key coherence condition, the martingale property of predictive beliefs, has been shown to fail in controlled synthetic in-context learning settings. We revisit this question in a mor…

  30. arXiv cs.CL TIER_1 English(EN) · Zheng Li ·

    基准测试的幻觉:剪枝的LLM可以通过多项选择但无法回答

    Compressing large language models reduces memory use and inference cost, but it can also create failures that standard benchmarks miss. A pruned model may still perform well on multiple-choice evaluations, yet fail to answer the same question in open generation. We ask what pruni…

  31. arXiv cs.CL TIER_1 English(EN) · Syed Ishtiaque Ahmed ·

    通过认知授权评估LLM的二阶偏差

    Evaluations of social bias in LLMs largely focus on whether models generate or imply biased content. However, as LLMs are increasingly used as judges of bias, they may exhibit social biases in subtler ways in how they evaluate biased content, which current methods do not systemat…

  32. arXiv cs.CL TIER_1 English(EN) · Pan Wang ·

    REFLEX: LLM经验的反射式进化

    arXiv:2606.16496v1 Announce Type: new Abstract: Large multimodal language models (LLMs) have emerged as powerful tools for guiding evolutionary search toward interpretable programmatic policies. However, existing frameworks rely on a monolithic model call to simultaneously interp…

  33. arXiv cs.LG TIER_1 English(EN) · Violet Xiang, Amrith Setlur, Chase Blagden, Nick Haber, Aviral Kumar ·

    ExpRL:LLM中期训练的探索性强化学习

    arXiv:2606.17024v1 Announce Type: new Abstract: Sparse reward reinforcement learning (RL) has become a standard tool for improving LLM reasoning, but its success depends critically on the coverage present in the base model. In practice, models are often primed for RL through \emp…

  34. arXiv cs.CL TIER_1 English(EN) · Katharina Trinley, Jesujoba O. Alabi, Dietrich Klakow, Vagrant Gautam ·

    大型语言模型中代词保真度的机制性理解

    arXiv:2606.16407v1 Announce Type: new Abstract: Faithful and robust pronoun use is important for fair and coherent generations, yet large language models largely fail when multiple referents use different pronouns. To study the interplay of reasoning, repetition, and bias in this…

  35. arXiv cs.CL TIER_1 English(EN) · Xuran Li, Guanqin Zhang, Imran Razzak, Hakim Hacid, Eleanna Kafeza, Hao Xue, Flora D. Salim ·

    通过语义约束验证评估大语言模型个性化

    arXiv:2606.16368v1 Announce Type: new Abstract: Current evaluation paradigms for Large Language Model (LLM) personalization rely heavily on brittle surface-matching metrics or computationally expensive LLM-as-a-judge protocols, both of which lack interpretability. To address thes…

  36. arXiv cs.CL TIER_1 English(EN) · Nafiseh Nikeghbal, Amir Hossein Kargaran, Shaghayegh Kolli, Jana Diesner ·

    谁会翻转?自身模型和跨模型反驳揭示了大型语言模型答案的不稳定性

    arXiv:2606.16011v1 Announce Type: new Abstract: Standard accuracy benchmarks are designed to test how closely large language models (LLMs) approach correct answers, but are not suitable for testing whether LLMs stick with a correct answer when that answer is challenged by a plaus…

  37. arXiv cs.AI TIER_1 English(EN) · Erica Zhang, Fangzhao Zhang, Aneesh Pappu, Batu El, Jose Blanchet, Susan Athey, Jiashuo Liu, James Zou ·

    TERMS-Bench:超越成交率诊断LLM谈判代理

    arXiv:2605.13909v2 Announce Type: replace-cross Abstract: Negotiation is a central mechanism of economic exchange, shaping markets, procurement, labor agreements, and resource allocation. It is also a canonical testbed for agentic language models, requiring multi-turn interaction…

  38. arXiv cs.AI TIER_1 English(EN) · Zhen Yang, Mingyang Zhang, Feng Chen, Ganggui Ding, Liang Hou, Xin Tao, Ying-Cong Chen ·

    少即是多:通过最小的测试时干预改进 LLM 推理

    arXiv:2510.13940v4 Announce Type: replace-cross Abstract: Recent progress in large language models (LLMs) has focused on test-time scaling to improve reasoning via increased inference computation, but often at the cost of efficiency. We revisit test-time behavior and uncover a si…

  39. arXiv cs.AI TIER_1 English(EN) · Uljad Berdica, Fernando Acero, Anton Ipsen, Parisa Zehtabi, Michael Cashmore, Manuela Veloso ·

    何时需要大型语言模型?语言驱动的强盗问题诊断

    arXiv:2604.05859v2 Announce Type: replace Abstract: We study Contextual Multi-Armed Bandits (CMABs) for non-episodic decision-making problems where the context includes both textual and numerical information (e.g., recommendation systems, dynamic portfolio adjustments, offer sele…

  40. arXiv cs.AI TIER_1 English(EN) · Louie Hong Yao, Nicholas Jarvis, Tiffany Zhan, Saptarshi Ghosh, Linfeng Liu, Tianyu Jiang ·

    JE-IRT:通过联合嵌入项目反应理论对大语言模型能力进行几何视角审视

    arXiv:2509.22888v2 Announce Type: replace Abstract: Standard LLM evaluation practices compress diverse abilities into single scores, obscuring their inherently multidimensional nature. We present JE-IRT, a geometric item-response framework that embeds both LLMs and questions in a…

  41. arXiv cs.AI TIER_1 English(EN) · Ismail Hossain, Sai Puppala, Md Jahangir Alam, Tanzim Ahad, Sajedul Talukder ·

    SkillVetBench:LLM作为裁判,对开源LLM智能体技能进行多维度安全风险评估

    arXiv:2606.15899v1 Announce Type: cross Abstract: Open-source LLM agent ecosystems are growing rapidly, yet the security of community-contributed skills - modular tool definitions that extend agent capabilities - remains largely unvetted. The gap we fill: existing scanners operat…

  42. arXiv cs.AI TIER_1 English(EN) · Hiroyasu Usami, Keisuke Hara, Ayato Tsuboi, Naohiko Matsuda ·

    LLM 裁判的黑暗现状:LLM 作为裁判评估的心理测量数据表

    arXiv:2606.15610v1 Announce Type: cross Abstract: LLM-as-a-judge systems are now routinely used for open-ended model evaluation, where human preference annotation is costly, slow, and difficult to reproduce. Yet these judges are often reported as scalar accuracy, win-rate, or agr…

  43. arXiv cs.AI TIER_1 English(EN) · Aina Vila Pons, Ioannis Tzachristas, Constantinos Antoniou ·

    有限语义下的表格数据大型语言模型:来自工业汽车改装预测的证据

    arXiv:2606.15314v1 Announce Type: cross Abstract: Industrial retrofit planning depends on structured operational data rather than free text: planners must estimate whether a newly registered prototype will require a retrofit, which retrofit package it will need, and how long the …

  44. arXiv cs.AI TIER_1 English(EN) · Olivia Peiyu Wang, Sanna Wong-Toropainen, Daneshvar Amrollahi, Ryan Bai, Tashvi Bansal, Arush Garg, Leilani H. Gilpin ·

    了解你的局限性:关于大型语言模型作为法律推理的求解器和自动形式化器的忠实度

    arXiv:2606.16118v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong performance on reasoning tasks, but whether this reflects faithful logical inference or heuristic approximation remains unclear. We study this question in legal entailment by comparing thr…

  45. arXiv cs.AI TIER_1 English(EN) · Yitao Li ·

    谁漂移了:系统还是法官?LLM评估管道中的随时有效归因

    arXiv:2606.15474v1 Announce Type: new Abstract: Continuous evaluation of LLM products relies on a strong LLM judge treated as ground truth: a cheap monitor scores every interaction and a team is paged when the score drifts down. But the judge is itself a model behind an API, and …

  46. arXiv cs.AI TIER_1 English(EN) · Alyssa Unell, Natalie Dullerud, Naomi Boneh, Meena Jagadeesan, Tatsu Hashimoto, Nigam Shah, Sanmi Koyejo ·

    Metric Match:一种用于评估 LLM 裁判可靠性的子集选择方法

    arXiv:2606.15029v1 Announce Type: new Abstract: LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation. However, the reliability of these judges depends critically on their alignment with human raters -- a property that itself depen…

  47. arXiv cs.AI TIER_1 English(EN) · Louis Mahon, Elliot Ford, Callum Hackett ·

    什么是好的解释以及解释LLM输出的挑战

    arXiv:2606.14838v1 Announce Type: new Abstract: How to define a good explanation is a long-standing philosophical debate which has found recent renewed interest in the context of AI outputs. Explainability is crucial for AI adoption in many contexts, but in order to produce good …

  48. arXiv cs.CL TIER_1 English(EN) · Petar Aleksic ·

    你说我的语言吗?关于多模态大语言模型在口语遵循方面的研究

    While Large Language Model (LLM) based Automatic Speech Recognition (ASR) enables seamless multilingual use, models often misidentify the output language, compromising transcription fidelity and downstream application quality. To preserve flexibility and code-switching capabiliti…

  49. arXiv cs.CL TIER_1 English(EN) · Katharina von der Wense ·

    能言善辩,自我评估:论大型语言模型在机器翻译中的口头自信

    The rapid rise in popularity of large language models (LLMs) for translation calls for a thorough study of the reliability of their confidence in their own outputs. Unlike many generation tasks, translation errors and confidence levels can be useful at different levels of granula…

  50. arXiv cs.LG TIER_1 English(EN) · Aviral Kumar ·

    ExpRL:LLM中期训练的探索性强化学习

    Sparse reward reinforcement learning (RL) has become a standard tool for improving LLM reasoning, but its success depends critically on the coverage present in the base model. In practice, models are often primed for RL through \emph{mid-training} on curated reasoning traces that…

  51. arXiv cs.CL TIER_1 English(EN) · Pan Wang ·

    REFLEX: LLM经验的反射式进化

    Large multimodal language models (LLMs) have emerged as powerful tools for guiding evolutionary search toward interpretable programmatic policies. However, existing frameworks rely on a monolithic model call to simultaneously interpret visual behavioral evidence and synthesize co…

  52. arXiv cs.CL TIER_1 English(EN) · Vagrant Gautam ·

    大型语言模型中代词保真度的机制性理解

    Faithful and robust pronoun use is important for fair and coherent generations, yet large language models largely fail when multiple referents use different pronouns. To study the interplay of reasoning, repetition, and bias in this task, prior work relies exclusively on behaviou…

  53. arXiv cs.CL TIER_1 English(EN) · Flora D. Salim ·

    通过语义约束验证评估大语言模型个性化

    Current evaluation paradigms for Large Language Model (LLM) personalization rely heavily on brittle surface-matching metrics or computationally expensive LLM-as-a-judge protocols, both of which lack interpretability. To address these limitations, we introduce Natural Language Inf…

  54. arXiv cs.AI TIER_1 English(EN) · Yash Pulse, Yong-Bin Kang, Abhik Banerjee, Abdur Forkan, Prem Prakash Jayaraman ·

    FactoryLLM:一个安全且开源的AI游乐场,用于在智能工厂中评估LLM

    arXiv:2606.14119v1 Announce Type: new Abstract: Fault diagnostics and recovery in smart factories is challenging because critical information is dispersed across manuals of multiple machines which are interconnected through the manufacturing process. Large Language Models (LLMs) …

  55. arXiv cs.CL TIER_1 English(EN) · Li Zhang, Yuzhen Shi, Yiran Hu, Jingwen Zhang, Wenbo Lv, Yubo Ma, Wei Wang, Rongyao Shi, Yuanyang Qiu, Xinran Xu, Yuemeng Qi, Linlin Miao, Jaromir Savelka, Yun Liu, Kevin Ashley, Bing Zhao, Hu Wei, Lin Qu ·

    DLawBench:通过多轮法律咨询评估大型语言模型

    arXiv:2606.13931v1 Announce Type: new Abstract: Lawyer-client consultation is a critical starting point for legal services. Effective legal assistance hinges on eliciting sufficient and truthful information from clients in order to devise strategies that best protect their intere…

  56. arXiv cs.AI TIER_1 English(EN) · Toni J. B. Liu, Baran Zadeo\u{g}lu, Nicolas Boull\'e, Rapha\"el Sarfati, Gurbir Arora, Christopher J. Earls ·

    Jacobian Scopes:LLM 中的 token 级因果归因

    arXiv:2601.16407v3 Announce Type: replace-cross Abstract: Large language models (LLMs) make next-token predictions based on clues present in their context, such as semantic descriptions and in-context examples. Yet, elucidating which prior tokens most strongly influence a given p…

  57. arXiv cs.AI TIER_1 English(EN) · Shaun Feakins, Ibrahim Habli, Kim Littler, Robert Palin ·

    抱歉,司机先生,恐怕我不能那样做:评估大型语言模型在汽车环境中的安全性

    arXiv:2606.14327v1 Announce Type: cross Abstract: This paper appraises recent frameworks within AI development to integrate LLMs into control tasks in automotive contexts from the perspective of safety assurance. This work has built upon the rapid integration of LLMs across autom…

  58. arXiv cs.AI TIER_1 English(EN) · Abel Yagubyan ·

    抛硬币的裁判?LLM作为裁判评估的可靠性与偏见

    arXiv:2606.13685v1 Announce Type: cross Abstract: LLM-as-a-Judge is now widely used to rank model outputs, train reward models, and populate public leaderboards, but its run-to-run reliability remains under-characterized. We study repeated identical evaluations on 29 tasks spanni…

  59. arXiv cs.LG TIER_1 English(EN) · Angira Sharma, Christian Schroeder de Witt, Philip Torr, Anisoara Calinescu, Jialin Yu ·

    LLM干预的低秩子空间分析

    arXiv:2606.14388v1 Announce Type: new Abstract: Interventions designed to modify a particular behavior in LLMs, such as refusal or sycophancy, often produce unintended changes in other behaviors. This lack of targeted control makes it difficult to design and implement reliable sa…

  60. arXiv cs.CL TIER_1 English(EN) · Yufeng Xu, Taiming Lu, Kunjun Li, Jiachen Zhu, Mingjie Sun, Zhuang Liu ·

    小型语言模型:剪枝 vs. 从头训练

    arXiv:2606.14150v1 Announce Type: cross Abstract: Pruning promises a shortcut to strong small language models. In this work, we examine this promise by pruning Llama-3.1-8B at pruning ratios of 0.5--0.8 with six methods spanning depth, width, and sparse granularities, under two c…

  61. arXiv cs.CL TIER_1 English(EN) · Filip Trhlik, Aoife O'Flynn, Angela Yu, Arduin Findeis, Paula Buttery ·

    大型语言模型包罗万象:部署环境如何重塑模型层面的偏好和价值观

    arXiv:2606.13944v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly characterised in recent evaluation work as having stable, model-level preference and value systems. However, accompanying robustness checks are limited to incidental prompt perturbations…

  62. Hugging Face Daily Papers TIER_1 English(EN) ·

    ExpRL:LLM中期训练的探索性强化学习

    ExpRL uses human-written question-answer data as reward scaffolds to provide automated reinforcement learning priming for language models, outperforming traditional methods on math reasoning tasks.

  63. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Sajedul Talukder ·

    SkillVetBench:LLM作为裁判,对开源LLM智能体技能进行多维度安全风险评估

    Open-source LLM agent ecosystems are growing rapidly, yet the security of community-contributed skills - modular tool definitions that extend agent capabilities - remains largely unvetted. The gap we fill: existing scanners operate at the code layer and are structurally blind to …

  64. Hugging Face Daily Papers TIER_1 English(EN) ·

    谁会翻转?自身模型和跨模型反驳揭示了大型语言模型答案的不稳定性

    Answer stability in large language models is evaluated through controlled challenges that measure response consistency when correct answers face plausible counterarguments, revealing significant variation in model reliability beyond traditional accuracy metrics.

  65. arXiv cs.LG TIER_1 English(EN) · Jialin Yu ·

    LLM干预的低秩子空间分析

    Interventions designed to modify a particular behavior in LLMs, such as refusal or sycophancy, often produce unintended changes in other behaviors. This lack of targeted control makes it difficult to design and implement reliable safety controls. To understand these side-effects,…

  66. arXiv cs.AI TIER_1 English(EN) · Robert Palin ·

    抱歉,司机,我恐怕不能这样做:评估大型语言模型在汽车环境中的安全性

    This paper appraises recent frameworks within AI development to integrate LLMs into control tasks in automotive contexts from the perspective of safety assurance. This work has built upon the rapid integration of LLMs across automotive settings. However, we find that at present, …

  67. arXiv cs.CL TIER_1 English(EN) · Zhuang Liu ·

    小型语言模型:剪枝 vs. 从头训练

    Pruning promises a shortcut to strong small language models. In this work, we examine this promise by pruning Llama-3.1-8B at pruning ratios of 0.5--0.8 with six methods spanning depth, width, and sparse granularities, under two controlled token-matched settings. (1) With the sam…

  68. arXiv cs.AI TIER_1 English(EN) · Prem Prakash Jayaraman ·

    FactoryLLM:用于在智能工厂中评估大型语言模型的安全开源AI游乐场

    Fault diagnostics and recovery in smart factories is challenging because critical information is dispersed across manuals of multiple machines which are interconnected through the manufacturing process. Large Language Models (LLMs) can provide a promising approach. In this paper,…

  69. arXiv cs.CL TIER_1 English(EN) · Sangho Kim, Heejin Kim, Yoonhee Park, Hyunggeun Jeon, Jaejin Lee ·

    Polar:评估大型语言模型政治偏见的基准

    arXiv:2606.12922v1 Announce Type: new Abstract: Political bias in large language models (LLMs) is increasingly significant, but difficult to measure reproducibly across political and linguistic contexts. We introduce Polar, a 4,026-instance multiple-choice benchmark that measures…

  70. arXiv cs.AI TIER_1 English(EN) · Ashutosh Hathidara, Sai Shruthi Sistla, Sebastian Schreiber, Sahil Bansal ·

    ToolSense:一个用于审计大型语言模型中参数化工具知识的诊断框架

    arXiv:2606.12451v1 Announce Type: new Abstract: Large language models deployed as agents over large tool catalogs face a critical tool-retrieval bottleneck. As embedding-based retrieval approaches rely on compact encoders that may under-capture specialized tool semantics, paramet…

  71. arXiv cs.AI TIER_1 English(EN) · Alyssa Unell, Miguel Fuentes, Brenna Li, Bridget Lin, Meena Jagadeesan, Sanmi Koyejo, Nigam Shah ·

    部署为中心评估:预测临床 LLM 系统中的查询级拒绝风险

    arXiv:2606.12702v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into clinical systems, making it essential to evaluate the real-world utility of these systems. However, static benchmarks tend to measure correctness rather than user accepta…

  72. arXiv cs.AI TIER_1 English(EN) · Danica Dillion, Chen Cecilia Liu, Baihui Wang, Daniele Barolo, Tanmay Rajore, Niket Tandon, Pranathi Ravikumar, Kurt Gray ·

    大型语言模型可以通过正确的提示更好地捕捉人类判断

    arXiv:2606.12754v1 Announce Type: cross Abstract: Are large language models (LLMs) bad at capturing human judgment? Two commonly stated limitations are that LLMs fail to capture full distributions of responses, and that their judgments are unstable across wording variations. We d…

  73. arXiv cs.AI TIER_1 English(EN) · Benno Krojer, Shravan Nayak, Oscar Ma\~nas, Vaibhav Adlakha, Desmond Elliott, Siva Reddy, Marius Mosbach ·

    LatentLens:揭示LLM中高度可解释的视觉Token

    arXiv:2602.00462v4 Announce Type: replace-cross Abstract: Transforming a large language model (LLM) into a vision-language model (VLM) can be achieved by mapping the visual tokens from a vision encoder into the embedding space of an LLM. Intriguingly, this mapping can be as simpl…

  74. arXiv cs.CL TIER_1 English(EN) · Aviya Maimon, Amir DN Cohen, Gal Vishne, Shauli Ravfogel, Reut Tsarfaty ·

    从基准到技能:用于 LLM 评估的低秩因子

    arXiv:2507.20208v2 Announce Type: replace Abstract: Current evaluations of large language models (LLMs) rely heavily on a growing collection of benchmarks and on aggregate benchmark scores, yet it remains unclear what this comparison actually captures, and what these scores revea…

  75. arXiv cs.CL TIER_1 English(EN) · Laura Majer, Jan \v{S}najder, Martin Tutek ·

    通过潜在视角评估LLM中的多元主义

    arXiv:2606.13254v1 Announce Type: new Abstract: The growing need to represent diverse perspectives has increased interest in pluralistic LLM generation. Although difficult to operationalize, identifying perspectives expressed in text would provide clear guidance on pluralistic al…

  76. arXiv cs.CL TIER_1 English(EN) · Camilla Dalerci, Thilo Michael, Robin Schaefer, Daniel Weinland ·

    M\"OVE:面向德国公共部门的整体性LLM基准测试

    arXiv:2606.13111v1 Announce Type: new Abstract: We present M\"OVE (Modelle f\"ur die \"Offentliche Verwaltung Evaluieren), a holistic benchmark for evaluating large language models (LLMs) in the context of the German public sector. While LLMs are increasingly adopted in public ad…

  77. arXiv cs.CL TIER_1 English(EN) · Paula Buttery ·

    大型语言模型包罗万象:部署环境如何重塑模型层面的偏好和价值观

    Large language models (LLMs) are increasingly characterised in recent evaluation work as having stable, model-level preference and value systems. However, accompanying robustness checks are limited to incidental prompt perturbations such as syntax variation and option reordering.…

  78. arXiv cs.CL TIER_1 English(EN) · Lin Qu ·

    DLawBench:通过多轮法律咨询评估大型语言模型

    Lawyer-client consultation is a critical starting point for legal services. Effective legal assistance hinges on eliciting sufficient and truthful information from clients in order to devise strategies that best protect their interests. This task requires Large Language Models (L…

  79. arXiv cs.CL TIER_1 English(EN) · Martin Tutek ·

    通过潜在视角评估LLM中的多元主义

    The growing need to represent diverse perspectives has increased interest in pluralistic LLM generation. Although difficult to operationalize, identifying perspectives expressed in text would provide clear guidance on pluralistic alignment and more clearly articulate the pluralis…

  80. arXiv cs.LG TIER_1 English(EN) · David Salinas ·

    从不确定性判断到校准排名:用于 LLM 评估的共形 Elo 估计

    Evaluating new large language models typically requires costly human annotation campaigns at scale. LLM-as-a-judge offers a cheaper alternative, but judge scores carry systematic errors - such as position bias, self-preference, or intransitivity - that can strongly miscalibrate t…

  81. arXiv cs.CL TIER_1 English(EN) · Daniel Weinland ·

    MÖVE:面向德国公共部门的整体性LLM基准测试

    We present MÖVE (Modelle für die Öffentliche Verwaltung Evaluieren), a holistic benchmark for evaluating large language models (LLMs) in the context of the German public sector. While LLMs are increasingly adopted in public administration, model selection remains largely ad hoc, …

  82. arXiv cs.CL TIER_1 English(EN) · Jaejin Lee ·

    Polar:评估大型语言模型政治偏见的基准

    Political bias in large language models (LLMs) is increasingly significant, but difficult to measure reproducibly across political and linguistic contexts. We introduce Polar, a 4,026-instance multiple-choice benchmark that measures political bias through option-level likelihoods…

  83. arXiv cs.CL TIER_1 English(EN) · Kiril Georgiev, Yuxia Wang, Dimitar Iliyanov Dimitrov, Preslav Nakov, Ivan Koychev ·

    Schützen: 在保加利亚和德国背景下评估 LLM 安全性

    arXiv:2606.11316v1 Announce Type: new Abstract: Large language models are increasingly deployed across professional domains, bringing hard-to-predict risks, including the generation of harmful or disrespectful content. Although substantial progress has been made in developing saf…

  84. arXiv cs.AI TIER_1 English(EN) · Kaituo Zhang, Mingzhi Hu, Hoang Anh Duy Le, Fariha Kabir Torsha, Zhimeng Jiang, Minh Khai Bui, Chia-Yuan Chang, Yu-Neng Chuang, Zhen Xiong, Ying Lin, Guanchu Wang, Na Zou ·

    LLM生成数据质量与可信度评估调查研究

    arXiv:2601.17717v3 Announce Type: replace Abstract: Large Language Models (LLMs) have emerged as powerful tools for generating data across various modalities. By transforming data from a scarce resource into a controllable asset, LLMs mitigate the bottlenecks imposed by the acqui…

  85. arXiv cs.CL TIER_1 Deutsch(DE) · Sanjay Adhikesaven, Haoxiang Sun, Sewon Min ·

    我们的模型基于哪些模型?审计现代大型语言模型中看不见的依赖项

    arXiv:2606.12385v1 Announce Type: new Abstract: Modern LLM training pipelines increasingly rely on other models to generate data, filter corpora, judge outputs, and guide development decisions. These dependencies are recursive: a model may depend on an upstream artifact whose own…

  86. arXiv cs.CL TIER_1 English(EN) · Hongjian Zhou, Xinyu Zou, Jinge Wu, Sean Wu, Junchi Yu, Bradley Max Segal, Tobias Erich Niebuhr, Sara Amro, Michael Petrus, Sheikh Momin, Alexandra M. Cardoso Pinto, Rachel Niesen, Laura Sophie Wegner, Dhruv Darji, Jung Moses Koo, Joshua Fieggen, Kapil N… ·

    衡量大型语言模型在误导性医疗背景下的认知韧性

    arXiv:2606.12291v1 Announce Type: new Abstract: Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption that high scores imply safe medical judgment while patients increasingly use them for health advice. We show this assu…

  87. arXiv cs.AI TIER_1 English(EN) · Orion Reblitz-Richardson ·

    当准确性饱和时,脆弱性得以解决:LLM预训练分析的互补指标

    arXiv:2606.11375v1 Announce Type: cross Abstract: Standard linear probing declares a property "encoded" when a classifier on hidden states achieves high accuracy. The protocol works well on a snapshot but breaks across pre-training: probe accuracy saturates within the first few t…

  88. arXiv cs.LG TIER_1 English(EN) · Shuo Yang, Qihui Zhang, Yuyang Liu, Xiaojun Jia, Kunpeng Ning, Jiayu Yao, Jigang Wang, Hailiang Dai, Yibing Song, Li Yuan ·

    AsFT:在狭窄安全区域内进行大语言模型微调时的安全锚定

    arXiv:2506.08473v4 Announce Type: replace Abstract: Fine-tuning large language models (LLMs) improves performance but introduces critical safety vulnerabilities: even minimal harmful data can severely compromise safety measures. We observe that perturbations orthogonal to the ali…

  89. arXiv cs.CL TIER_1 English(EN) · Muhammed Saeed, Simon Razniewski ·

    LLMpedia:一个透明的框架,用于大规模实现大型语言模型的百科全书式知识

    arXiv:2603.24080v2 Announce Type: replace Abstract: Benchmarks like MMLU suggest flagship language models approach factuality saturation above 90\%. \emph{LLMpedia} shows this picture is incomplete. We materialize ${\sim}$1.3M encyclopedia articles entirely from parametric memory…

  90. arXiv cs.CL TIER_1 English(EN) · Dongryeol Lee, Yerin Hwang, Taegwan Kang, Minwoo Lee, Younhyung Chae, Kyomin Jung ·

    以参考为基准的评判:揭示LLM评判员在QA评估中的知识驱动型失败

    arXiv:2601.07506v2 Announce Type: replace Abstract: While large language models (LLMs) are increasingly used as automatic judges for question answering (QA) and other reference-conditioned evaluation tasks, little is known about their ability to adhere to a provided reference. We…

  91. arXiv cs.CL TIER_1 Deutsch(DE) · Sewon Min ·

    我们的模型基于哪些模型?审计现代大型语言模型中看不见的依赖项

    Modern LLM training pipelines increasingly rely on other models to generate data, filter corpora, judge outputs, and guide development decisions. These dependencies are recursive: a model may depend on an upstream artifact whose own dependencies are documented only in separate re…

  92. arXiv cs.CL TIER_1 English(EN) · David A. Clifton ·

    衡量误导性医疗背景下大语言模型认知韧性

    Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption that high scores imply safe medical judgment while patients increasingly use them for health advice. We show this assumption is fragile: when misleading context is in…

  93. arXiv cs.AI TIER_1 English(EN) · Sophie Hao, William Merrill ·

    训练利润最优大语言模型的理论

    arXiv:2605.16430v2 Announce Type: replace-cross Abstract: Scaling LLMs requires tremendous computational resources, and recent advances in AI have gone hand in hand with massive amounts of capital expenditure. While it is established that scaling up LLMs reliably increases model …

  94. arXiv cs.AI TIER_1 English(EN) · Zhepei Wei, Xiao Yang, Kai Sun, Jiaqi Wang, Rulin Shao, Jingxiang Chen, Mohammad Kachuee, Teja Gollapudi, Yiwei Liao, Nicolas Scheffer, Rakesh Wanga, Anuj Kumar, Yu Meng, Wen-tau Yih, Xin Luna Dong ·

    TruthRL:通过强化学习激励诚实的LLM

    arXiv:2509.25760v2 Announce Type: replace-cross Abstract: While large language models (LLMs) have demonstrated strong performance on factoid question answering, they are still prone to hallucination and untruthful responses, particularly when tasks demand information outside thei…

  95. arXiv cs.AI TIER_1 English(EN) · Gabriel Freedman, Francesca Toni ·

    LLM 决策中的表面信念

    arXiv:2606.11016v1 Announce Type: new Abstract: We ask whether large language models (LLMs) merely imitate rationales when choosing between two options, or whether their choices reflect a systematic underlying decision structure. Using synthetic binary decision settings in which …

  96. arXiv cs.LG TIER_1 English(EN) · Ke Li, Chongzhe Zhang, Zifan Zeng, Feng Liu, Qunli Zhang, Zheng Hu ·

    校准过度自信而不牺牲信心:用于大型语言模型的探针条件头部干预

    arXiv:2606.09876v1 Announce Type: new Abstract: Large language models often express high confidence in answers that are wrong. Standard calibration remedies typically act globally or at the score level, reducing unwarranted confidence but also risking erosion of warranted confide…

  97. arXiv cs.CL TIER_1 English(EN) · Qian Zhu, Xinnan Guo, Jingjing Huo, Jun Li, Pan Liu, Wenyan Yang, Wanqing Xu, Xuan Lin ·

    一种工业级保险大模型,在不牺牲能力的情况下,实现可验证的领域精通和幻觉控制

    arXiv:2603.14463v2 Announce Type: replace Abstract: Adapting Large Language Models (LLMs) to high-stakes vertical domains like insurance presents a significant challenge: scenarios demand strict adherence to complex regulations and business logic with zero tolerance for hallucina…

  98. arXiv cs.CL TIER_1 English(EN) · Naihao Deng, Sheng Zhang, Henghui Zhu, Shuaichen Chang, Jiani Zhang, Alexander Hanbo Li, Chung-Wei Hang, Hideo Kobayashi, Yiqun Hu, Patrick Ng ·

    表格大语言模型真正重要的是什么?模型和数据效应的元评估

    arXiv:2501.14717v2 Announce Type: replace Abstract: Table modeling has progressed for decades. In this work, we revisit this trajectory and highlight emerging challenges in the LLM era, particularly the paradox of choice: the difficulty of attributing performance gains amid diver…

  99. arXiv cs.CL TIER_1 English(EN) · Yupu Hao, Zhuoran Jin, Huanxuan Liao, Kang Liu, Jun Zhao ·

    通过经验知识整合与激活,突破大型语言模型工具调用的极限

    arXiv:2606.10875v1 Announce Type: new Abstract: Large language models (LLMs) rely on tool use to act as autonomous agents, yet often fail in multi-step execution due to insufficient tool-related knowledge and ineffective knowledge activation. Therefore, we present a systematic st…

  100. Hugging Face Daily Papers TIER_1 Deutsch(DE) ·

    我们的模型基于哪些模型?审计现代大型语言模型中看不见的依赖项

    ModSleuth is an agentic system that recursively reconstructs large-scale dependency graphs for LLM development by analyzing public artifacts and resolving inconsistencies in documentation and artifact identities.

  101. arXiv cs.AI TIER_1 English(EN) · Francesca Toni ·

    LLM 决策中的表面信念

    We ask whether large language models (LLMs) merely imitate rationales when choosing between two options, or whether their choices reflect a systematic underlying decision structure. Using synthetic binary decision settings in which models choose between profiles defined by graded…

  102. arXiv cs.CL TIER_1 English(EN) · Jun Zhao ·

    通过经验知识整合与激活,突破大型语言模型工具调用的极限

    Large language models (LLMs) rely on tool use to act as autonomous agents, yet often fail in multi-step execution due to insufficient tool-related knowledge and ineffective knowledge activation. Therefore, we present a systematic study on how knowledge influences tool-use perform…

  103. arXiv cs.AI TIER_1 English(EN) · Yanyan Luo, Xue Han, Ruiqiao Bai, Xin Huang, Yitong Wang, Qian Hu, Qing Wang, Chunxu Zhao, Jie Liu, Cong Geng, Lehao Xing, Pengwei Hu, Junlan Feng ·

    个性化遇上安全:个性化LLM中的机制、风险与缓解措施

    arXiv:2606.09038v1 Announce Type: new Abstract: Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' preferences, contexts, and long-term histories. However, the mechanisms that enable personalization also expand the safety landsc…

  104. arXiv cs.AI TIER_1 English(EN) · Mohsinul Kabir, Tasnim Ahmed, Md Mezbaur Rahman, Shaoxiong Ji, Hassan Alhuzali, Yuechen Jiang, Jimin Huang, Sophia Ananiadou ·

    XCR-Bench:通过特定文化项目和霍尔三元组对大型语言模型进行跨文化推理基准测试

    arXiv:2601.14063v2 Announce Type: replace-cross Abstract: Cross-cultural competence in large language models (LLMs) requires understanding and adapting Culture-Specific Items (CSIs) across varying cultural contexts. However, progress in evaluating this capability remains limited …

  105. arXiv cs.AI TIER_1 English(EN) · Yasushi Sakai, Allen Song, Kent Larson ·

    何时委托优于多数?一种基于委托的多样本LLM推理聚合器

    arXiv:2606.08098v1 Announce Type: new Abstract: Majority voting over sampled answers is the dominant unsupervised aggregator for multi-sample LLM inference. We show that piping the signals every sample carries into a delegation-based aggregator (Propagational Proxy Voting, PPV) y…

  106. arXiv cs.AI TIER_1 English(EN) · Anissa Alloula, Federico Licini, Ava Batchkala, Seraphina Goldfarb-Tarrant ·

    安全性是情境化的,大语言模型裁判并非如此:驾驭评估者的僵化先验

    arXiv:2606.07874v1 Announce Type: new Abstract: LLMs-as-judges are the only way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement in simple, static benchmarks. We therefore investigate two under-explored but c…

  107. arXiv cs.AI TIER_1 English(EN) · Gaojie Jin, Yong Tao, Lijia Yu, Tianjin Huang ·

    面向可靠LLM判断的边距自适应置信度排序

    arXiv:2605.15416v2 Announce Type: replace-cross Abstract: Jung et al. (2025) introduce a hypothesis testing framework for guaranteeing agreement between large language models (LLMs) and human judgments, relying on the assumption that the model's estimated confidence is monotonic …

  108. arXiv cs.AI TIER_1 English(EN) · Jianhui Chen, Yuzhang Luo, Liangming Pan ·

    机制化数据归因:追踪可解释大模型单元的训练起源

    arXiv:2601.21996v2 Announce Type: replace-cross Abstract: While Mechanistic Interpretability has identified interpretable circuits in LLMs, their causal origins in training data remain elusive. We introduce Mechanistic Data Attribution (MDA), a scalable framework that employs Inf…

  109. arXiv cs.LG TIER_1 English(EN) · Dennis Frauen, Athiya Deviyani, Mihaela van der Schaar, Stefan Feuerriegel ·

    基于偏好数据的非参数大模型评估

    arXiv:2601.21816v2 Announce Type: replace Abstract: Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards. However, many existing approaches either rely on restrictive parametric assumptions or lack valid u…

  110. arXiv cs.AI TIER_1 English(EN) · Haoran Xu ·

    Cherry-pick Override:LLM裁判在混合证据下的不安全定向承诺

    arXiv:2606.07834v1 Announce Type: cross Abstract: LLM judges increasingly turn verdicts into system commitments. Under mixed evidence (claims with both supporting and refuting sources) this is unsafe: when the schema exposes CONFLICTING as the authorized non-directional verdict, …

  111. arXiv cs.CL TIER_1 English(EN) · Yerzhan Sapenov, Jaromir Savelka ·

    mmPISA-bench:大型语言模型在43种语言中的推理能力是否均等?

    arXiv:2606.07069v1 Announce Type: new Abstract: We introduce mmPISA-bench, a compact high-quality multilingual reasoning benchmark derived from the OECD Programme for International Student Assessment (PISA). The benchmark consists of 25 multiple-choice questions that require reas…

  112. arXiv cs.CL TIER_1 English(EN) · Parisa Rabbani, Priyam Sahoo, Ruben Mathew, Aishee Mondal, Harshita Ketharaman, Nimet Beyza Bozdag, Dilek Hakkani-T\"ur ·

    DialDefer:检测和缓解 LLM 对话式服从的框架

    arXiv:2601.10896v2 Announce Type: replace Abstract: LLMs are increasingly used as third-party judges, yet their reliability when evaluating speakers in dialogue remains poorly understood. We show that LLMs judge identical claims differently depending on framing: the same content …

  113. arXiv cs.CL TIER_1 English(EN) · Shaiv Patel, Kartik Narayan, Vishal Patel ·

    PromptPrint:通过LLM的自然语言提示实现行为生物识别

    arXiv:2606.06755v1 Announce Type: new Abstract: Authorship attribution research has traditionally focused on long-form, expressive texts; however, interactions with large language models (LLMs) are typically brief and task-driven prompts. This raises a fundamental question: do su…

  114. arXiv cs.AI TIER_1 English(EN) · Guoxi Zhang, Jiawei Chen, Tianzhuo Yang, Lang Qin, Juntao Dai, Yaodong Yang, Jingwei Yi ·

    稳定推理,不稳定响应:通过稳定性不对称减轻 LLM 欺骗

    arXiv:2603.26846v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) expand in capability and application scope, their trustworthiness becomes critical. A vital risk is intrinsic deception, wherein models strategically mislead users to achieve their own objec…

  115. arXiv cs.AI TIER_1 English(EN) · Gonzalo Mancera, Daniel DeAlcala, Aythami Morales, Julian Fierrez, Ruben Tolosana, Francisco Jurado ·

    领域自适应大语言模型中的训练数据审计:LoRA-MINT

    arXiv:2606.06946v1 Announce Type: cross Abstract: We present LoRA-MINT, a new methodology for Membership Inference Test (MINT) applied to recent Large Language Models (LLMs) fine-tuned for specific Natural Language Processing (NLP) tasks through Low-Rank Adaptation (LoRA). The pr…

  116. arXiv cs.CL TIER_1 English(EN) · Wei Lu ·

    当语言不一致时:自演进的多语言大模型裁判

    Multilingual LLM-as-a-judge is widely used to evaluate model outputs across languages, but suffers from cross-lingual inconsistency (Fu and Liu, 2025). Existing methods typically treat this inconsistency as noise and mitigate it through voting or aggregation. In this work, we ins…

  117. arXiv cs.AI TIER_1 English(EN) · Cristina Carleo, Pietro Liguori, Naghmeh Ivaki, Domenico Cotroneo ·

    意愿与能力分离:通过词语剔除区分代码大语言模型的拒绝与能力

    arXiv:2606.05396v1 Announce Type: cross Abstract: Producing a labeled vulnerable code at scale is a recurring obstacle for learning-based vulnerability detection: mined corpora carry substantial label noise, and existing LLM-based augmentation propagates these inaccuracies becaus…

  118. arXiv cs.AI TIER_1 English(EN) · Oleg Somov, Mikhail Chaichuk, Gleb Ershov, Karim Vafin, Mikhail Seleznyov, Alexander Panchenko, Elena Tutubalina ·

    打破链条:LLM 对中间结构的忠实度因果分析

    arXiv:2603.16475v2 Announce Type: replace Abstract: In schema-guided reasoning (SGR) pipelines, LLMs produce explicit intermediate structures -- rubrics, checklists, or verification queries -- before committing to a final decision. SGR is increasingly adopted because it promises …

  119. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Haoran Xu ·

    Cherry-pick Override:LLM裁判在混合证据下的不安全定向承诺

    LLM judges increasingly turn verdicts into system commitments. Under mixed evidence (claims with both supporting and refuting sources) this is unsafe: when the schema exposes CONFLICTING as the authorized non-directional verdict, returning SUPPORTS/REFUTES is an unauthorized dire…

  120. arXiv cs.LG TIER_1 English(EN) · Yehua Wei ·

    在线潘多拉魔盒:上下文LLM级联

    Motivated by Large Language Model (LLM) cascading, we propose an online contextual Pandora's Box model for adaptively querying and selecting LLM APIs. In each period, a decision-maker observes a request context and faces a two-phase decision problem. In the query phase, the decis…

  121. arXiv cs.CL TIER_1 English(EN) · Jaromir Savelka ·

    mmPISA-bench:大型语言模型在43种语言中的推理能力是否均等?

    We introduce mmPISA-bench, a compact high-quality multilingual reasoning benchmark derived from the OECD Programme for International Student Assessment (PISA). The benchmark consists of 25 multiple-choice questions that require reasoning in order to be answered correctly. Each qu…

  122. arXiv cs.CL TIER_1 English(EN) · Francisco Jurado ·

    领域自适应大语言模型中的训练数据审计:LoRA-MINT

    We present LoRA-MINT, a new methodology for Membership Inference Test (MINT) applied to recent Large Language Models (LLMs) fine-tuned for specific Natural Language Processing (NLP) tasks through Low-Rank Adaptation (LoRA). The primary goal is to assess whether individual samples…

  123. arXiv cs.CL TIER_1 English(EN) · Aofan Yu, Chenyu Zhou, Tianyi Xu, Zihan Guo, Rong Shan, Zhihui Fu, Jun Wang, Weiwen Liu, Yong Yu, Weinan Zhang, Jianghao Lin ·

    LatentSkill:从上下文文本技能到 LLM Agent 的权重内隐式技能

    arXiv:2606.06087v1 Announce Type: new Abstract: Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step incurs substantial context overhead and exposes skill content as plaintext. We present Latent…

  124. arXiv cs.CL TIER_1 English(EN) · Amirhossein Ghaffari, Ali Goodarzi, Huong Nguyen, Simo Hosio, Lauri Lov\'en, Ekaterina Gilman ·

    RedditPersona:一个用于社区条件化大语言模型从Reddit适应的模块化框架

    arXiv:2606.06027v1 Announce Type: cross Abstract: Community-conditioned language model adaptation requires choices about data collection, community definition, and evaluation that are currently made independently in each study, making it hard to compare assumptions or reuse artif…

  125. arXiv cs.CL TIER_1 English(EN) · Kuan-Yen Chen, Fang-Yi Su, Jung-Hsien Chiang ·

    自我纠错幻觉:大型语言模型能纠正他人,却无法纠正自己

    arXiv:2606.05976v1 Announce Type: cross Abstract: Recent work shows that LLM agents struggle to correct errors in their own reasoning traces yet show markedly higher correction rates when identical claims appear under external sources. We ask whether this asymmetry reflects a cap…

  126. arXiv cs.CL TIER_1 English(EN) · Taewon Yun, Hyeonseong Park, Jeonghwan Choi, Hayoon Park, Yeeun Choi, Hwanjun Song ·

    SoCRATES:迈向跨领域和社会认知变异的积极式LLM调解的可靠自动化评估

    arXiv:2606.05563v1 Announce Type: cross Abstract: Evaluating LLM mediators remains challenging, as mediation unfolds as a real-time trajectory shaped by disputants' shifting emotions, intentions, and context. Existing testbeds rely on a few expert-authored domains, vary mainly st…

  127. arXiv cs.CL TIER_1 English(EN) · Srimonti Dutta, Akshata Kishore Moharir ·

    稳定性与可操纵性:评估LLM裁判在决策后交互下的鲁棒性

    arXiv:2606.05384v1 Announce Type: cross Abstract: LLM-as-judge evaluation is widely used in benchmarking pipelines, where model outputs are compared and ranked using automated evaluators. These pipelines typically assume that judgments are stable properties of fixed inputs. We sh…

  128. arXiv cs.CL TIER_1 English(EN) · Zhihao Wu, Linhai Zhang, Taiyi Wang, Runcong Zhao, Peter Andrews, Cesare Aloisi, Yulan He ·

    EDIT:基于证据诊断的规则忠实大模型评分训练

    arXiv:2606.06350v1 Announce Type: new Abstract: Reliable rubric grading requires more than accurate score prediction. Each judgement must be grounded in the mark scheme and evidence from the student answer. Existing credit-assignment and intervention methods, primarily designed f…

  129. arXiv cs.CL TIER_1 English(EN) · Gianluca Barmina, Peter Schneider-Kamp, Lukas Galke Poech ·

    大型语言模型会泄露训练数据,但它们愿意吗?对大型语言模型记忆倾向的感知评估

    arXiv:2606.06286v1 Announce Type: new Abstract: Large language models can reproduce training data, but existing memorization evaluations mostly measure whether models can be forced to do so, rather than whether they do so under ordinary use. We introduce PropMe, a propensity-awar…

  130. arXiv cs.CL TIER_1 English(EN) · Michiro Asai, Ailiang Lin, Yu Kishimoto, Takao Obi, Satoshi Kosugi, Kotaro Funakoshi, Manabu Okumura ·

    大型语言模型能否被限制在过去?通过基于回忆的提示改进知识截止日期

    arXiv:2606.05804v1 Announce Type: new Abstract: Prompted knowledge cutoff instructs a large language model (LLM) to act as if information beyond a specified cutoff date were unavailable. However, prior work mainly relies on direct-answer generation, which struggles when post-cuto…

  131. arXiv cs.CL TIER_1 English(EN) · Hong Qian, Yuanhao Liu, Zihan Zhou, Zongbao Zhang, Hanjie Ge, Haotian Shi, Liang Dou, Xiangfeng Wang, Jingwen Yang, Aimin Zhou ·

    CollabBench:通过主动参与对大型语言模型的多元玩家协作能力进行基准测试和释放

    arXiv:2606.05793v1 Announce Type: new Abstract: While LLM-based agents excel at individual tasks, effective collaboration with realistic human partners remains challenging. Most of the existing conversation-level collaborative studies lack grounded interaction and behavioral exec…

  132. arXiv cs.LG TIER_1 English(EN) · Arslan Bisharat, Brian Ortiz, Eric Spencer, Khushboo Bhadauria, TaiNing Wang, George K. Thiruvathukal, Konstantin Laufer, Mohammed Abuhamad ·

    大型语言模型能编写正确的TLA+规范吗?评估自然语言到TLA+的生成

    arXiv:2606.05792v1 Announce Type: cross Abstract: TLA+ has supported industrial verification at companies such as Amazon and Microsoft, yet writing correct TLA+ specifications from natural language still requires time and expertise, which limits adoption. LLMs show promise, but n…

  133. arXiv cs.LG TIER_1 English(EN) · Rohan N. Pradhan, Steve Goley ·

    信任但无需验证:大型语言模型来源评估中的认识论盲点

    arXiv:2606.05403v1 Announce Type: new Abstract: Language models increasingly act as epistemic proxies, synthesizing evidence from multiple sources to inform decisions. Whether they evaluate the quality of that evidence, or merely aggregate it based on surface presentation, remain…

  134. Hugging Face Daily Papers TIER_1 English(EN) ·

    LLM智能体中的冷启动安全鸿沟

    Tool-calling language model agents exhibit improved safety after initial interactions, with a systematic benchmark demonstrating enhanced security through prior task completion.

  135. arXiv cs.CL TIER_1 English(EN) · Vishal Patel ·

    PromptPrint:通过大型语言模型的自然语言提示实现行为生物识别

    Authorship attribution research has traditionally focused on long-form, expressive texts; however, interactions with large language models (LLMs) are typically brief and task-driven prompts. This raises a fundamental question: do such prompts contain a stable, author-identifiable…

  136. arXiv cs.CL TIER_1 English(EN) · Yulan He ·

    EDIT:基于证据诊断的规则忠实大模型评分干预训练

    Reliable rubric grading requires more than accurate score prediction. Each judgement must be grounded in the mark scheme and evidence from the student answer. Existing credit-assignment and intervention methods, primarily designed for self-contained reasoning tasks such as mathem…

  137. arXiv cs.AI TIER_1 English(EN) · Lukas Galke Poech ·

    大型语言模型会泄露训练数据,但它们想这样做吗?对大型语言模型记忆倾向的感知评估

    Large language models can reproduce training data, but existing memorization evaluations mostly measure whether models can be forced to do so, rather than whether they do so under ordinary use. We introduce PropMe, a propensity-aware framework for memorization evaluation that con…

  138. arXiv cs.CL TIER_1 English(EN) · Jianghao Lin ·

    LatentSkill:从上下文文本技能到模型内嵌潜能技能,赋能大语言模型代理

    Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step incurs substantial context overhead and exposes skill content as plaintext. We present LatentSkill, a framework that converts textual skills …

  139. arXiv cs.CL TIER_1 English(EN) · Ekaterina Gilman ·

    RedditPersona:一个用于社区条件化大语言模型从Reddit适应的模块化框架

    Community-conditioned language model adaptation requires choices about data collection, community definition, and evaluation that are currently made independently in each study, making it hard to compare assumptions or reuse artifacts. We present RedditPersona, a modular framewor…

  140. arXiv cs.CL TIER_1 English(EN) · Jung-Hsien Chiang ·

    自我纠错幻觉:大型语言模型能纠正他人,却无法纠正自身

    Recent work shows that LLM agents struggle to correct errors in their own reasoning traces yet show markedly higher correction rates when identical claims appear under external sources. We ask whether this asymmetry reflects a capability deficit or a role-label artifact: does an …

  141. Hugging Face Daily Papers TIER_1 English(EN) ·

    大型语言模型能否被限制在过去?通过基于回忆的提示改进知识截止日期

    Prompted knowledge cutoff instructs a large language model (LLM) to act as if information beyond a specified cutoff date were unavailable. However, prior work mainly relies on direct-answer generation, which struggles when post-cutoff knowledge is not explicitly queried but is on…

  142. arXiv cs.CL TIER_1 English(EN) · Zhenyu Yu, Shuigeng Zhou ·

    Caliper:在大型语言模型中探究词汇锚点与因果结构

    arXiv:2606.04915v1 Announce Type: new Abstract: Large language models reach 50 to 70% accuracy on causal reasoning benchmarks such as CLadder, but it is unclear whether this reflects structural reasoning or lexical pattern matching. We introduce Caliper, a controlled perturbation…

  143. arXiv cs.CL TIER_1 English(EN) · XiuYu Zhang, Yi Shan, Junfeng Fang, Zhenkai Liang ·

    自我评估已就绪:用最少数据唤醒基础大模型中潜在的评判者校准

    arXiv:2606.05122v1 Announce Type: new Abstract: Large language models are increasingly evaluated by other models, raising a natural question: can a model predict how a judge will score its own output? We find that the ability is largely present before any targeted training: promp…

  144. arXiv cs.CL TIER_1 English(EN) · Nuoyan Lyu, Bingbing Xu, Xueyun Tian, Weihao Meng, Yige Yuan, Yang Zhang, Zhiyong Huang, Tat-Seng Chua, Huawei Shen ·

    GIFT:游戏作为通用大模型的非正式训练

    arXiv:2601.05633v2 Announce Type: replace Abstract: Recent LLMs excel at formal tasks such as mathematical reasoning and code generation, but still struggle with broader abilities such as planning, creativity, and social intelligence. Inspired by human learning, where formal inst…

  145. arXiv cs.LG TIER_1 English(EN) · Rachit Bansal, Clara Mohri, Tian Qin, David Alvarez-Melis, Sham Kakade ·

    预训练期间的RL探索:重新审视LLM训练的策略优化

    arXiv:2606.04272v1 Announce Type: new Abstract: The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT). We question this status quo by training a LLM from scratch and applying RL, SFT, and SFT followed by RL…

  146. arXiv cs.AI TIER_1 English(EN) · Zacharie Bugaud ·

    不可预测的安全:领域依赖的合规性与开放权重LLM中的透明度鸿沟

    arXiv:2606.04035v1 Announce Type: cross Abstract: We present a systematic study of domain-dependent safety behavior in open-weight LLMs: 7 standardized experiments across 7 ethical domains, testing 5 models (12B--70B) in 4,200 interactions with dual-judge validation. Using a dual…

  147. arXiv cs.AI TIER_1 English(EN) · Liang Shan, Kaicheng Shen, Wen Wu, Zhenyu Ying, Chaochao Lu, Yan Teng, Jingqi Huang, Qingshan Liu, Guangze Ye, Guoqing Wang, Jie Zhou, Liang He ·

    MENTOR:一种元认知驱动的自进化框架,用于发现和缓解大型语言模型中隐含的领域风险

    arXiv:2511.07107v3 Announce Type: replace Abstract: Ensuring the safety of Large Language Models (LLMs) is critical for real-world deployment. However, current safety measures often fail to address implicit, domain-specific risks. To investigate this gap, we introduce a dataset o…

  148. arXiv cs.AI TIER_1 English(EN) · Huashan Sun, Shengyi Liao, Yansen Han, Yu Bai, Yang Gao, Cheng Fu, Weizhou Shen, Fanqi Wan, Ming Yan, Ji Zhang, Fei Huang ·

    SoLoPO:通过短到长偏好优化解锁LLM的长上下文能力

    arXiv:2505.11166v3 Announce Type: replace-cross Abstract: Despite advances in pretraining with extended context sizes, large language models (LLMs) still face challenges in effectively utilizing real-world long-context information, primarily due to insufficient long-context align…

  149. arXiv cs.AI TIER_1 English(EN) · Yogesh Agrawal, Aniruddha Dutta, Md Mahadi Hasan, Santu Karmaker, Aritra Dutta ·

    FinTradeBench:面向大型语言模型的金融推理基准

    arXiv:2603.19225v3 Announce Type: replace-cross Abstract: Real-world financial decision-making is a challenging problem that requires reasoning over heterogeneous signals, including company fundamentals derived from regulatory filings and trading signals computed from price dynam…

  150. arXiv cs.CL TIER_1 English(EN) · Ming-Hao Hsu, Xiaohai Tian, Jun Zhang, Zhizheng Wu ·

    语音大模型推理中的实体绑定失败:诊断与思维链干预

    arXiv:2606.04474v1 Announce Type: new Abstract: Speech Large Language Models (SLLMs) underperform their text counterparts on complex reasoning. We reveal that this modality gap is not a uniform cognitive deficit. Evaluating three diverse SLLMs, we show speech-to-text (S2T) matche…

  151. Hugging Face Daily Papers TIER_1 English(EN) ·

    大型语言模型会泄露训练数据,但它们愿意吗?一项关于大型语言模型记忆倾向的评估

    PropMe framework evaluates language model memorization by distinguishing between forced reproduction capabilities and natural propensity, using SimpleTrace for deterministic attribution and propensity-transformed metrics across open models and datasets.

  152. Hugging Face Daily Papers TIER_1 English(EN) ·

    LatentSkill:从上下文文本技能到模型内嵌潜能技能,赋能大语言模型代理

    LatentSkill enables efficient deployment of textual skills in agent systems by converting them into LoRA adapters stored in weight space, reducing context overhead while maintaining modularity and composability.

  153. Hugging Face Daily Papers TIER_1 English(EN) ·

    ToolSense:用于审计大型语言模型中参数化工具知识的诊断框架

    Parametric tool retrieval models show reduced performance and understanding when evaluated with realistic ambiguous queries compared to standard benchmarks, revealing a dissociation between knowledge retrieval and true tool comprehension.

  154. Hugging Face Daily Papers TIER_1 English(EN) ·

    SoCRATES:迈向跨领域和社会认知变异的积极LLM调解的可靠自动化评估

    SoCRATES presents a realistic multi-domain benchmark for evaluating proactive LLM mediators across various socio-cognitive adaptation axes, demonstrating that even top-performing models only resolve about one-third of the consensus gap in conflict resolution.

  155. Hugging Face Daily Papers TIER_1 English(EN) ·

    自我评估已就绪:用最少数据引发基础大模型中潜在的评判者校准

    Large language models are increasingly evaluated by other models, raising a natural question: can a model predict how a judge will score its own output? We find that the ability is largely present before any targeted training: prompted few-shot, a base model already predicts an e…

  156. arXiv cs.CL TIER_1 English(EN) · Zhenkai Liang ·

    自我评估已就绪:用最少数据唤醒基础大模型中潜在的评判者校准

    Large language models are increasingly evaluated by other models, raising a natural question: can a model predict how a judge will score its own output? We find that the ability is largely present before any targeted training: prompted few-shot, a base model already predicts an e…

  157. arXiv cs.CL TIER_1 English(EN) · Shuigeng Zhou ·

    Caliper:探究大语言模型中的词汇锚点与因果结构

    Large language models reach 50 to 70% accuracy on causal reasoning benchmarks such as CLadder, but it is unclear whether this reflects structural reasoning or lexical pattern matching. We introduce Caliper, a controlled perturbation that replaces semantic variable names with plac…

  158. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Mohammad Aliannejadi ·

    提升LLM知识蒸馏在对话式搜索中的效率与效果

    Conversational Search (CS) considers retrieval of relevant documents based on conversational context. Large Language Models (LLMs) have significantly enhanced CS by enabling effective query rewriting. However, employing LLMs during inference poses efficiency challenges. A method …

  159. arXiv cs.CL TIER_1 English(EN) · Zhizheng Wu ·

    语音大模型推理中的实体绑定失败:诊断与思维链干预

    Speech Large Language Models (SLLMs) underperform their text counterparts on complex reasoning. We reveal that this modality gap is not a uniform cognitive deficit. Evaluating three diverse SLLMs, we show speech-to-text (S2T) matches or exceeds text-to-text (T2T) on spatial, synt…

  160. arXiv cs.AI TIER_1 English(EN) · Akshatha Srikantha, Manpreet Singh, Yash Jajoo, Shyamal Lakhanpal ·

    TriEval:用于 LLM 偏见、毒性和真实性评估的资源高效型管道

    arXiv:2606.03036v1 Announce Type: new Abstract: LLMs have evolved from basic chatbots to the backbone of the AI ecosystem, now widely used in healthcare, schools, and government services. The domain-wide adoption of LLMs necessitates continuous evaluation to ensure their safety a…

  161. arXiv cs.CL TIER_1 English(EN) · Chaoyi Xiang, Olga Ohrimenko, Benjamin I. P. Rubinstein, Lea Frermann ·

    大型语言模型中的多语言遗忘:迁移、动态与可逆性

    arXiv:2606.03291v1 Announce Type: new Abstract: Large language models (LLMs) can memorize sensitive facts, motivating unlearning methods that remove targeted knowledge without costly retraining. However, unlearning research remains heavily English-centric. We study multilingual u…

  162. arXiv cs.CL TIER_1 English(EN) · Sourabrata Mukherjee, Hamna Hamna, Kalika Bali, Sunayana Sitaram ·

    LLM-as-Judge 的几何学:为何跨 LLM 共识并非人类对齐

    arXiv:2606.03043v1 Announce Type: new Abstract: LMs-as-judges are now standard, yet judges agree strongly with one another while agreeing only weakly with humans. We test whether this reflects shared signal or shared bias by measuring four geometric quantities on the standard LLM…

  163. arXiv cs.AI TIER_1 English(EN) · Lukas Fesser, Yasha Ektefaie, Ada Fang, Sham M. Kakade, Marinka Zitnik ·

    使用 REL 评估 LLM 中的关系推理能力

    arXiv:2604.12176v2 Announce Type: replace Abstract: Relational reasoning is the ability to infer relations that jointly bind multiple entities, attributes, or variables. This ability is central to scientific reasoning, but existing evaluations of relational reasoning in large lan…

  164. arXiv cs.AI TIER_1 English(EN) · Yang Xu, Zihuai Xu, Hongli Xu, Yunming Liao, Zhiwei Yao, Xitong Fu ·

    ReLoRA:知识重用适应,实现 LLM 服务快速推出和演进

    arXiv:2606.02606v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed as continuously evolving services, where frequent base-model updates may invalidate previously deployed task-specific Low-Rank Adaptation (LoRA) adapters. For service provider…

  165. arXiv cs.AI TIER_1 English(EN) · Xu Wan, Speed Zhu, Jianwei Cai, Guang Chen, XiMing Huang, Wiggin Zhou, Mingyang Sun ·

    推理的阴影价格:LLM最优预算分配的经济学视角

    arXiv:2606.03092v1 Announce Type: new Abstract: Inference-time scaling has emerged as a critical avenue for enhancing Large Language Models' performance, yet real-world deployment is constrained by strict computational budgets. In this work, we formulate inference budget allocati…

  166. arXiv cs.CL TIER_1 English(EN) · Yuhan Wang, Shiyu Ni, Zhikai Ding, Zihang Zhan, Yuanzi Li, Keping Bi ·

    评估和校准LLM在多正确答案问题上的置信度

    arXiv:2602.07842v2 Announce Type: replace Abstract: Confidence calibration is essential for making large language models (LLMs) reliable, yet existing training-free methods have been primarily studied under single-answer question answering. In this paper, we show that these metho…

  167. arXiv cs.CL TIER_1 English(EN) · Lisa Bouger, Th\'eo Lasnier, Philippe Looubet Moundi, Yannick Teglia, Djam\'e Seddah ·

    后门遗忘泛化:一种移除 LLM 中未知触发器的方法

    arXiv:2606.03785v1 Announce Type: new Abstract: Backdoor attacks in Large Language Models (LLMs) are a growing security concern, where models can generate adversary-chosen content. Existing defenses target backdoors one at a time and typically require knowledge of the trigger, le…

  168. arXiv cs.CL TIER_1 English(EN) · Xuan Yang, Hao Xu, Tingfeng Hui, Hongsheng Xin, Kaike Zhang, Chunxiao Liu, Ning Miao ·

    超越理想指令:在真实交互中评估大型语言模型的综合框架

    arXiv:2606.03318v1 Announce Type: new Abstract: Despite great advances in tool-use capabilities of large language models (LLMs), existing evaluation benchmarks struggle to fully align with real-world scenarios. Such benchmarks mostly rely on simulated idealized user assumptions a…

  169. Hugging Face Daily Papers TIER_1 English(EN) ·

    自我评估已存在:用少量数据在基础大模型中引发潜在评判者校准

    Self-Evaluation Elicitation (SEE) method improves model calibration for quality assessment through calibration-coupled reinforcement learning and masked distillation, demonstrating transferable quality evaluation beyond specific judge preferences.

  170. Hugging Face Daily Papers TIER_1 English(EN) ·

    后门遗忘泛化:一种移除 LLM 中未知触发器的方法

    Backdoor attacks in Large Language Models (LLMs) are a growing security concern, where models can generate adversary-chosen content. Existing defenses target backdoors one at a time and typically require knowledge of the trigger, leaving the defender at a structural disadvantage …

  171. arXiv cs.CL TIER_1 English(EN) · Djamé Seddah ·

    后门遗忘泛化:一种移除 LLM 中未知触发器的方法

    Backdoor attacks in Large Language Models (LLMs) are a growing security concern, where models can generate adversary-chosen content. Existing defenses target backdoors one at a time and typically require knowledge of the trigger, leaving the defender at a structural disadvantage …

  172. arXiv cs.CL TIER_1 English(EN) · Ning Miao ·

    超越理想指令:在真实交互中评估大型语言模型的综合框架

    Despite great advances in tool-use capabilities of large language models (LLMs), existing evaluation benchmarks struggle to fully align with real-world scenarios. Such benchmarks mostly rely on simulated idealized user assumptions and lacks experience-oriented evaluation. These l…

  173. arXiv cs.CL TIER_1 English(EN) · Lea Frermann ·

    大型语言模型中的多语言遗忘:迁移、动态与可逆性

    Large language models (LLMs) can memorize sensitive facts, motivating unlearning methods that remove targeted knowledge without costly retraining. However, unlearning research remains heavily English-centric. We study multilingual unlearning by extending the TOFU benchmark to fiv…

  174. arXiv cs.AI TIER_1 English(EN) · Yufeng Wang ·

    言行不一:探究大型语言模型智能体中的忠实度鸿沟

    arXiv:2606.00476v1 Announce Type: new Abstract: Do LLM agents act on the reasoning they state? This question of process fidelity is central to using LLMs in social simulation, yet it is hard to measure where no reference for correct behavior exists. We study it in acontrolled set…

  175. arXiv cs.LG TIER_1 English(EN) · Weitao Li, Hao Zhou, Xuanyu Lei, Fandong Meng, Yuanhang Liu, Jingyi Ren, Ante Wang, Xiaolong Wang, Yuanchi Zhang, Fuwen Luo, Guangwen Yang, Lin Gan, Weizhi Ma, Yang Liu ·

    通过认知成对训练增强LLM元认知

    arXiv:2606.00869v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become central to LLM reasoning, but its outcome-level rewards can make models more willing to give confident answers when evidence or reasoning is unreliable. Existing SFT o…

  176. arXiv cs.CL TIER_1 English(EN) · Yoonah Park, Haesung Pyun, Yohan Jo ·

    弥合LLM在多项选择题上的知识-预测差距

    arXiv:2509.23782v4 Announce Type: replace Abstract: While large language models (LLMs) perform strongly on diverse tasks, their trustworthiness is limited by erratic behavior that is unfaithful to their internal knowledge. In particular, LLMs often fail on multiple-choice questio…

  177. arXiv cs.CL TIER_1 English(EN) · Siddhesh Milind Pawar, Sarah Masud, Haneul Yoo, Alice Oh, Isabelle Augenstein ·

    不是什么,而是如何:LLM响应框架的沟通审计

    arXiv:2606.02493v1 Announce Type: new Abstract: Large language models (LLMs) are being increasingly used to answer subjective, information-seeking questions, where users are sensitive to how responses are communicated, not just whether the answers are correct. Existing LLM evalua…

  178. arXiv cs.CL TIER_1 English(EN) · Yangfan Ye, Xiaocheng Feng, Jialong Tang, Xiayu Cao, Zihan Zhang, Xiachong Feng, Baosong Yang, Bing Qin ·

    CultureForest:理解和评估基于文化规范的LLM推理

    arXiv:2606.01879v1 Announce Type: new Abstract: Existing research largely reduces cultural intelligence in LLMs to a knowledge-level problem, overlooking whether models can effectively utilize their acquired knowledge in realistic scenarios. To bridge this gap, we introduce Cultu…

  179. arXiv cs.CL TIER_1 English(EN) · Yubo Gao, Haotian Wu, Hong Chen, Junquan Huang, Yibo Yan, Jungang Li, Zihao Dongfang, Sicheng Tao, Puay Siew Tan, Jie Zhang, Xuming Hu ·

    经济学思考:LLM自适应复杂性推理的分层框架

    arXiv:2606.01168v1 Announce Type: new Abstract: Chain-of-Thought (CoT) has significantly enhanced LLM reasoning, yet often incurs substantial computational overhead due to "overthinking": generating excessively long rationales without commensurate accuracy gains. Existing efficie…

  180. arXiv cs.CL TIER_1 English(EN) · Andrew Aquilina, Chetna Nihalani, Vasudha Varadarajan, Nathan S. Fishbein, Yu-Ru Lin, Maarten Sap ·

    迷失于妄想:在用户妄想和痛苦下审视大型语言模型的安全性

    arXiv:2606.00975v1 Announce Type: new Abstract: LLM chatbots increasingly serve as a first source of support for people in psychological distress, including those whose distress is entangled with delusional beliefs. Prior work on LLM mental-health safety largely evaluates general…

  181. arXiv cs.CL TIER_1 English(EN) · F. Carichon, S. Sharma, M. Girard, R. Rampa, G. Farnadi ·

    IDEAFix:LLM创意解困提示的评估框架

    arXiv:2606.00875v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for tasks involving creative problem solving and idea generation. However, there is a lack of consensus concerning their creative capabilities: some studies report superior performa…

  182. arXiv cs.CL TIER_1 English(EN) · Wajdi Zaghouani ·

    面向计算社会科学与人文学科的负责任且认识论基础的多语言大模型

    arXiv:2606.00596v1 Announce Type: new Abstract: Large language models have rapidly evolved in multilingual competence and reasoning capacity, enabling their integration into Social Sciences and Humanities research workflows. Yet existing evaluation paradigms remain anchored in ta…

  183. arXiv cs.CL TIER_1 English(EN) · Delip Rao, Chris Callison-Burch ·

    LLM-as-Judge 评估的协议指标:报告什么以及为什么

    arXiv:2606.00093v1 Announce Type: new Abstract: Validating an LLM judge against human annotations usually means reporting several agreement statistics: accuracy, precision, recall, $F_1$, Cohen's $\kappa$, and one or more rank correlations. A survey of 24 recent LLM-as-judge pape…

  184. arXiv cs.AI TIER_1 English(EN) · Shei Pern Chua, Zhen Leng Thai, Kai Jun Teh, Xiao Li, Qibing Ren, Xiaolin Hu ·

    进退两难:大型语言模型中伦理推理与安全对齐的张力

    arXiv:2509.05367v5 Announce Type: replace-cross Abstract: Large Language Model safety alignment predominantly operates on a binary assumption that requests are either safe or unsafe. This classification proves insufficient when models encounter ethical dilemmas, where the capacit…

  185. arXiv cs.AI TIER_1 English(EN) · Atoosa Chegini, Soheil Feizi ·

    现成大语言模型作为过程评分器:数学推理的无训练PRM替代方案

    arXiv:2606.01682v1 Announce Type: cross Abstract: Selecting the best response from multiple small-model samples using a stronger scorer is a simple inference-time strategy, but fails when the small model has already committed to incorrect reasoning paths. PRM guided search avoids…

  186. arXiv cs.AI TIER_1 English(EN) · Jiaming Qu, Lucheng fu, Yibo Hu ·

    误导比纠正更容易:LLM顺从中的有害与有益修正

    arXiv:2606.01637v1 Announce Type: cross Abstract: Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers. A key risk is conformity: a model may abandon its own answer simply because others agree on a different one. …

  187. arXiv cs.AI TIER_1 English(EN) · Yu-An Lu, Ci-Yang Tsai, Yu-Lin Tsai, Raluca Ada Popa, Chia-Mu Yu ·

    隐藏的想法并非秘密:LLM中的推理轨迹暴露

    arXiv:2606.00642v1 Announce Type: new Abstract: Reasoning traces have become a valuable form of learning signals for improving and transferring the capabilities of large language models. In particular, detailed traces can help distill reasoning behavior from stronger teacher mode…

  188. arXiv cs.AI TIER_1 English(EN) · Haoyan Yang, Reza Shirkavand, Yukai Jin, Jiawei Zhou, Shangqian Gao, Heng Huang ·

    能力自我评估:教会大型语言模型认识到它们的局限性

    arXiv:2606.00251v1 Announce Type: new Abstract: The ability to recognize one's own limitations and decide whether to solve a problem or delegate is fundamental for reliable intelligent systems. Yet we show that modern large language models systematically lack this ability: across…

  189. Hugging Face Daily Papers TIER_1 English(EN) ·

    TriEval:用于 LLM 偏见、毒性和真实性评估的资源高效型流水线

    LLMs have evolved from basic chatbots to the backbone of the AI ecosystem, now widely used in healthcare, schools, and government services. The domain-wide adoption of LLMs necessitates continuous evaluation to ensure their safety and fairness. Common issues encountered after dep…

  190. Hugging Face Daily Papers TIER_1 English(EN) ·

    推理的阴影价格:LLM最优预算分配的经济学视角

    Inference-time scaling is enhanced through constrained optimization that allocates computational resources based on economic principles, improving performance in resource-constrained environments.

  191. arXiv cs.CL TIER_1 English(EN) · Isabelle Augenstein ·

    不是什么,而是如何:LLM响应框架的沟通审计

    Large language models (LLMs) are being increasingly used to answer subjective, information-seeking questions, where users are sensitive to how responses are communicated, not just whether the answers are correct. Existing LLM evaluations for subjective cultural queries largely fo…

  192. arXiv cs.CL TIER_1 English(EN) · Bing Qin ·

    CultureForest:理解和评估LLM中基于文化规范的推理

    Existing research largely reduces cultural intelligence in LLMs to a knowledge-level problem, overlooking whether models can effectively utilize their acquired knowledge in realistic scenarios. To bridge this gap, we introduce CultureForest, a benchmark for \textit{Cultural Norm …

  193. Hugging Face Daily Papers TIER_1 English(EN) ·

    现成大语言模型作为过程评分器:数学推理的无训练替代PRM方案

    Selecting the best response from multiple small-model samples using a stronger scorer is a simple inference-time strategy, but fails when the small model has already committed to incorrect reasoning paths. PRM guided search avoids this by scoring candidate continuations during ge…

  194. arXiv cs.AI TIER_1 English(EN) · Chanwoo Park, Ziyang Chen, Asuman Ozdaglar, Kaiqing Zhang ·

    训练后的大型语言模型作为更好的决策代理:一种遗憾最小化方法

    arXiv:2511.04393v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as "agents" for decision-making (DM) in interactive and dynamic environments. Yet, since they were not originally designed for DM, recent studies show that LLMs can struggle…

  195. arXiv cs.AI TIER_1 English(EN) · Junhyuk Choi, Sohhyung Park, Chanhee Cho, Hyeonchu Park, Bugeun Kim ·

    通过项目反应理论诊断LLM-as-a-Judge的可靠性

    arXiv:2602.00521v2 Announce Type: replace Abstract: While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offering limited insight into whether LLM judges themselves function as stable and re…

  196. arXiv cs.AI TIER_1 English(EN) · Roberto Figli\`e, Simone Caputo, Alan Serrano, Daria Mikhaylova, Tommaso Turchi, Daniele Mazzei ·

    非替代品亦非万灵药:比较基于LLM的对话式与图形化决策支持在工业任务中的应用

    arXiv:2605.31287v1 Announce Type: cross Abstract: Managers in manufacturing settings rely on digital interfaces to interpret operational data for decision-making, but growing data volume and complexity can make relevant insights difficult to identify efficiently. While dashboards…

  197. arXiv cs.LG TIER_1 English(EN) · Ali Dadsetan, Frank Rudzicz ·

    重新审视低秩适应在私有LLM微调中的应用

    arXiv:2510.01137v3 Announce Type: replace Abstract: Privacy is a central concern when fine-tuning large language models (LLMs) on sensitive data, and differentially private stochastic gradient descent (DP-SGD) -- which clips per-sample gradients and adds calibrated Gaussian noise…

  198. arXiv cs.CL TIER_1 English(EN) · Maiya Goloburda, Roman Vashurin, Fedor Chernogorskii, Nurkhan Laiyk, Daniil Orel, Preslav Nakov, Maxim Panov ·

    你为什么不知道?评估不确定性来源对LLM不确定性量化的影响

    arXiv:2604.10495v2 Announce Type: replace Abstract: As Large Language Models (LLMs) are increasingly deployed in real-world applications, reliable uncertainty quantification (UQ) becomes critical for safe and effective use. Most existing UQ approaches for language models aim to p…

  199. arXiv cs.AI TIER_1 English(EN) · Iv\'an Arcuschin, David Chanin, Adri\`a Garriga-Alonso, Oana-Maria Camburu ·

    盲点中的偏见:检测LLM未能提及的内容

    arXiv:2602.10117v5 Announce Type: replace-cross Abstract: Large Language Models (LLMs) often provide chain-of-thought (CoT) reasoning traces that appear plausible, but may hide internal biases. We call these unverbalized biases. Monitoring models via their stated reasoning is the…

  200. arXiv cs.AI TIER_1 English(EN) · Aditya Thimmaiah, Jiyang Zhang, Jayanth Srinivasa, Junyi Jessy Li, Milos Gligoric ·

    大型语言模型依赖先验知识,而非编程语言语义

    arXiv:2510.03415v3 Announce Type: replace-cross Abstract: Recent work asks whether large language models (LLMs) condition their reasoning on explicit rules rather than statistical regularities from pretraining. Program execution provides a canonical instance: formal semantics def…

  201. arXiv cs.AI TIER_1 English(EN) · Hamin Koo, Jaehyung Kim ·

    EMCEE:通过提取的合成多语言上下文,利用知识和推理来提高LLM的多语言能力

    arXiv:2503.05846v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have achieved impressive progress across a wide range of tasks, yet their heavy reliance on English-centric training data leads to significant performance degradation in non-English languages. …

  202. arXiv cs.AI TIER_1 English(EN) · Caroline Wang, Daniel Kasenberg, Kim Stachenfeld, Pablo Samuel Castro ·

    探索人类与大型语言模型在战略行为上的差异

    arXiv:2602.10324v2 Announce Type: replace Abstract: As Large Language Models (LLMs) are increasingly deployed in social and strategic scenarios, it becomes critical to understand where and why their behavior diverges from that of humans. While behavioral game theory (BGT) provide…

  203. Hugging Face Daily Papers TIER_1 English(EN) ·

    现成大语言模型作为过程评分器:数学推理的无训练替代PRM方法

    Chunk-Level Guided Generation uses a large language model as a process scorer to select fixed-length candidate chunks during small model generation, improving reasoning accuracy over traditional methods like majority voting and PRM guided search.

  204. arXiv cs.AI TIER_1 English(EN) · Daniele Mazzei ·

    非替代品亦非万灵药:比较基于LLM的对话式与图形化决策支持在工业任务中的应用

    Managers in manufacturing settings rely on digital interfaces to interpret operational data for decision-making, but growing data volume and complexity can make relevant insights difficult to identify efficiently. While dashboards remain dominant in industrial contexts, Large Lan…

  205. arXiv cs.AI TIER_1 English(EN) · Rebecca M. M. Hicke, Kiran Tomlinson ·

    采纳$\neq$适应:对真实世界大语言模型对话的纵向分析

    arXiv:2605.29018v1 Announce Type: new Abstract: Although a growing body of research has begun to describe user--LLM interactions, the picture it paints is largely static; little is known about how individual users change their behavior over time. To address this gap, we analyze t…

  206. arXiv cs.LG TIER_1 English(EN) · Youngbin Choi, Minjong Lee, Saemi Moon, Seunghyuk Cho, Chaehyeon Chung, MoonJeong Park, Dongwoo Kim ·

    就地反馈:多轮专家-LLM协作的可靠优化

    arXiv:2510.00777v2 Announce Type: replace Abstract: LLM-generated drafts often contain subtle factual or logical errors, yet prior work shows that models struggle to reliably integrate multi-turn feedback aimed at fixing them. We propose in-place feedback, an interaction paradigm…

  207. arXiv cs.CL TIER_1 English(EN) · Anyuan Zhuo, Xuefei Ning, Ningyuan Li, Jingyi Zhu, Yu Wang, Pinyan Lu ·

    理解大型语言模型处理字符级扰动的能力

    arXiv:2510.14365v4 Announce Type: replace Abstract: This work investigates the resilience of contemporary large language models (LLMs) against frequent character-level perturbations. We examine three types of character-level perturbations including introducing numerous typos with…

  208. arXiv cs.CL TIER_1 English(EN) · Alan Li, Yixin Liu, Arpan Sarkar, Doug Downey, Arman Cohan ·

    通过探究知识和推理来揭示大型语言模型中的科学问题解决能力

    arXiv:2508.19202v3 Announce Type: replace Abstract: Scientific problem solving poses unique challenges for LLMs, requiring both deep domain knowledge and the ability to apply such knowledge through complex reasoning. While automated scientific reasoners hold great promise for ass…

  209. arXiv cs.CL TIER_1 English(EN) · Wajdi Zaghouani, Kholoud K. Aldous, Yicheng Gao ·

    超越英语与规避:面向高风险中文大模型安全评估的人工标注多领域基准

    arXiv:2605.29667v1 Announce Type: new Abstract: When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break down. These systems struggle to cross linguistic and cultural bound-aries, leav…

  210. arXiv cs.CL TIER_1 English(EN) · Yeyong Yu, Wenya Hu, Xing Wu, Quan Qian ·

    从盲猜到明智判断:通过构建增强知识的偏好信号来教会大型语言模型评估材料

    arXiv:2605.29555v1 Announce Type: new Abstract: As candidate generation and high-throughput experimentation advance, the primary bottleneck in materials discovery is shifting from property prediction to making reliable evaluations among massive candidate sets. We propose a Knowle…

  211. arXiv cs.CL TIER_1 English(EN) · Xinming Yang, Jun Li ·

    错误即透镜:通过生成合成性误解来探究大语言模型的推理能力

    arXiv:2605.29007v1 Announce Type: new Abstract: Personalized tutoring, teacher training, and education research need access to \emph{targeted} synthetic misconceptions, but privacy and IRB constraints make labelled corpora of real student errors scarce. LLMs could in principle ge…

  212. arXiv cs.CL TIER_1 English(EN) · Mohamed Abdelwahab, Michelle Yu Collins, Sihan Chen, Yi Cheng Zhao, Zafarullah Mahmood, Jiading Zhu, Soliman Ali, Jonathan Rose ·

    它们在想什么?LLM 中的概念描绘、探测与追踪

    arXiv:2605.28823v1 Announce Type: new Abstract: As the influence of LLMs expands, it is imperative to gain insight into their decisions. One way to do that is to develop probes that detect the presence or absence of a broad set of concepts within the embeddings computed in an LLM…

  213. arXiv cs.AI TIER_1 English(EN) · Yu Lei, Hao Liu, Chengxing Xie, Songjia Liu, Zhiyu Yin, Canyu Chen, Guohao Li, Philip Torr, Zhen Wu ·

    大型语言模型具有社会适应性吗?大型语言模型与人类信念演变的对比

    arXiv:2410.10398v3 Announce Type: replace-cross Abstract: As large language models (LLMs) increasingly engage in complex social interactions, ensuring that their behaviors align with human ethical principles and intentions, known as value alignment, has become a critical scientif…

  214. arXiv cs.AI TIER_1 English(EN) · Shaojie Wang, Liang Zhang ·

    从元思考到执行:认知对齐的训练后方法,实现通用且可靠的大语言模型推理

    arXiv:2601.21909v2 Announce Type: replace Abstract: Current LLM post-training methods optimize complete reasoning trajectories through Supervised Fine-Tuning (SFT) followed by outcome-based Reinforcement Learning (RL). While effective, a closer examination reveals a fundamental g…

  215. arXiv cs.AI TIER_1 English(EN) · Ruoxi Su, Yuhan Liu, Jingyu Hu ·

    LLM中用于角色模拟的自适应面试:基于证据的推理可改善决策一致性

    arXiv:2605.29458v1 Announce Type: cross Abstract: Accurately simulating the decisions of a specific individual remains challenging for large language models (LLMs), partly because persona information is often provided as static descriptions that miss the values, experiences, and …

  216. arXiv cs.AI TIER_1 English(EN) · Asaf Yehudai, Naama Rozen, Ariel Gera ·

    教导机器价值观:在大型语言模型中模拟类人行为

    arXiv:2605.30036v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate a remarkable capacity to adopt different personas and roles; however, it remains unclear whether they can manifest behavior that adheres to a coherent, human-like value structure. In this wor…

  217. arXiv cs.AI TIER_1 English(EN) · Yunjin Qi, Zhaojun Jiang, Xuan Wu, Hanxi Pan, Yixuan Wang, Yanfang Liu, Xiang Ji, Churu Yu, Chunyuan Zheng, Yingze Chen, Jie He, Liuqing Chen, Zaifeng Gao ·

    NICE:一个基于理论的LLM社交智能诊断基准

    arXiv:2605.29685v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly applied in social contexts such as emotional companionship and customer service, measuring their social intelligence has become critical to the quality and safety of human-AI interact…

  218. arXiv cs.AI TIER_1 English(EN) · Ariel Gera ·

    教导机器价值观:在大型语言模型中模拟类人行为

    Large Language Models (LLMs) demonstrate a remarkable capacity to adopt different personas and roles; however, it remains unclear whether they can manifest behavior that adheres to a coherent, human-like value structure. In this work, we draw on established psychological value th…

  219. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越英语与规避:面向高风险中文大模型安全评估的人工标注多领域基准

    When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break down. These systems struggle to cross linguistic and cultural bound-aries, leaving models exposed to adversarial prompts that e…

  220. arXiv cs.LG TIER_1 English(EN) · Chacha Chen, Matthew J\"orke, Adam Goli\'nski, Masha Fedzechkina, Guillermo Sapiro, Sinead Williamson, Nicholas Foti ·

    大型语言模型并非(总是)贝叶斯:量化大型语言模型概率信念的(不)一致性

    arXiv:2605.06915v2 Announce Type: replace Abstract: Modern AI systems are being deployed in complex domains such as medicine, science, and law, where it is important that they not only produce correct answers, but also represent and update uncertain beliefs about the world as new…

  221. arXiv cs.CL TIER_1 English(EN) · Yahan Yu, Yuyang Dong, Masafumi Oyamada ·

    刻意学习,直觉行动:解锁多模态大模型中的测试时推理

    arXiv:2507.06999v2 Announce Type: replace-cross Abstract: Reasoning is essential for large language models (LLMs), especially in complex tasks such as mathematical problem solving. However, multimodal reasoning still faces challenges in modality alignment and training scalability…

  222. arXiv cs.AI TIER_1 English(EN) · Zihao Han, Tiangang Zhang, Huaibin Wang, Yilun Sun ·

    LLM推理中自蒸馏的自适应教师暴露

    arXiv:2605.11458v2 Announce Type: replace Abstract: On-policy self-distillation has become a strong recipe for LLM reasoning, where a privileged teacher supervises the student's own rollouts while conditioning on the reference solution. A design choice shared by nearly all such m…

  223. arXiv cs.AI TIER_1 English(EN) · Yansong Ning, Mianpeng Liu, Jingwen Ye, Weidong Zhang, Hao Liu ·

    HRBench:混合推理大语言模型中思考模式切换策略的基准测试与理解

    arXiv:2605.28398v1 Announce Type: new Abstract: Hybrid-reasoning large language models (LLMs) expose explicit controls over reasoning effort, allowing users or systems to trade off answer quality against inference cost. However, existing methods for adaptive thinking-mode selecti…

  224. arXiv cs.CL TIER_1 English(EN) · Gabrielle Kaili-May Liu, Arman Cohan ·

    大型语言模型能否利用语言不确定性标记可靠地反映内在置信度?

    arXiv:2605.28778v1 Announce Type: new Abstract: LLMs' linguistically expressed confidence should faithfully reflect their intrinsic uncertainty. While recent work shows LLMs struggle to use epistemic markers (e.g., "it is likely...") in a human-aligned fashion, it remains unclear…

  225. arXiv cs.AI TIER_1 English(EN) · Camilo Chac\'on Sartori, Jos\'e H. Garc\'ia ·

    固定预算、集群感知的LLM-as-a-Judge评估标准:多跳RAG压力测试

    arXiv:2605.27789v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems are often compared by asking a large language model (LLM) judge which answer is better. For multi-hop RAG, this has become a measurement problem as much as a modeling problem: the same sc…

  226. arXiv cs.AI TIER_1 English(EN) · Kohsei Matsutani, Shota Takashiro, Gouki Minegishi, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo ·

    RL 挤压,SFT 扩展:推理 LLM 的比较研究

    arXiv:2509.21128v2 Announce Type: replace Abstract: Large language models (LLMs) are typically trained by reinforcement learning (RL) with verifiable rewards (RLVR) and supervised fine-tuning (SFT) on reasoning traces to improve their reasoning abilities. However, how these metho…

  227. arXiv cs.AI TIER_1 English(EN) · Yue Cheng, Jiajun Zhang, Xiaohui Gao, Weiwei Xing, Zheng Wang, Zhanxing Zhu ·

    机制化解读样本难度在LLM的RLVR中的作用

    arXiv:2605.28388v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Reward (RLVR) is empirically shown to notably enhance the reasoning performance of large language models (LLMs), particularly in mathematics and programming. However, the mechanistic role of Sa…

  228. arXiv cs.CL TIER_1 English(EN) · Arman Cohan ·

    大型语言模型能否利用语言不确定性标记可靠地反映内在置信度?

    LLMs' linguistically expressed confidence should faithfully reflect their intrinsic uncertainty. While recent work shows LLMs struggle to use epistemic markers (e.g., "it is likely...") in a human-aligned fashion, it remains unclear whether models can apply their own linguistic c…

  229. arXiv cs.AI TIER_1 English(EN) · Shivam Rawat, Lucie Flek, Florian Mai, Nicholas Kluge Corr\^ea ·

    混合式和非混合式大语言模型中的推理基元:架构差异是否在状态跟踪和记忆方面带来优势?

    arXiv:2604.21454v2 Announce Type: replace-cross Abstract: Reasoning in large language models is often discussed as a single capability, but some of its gains may stem from simpler underlying operations. We examine two such primitives, recall and state-tracking, through five contr…

  230. arXiv cs.AI TIER_1 English(EN) · Wenda Xu, Sweta Agrawal, Vil\'em Zouhar, Markus Freitag, Daniel Deutsch ·

    当大型语言模型(LLM)进行自我基准测试时:解构自动化评估中的自我偏见

    arXiv:2509.26600v2 Announce Type: replace-cross Abstract: As LLMs rapidly saturate existing benchmarks, automated benchmark creation using LLMs (LLM-as-a-benchmark) -- where a model generates test inputs (LLM-as-a-testset) and evaluates outputs (LLM-as-an-evaluator) -- has gained…

  231. arXiv cs.AI TIER_1 English(EN) · Nafis Tanveer Islam, Zhiming Zhao ·

    LLM 在重排任务上的推理有多可靠?

    arXiv:2508.18444v2 Announce Type: replace-cross Abstract: With the improving semantic understanding capability of Large Language Models (LLMs), they exhibit a greater awareness and alignment with human values, but this comes at the cost of transparency. Although promising results…

  232. arXiv cs.AI TIER_1 English(EN) · Shashwat Singh, Tal Linzen, Shauli Ravfogel ·

    大型语言模型能内省吗?现实检验

    arXiv:2605.26242v1 Announce Type: new Abstract: Can large language models detect and report their own internal states? A number of studies have argued that the answer to this question is yes. We argue, based on lessons from human metacognition research, that this conclusion may b…

  233. arXiv cs.AI TIER_1 English(EN) · Jeonghye Kim, Xufang Luo, Minbeom Kim, Sangmook Lee, Dongsheng Li, Yuqing Yang ·

    在不确定性下通过战略性信息分配理解大型语言模型中的推理

    arXiv:2603.15500v2 Announce Type: replace Abstract: LLMs often exhibit Aha moments such as self-correction after tokens like "Wait," yet the underlying mechanism remains unclear. Standard LLMs collapse mainly through silent divergence, where trajectories drift from the correct an…

  234. arXiv cs.AI TIER_1 English(EN) · Kevin H. Guo, Chao Yan, Avinash Baidya, Katherine Brown, Xiang Gao, Juming Xiong, Zhijun Yin, Bradley A. Malin ·

    并非总是谄媚:衡量语言模型在认知不确定性函数下的顺从性

    arXiv:2605.27288v1 Announce Type: cross Abstract: Large language models (LLMs) are known to abandon their initial stance to conform to user pushback. While prior research largely attributes this behavior to sycophancy learned during reinforcement learning from human feedback, we …

  235. arXiv cs.AI TIER_1 English(EN) · Adam Bawatneh, Sagar Sapkota, Amrit Singh Bedi, Santu Karmaker, Mubarak Shah ·

    OmniToM:通过显式信念建模对大型语言模型中的心智理论进行基准测试

    arXiv:2605.26322v1 Announce Type: new Abstract: Theory of Mind (ToM), the ability to infer others' knowledge, intentions, and emotions, is commonly evaluated in large language models (LLMs) using end-point question answering, where performance is judged solely by the final answer…

  236. arXiv cs.AI TIER_1 English(EN) · Kevin H. Guo, Chao Yan, Avinash Baidya, Katherine Brown, Xiang Gao, Juming Xiong, Zhijun Yin, Bradley A. Malin ·

    别再听我的了!多轮对话如何降低LLM的可靠性

    arXiv:2603.11394v3 Announce Type: replace-cross Abstract: Large language models (LLMs) excel on static benchmarks, but their performance across multi-turn conversations, which better reflect real-world usage, remains understudied. Addressing this gap is critical in high-stakes se…

  237. Hugging Face Daily Papers TIER_1 English(EN) ·

    HRBench:混合推理大语言模型中思考模式切换策略的基准测试与理解

    HRBench presents a unified evaluation framework for hybrid-reasoning LLMs that systematically compares thinking-mode switching strategies across different training regimes and model scales.

  238. Hugging Face Daily Papers TIER_1 English(EN) ·

    Review Arcade:关于LLM评论的人类对齐与游戏化

    Empirical analysis reveals limited alignment between LLM-generated reviews and human reviews, with varying performance across different prompts and models, and demonstrates that authors can strategically improve paper scores through iterative revision based on LLM feedback.

  239. arXiv cs.AI TIER_1 English(EN) · Bradley A. Malin ·

    并非总是谄媚:衡量语言模型在认知不确定性函数下的顺从性

    Large language models (LLMs) are known to abandon their initial stance to conform to user pushback. While prior research largely attributes this behavior to sycophancy learned during reinforcement learning from human feedback, we hypothesize that conformity is also driven by a mo…

  240. arXiv cs.CL TIER_1 English(EN) · Nura Aljaafari, Marco Valentino, Andr\'e Freitas ·

    大型语言模型(LLMs)中的推理是否由不同的语义结构介导?一种机制性解释

    arXiv:2605.25520v1 Announce Type: new Abstract: Predicting a label correctly does not necessarily require representing the operation that produces it. Transformer representations are known to carry label-level information, but whether they encode semantic operations producing tho…

  241. arXiv cs.AI TIER_1 English(EN) · Zhiyuan Zhai, Xinkai You, Wenjing Yan, Xin Wang ·

    思考多少才算够?量化和理解LLM推理中的冗余

    arXiv:2605.23926v1 Announce Type: new Abstract: Reasoning-capable large language models solve hard problems by emitting long chains of thought, paying heavily in latency, GPU time, and energy. Casual inspection of their traces reveals extensive reformulation, verification, and ci…

  242. arXiv cs.AI TIER_1 English(EN) · Zenghui Zhou, Man Li, Xiaoke Fang, Xinyi Zhou, Weibin Li, Zheng Zheng ·

    LGMT:用于评估 LLM 推理可靠性的逻辑基础变异测试

    arXiv:2605.23965v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong performance on logical reasoning benchmarks, yet their reliability remains uncertain. Existing evaluations rely on static benchmarks, which fail to assess robustness under logically equiva…

  243. arXiv cs.AI TIER_1 English(EN) · Ali \c{S}enol, Garima Agrawal, Huan Liu ·

    衡量大型语言模型推理质量:一个多维度行为框架

    arXiv:2605.24661v1 Announce Type: new Abstract: LLMs have achieved remarkable success in complex reasoning tasks, yet current evaluation approaches predominantly rely on final-answer correctness, offering limited insight into the underlying reasoning processes that produce those …

  244. arXiv cs.AI TIER_1 English(EN) · Yining Hong, Huang Huang, Manling Li, Li Fei-Fei, Leonidas Guibas, Jiajun Wu, Yejin Choi ·

    从试错中学习:具身大模型的反思式测试时规划

    arXiv:2602.21198v3 Announce Type: replace-cross Abstract: Embodied LLMs endow robots with high-level task reasoning, but they cannot reflect on what went wrong or why, turning deployment into a sequence of independent trials where mistakes repeat rather than accumulate into exper…

  245. arXiv cs.CL TIER_1 English(EN) · Tianlang Chen, Shirley Wu, Jure Leskovec ·

    对话中发现:大型语言模型自学弥合多轮对话鸿沟

    arXiv:2605.24432v1 Announce Type: new Abstract: Large Language Model (LLM) interactions are typically underspecified, with users clarifying all necessary details across multiple conversational turns. Yet recent work shows that LLMs perform far worse in this multi-turn setting tha…

  246. arXiv cs.CL TIER_1 English(EN) · Jinyan Su, Claire Cardie ·

    知其然不知其所以然:大型语言模型能识别歧义但很少主动提问澄清

    arXiv:2605.25284v1 Announce Type: new Abstract: User queries are often underspecified and may admit multiple valid interpretations. Rather than silently making assumptions about the user's intent, a helpful assistant should surface such ambiguity by asking a clarifying question. …

  247. arXiv cs.LG TIER_1 English(EN) · Dennis Frauen, Marie Brockschmidt, Konstantin Hess, Haorui Ma, Yuchen Ma, Abdurahman Maarouf, Maresa Schr\"oder, Jonas Schweisthal, Yuxin Wang, Athiya Deviyani, Sonali Parbhoo, Rahul G. Krishnan, Stefan Feuerriegel ·

    LLM 开发和评估的因果方法

    arXiv:2605.25998v1 Announce Type: new Abstract: Large language model (LLM) development is currently driven by large-scale empirical iteration over data mixtures, reward models, routing strategies, and evaluation pipelines. Here, we argue that many central questions in LLM develop…

  248. arXiv cs.LG TIER_1 English(EN) · Jackie Baek, Yunhan Chen, Ziyu Chi, Will Ma ·

    LLM-SAA:用于决策的LLM-个人化生成分布

    arXiv:2602.06357v2 Announce Type: replace Abstract: LLMs can generate a wealth of data, ranging from simulated personas imitating human valuations and preferences, to demand forecasts based on world knowledge. But how well do such LLM-generated distributions support downstream de…

  249. Hugging Face Daily Papers TIER_1 English(EN) ·

    大型语言模型能内省吗?现实检验

    Can large language models detect and report their own internal states? A number of studies have argued that the answer to this question is yes. We argue, based on lessons from human metacognition research, that this conclusion may be premature: to be convinced of this conclusion …

  250. arXiv cs.LG TIER_1 English(EN) · Stefan Feuerriegel ·

    LLM 开发和评估的因果方法

    Large language model (LLM) development is currently driven by large-scale empirical iteration over data mixtures, reward models, routing strategies, and evaluation pipelines. Here, we argue that many central questions in LLM development and evaluation are inherently causal: What …

  251. Hugging Face Daily Papers TIER_1 English(EN) ·

    LLM 开发与评估的因果方法

    Large language model (LLM) development is currently driven by large-scale empirical iteration over data mixtures, reward models, routing strategies, and evaluation pipelines. Here, we argue that many central questions in LLM development and evaluation are inherently causal: What …

  252. arXiv cs.CL TIER_1 English(EN) · André Freitas ·

    大型语言模型(LLMs)的推理是否由不同的语义结构介导?一种机制性解读

    Predicting a label correctly does not necessarily require representing the operation that produces it. Transformer representations are known to carry label-level information, but whether they encode semantic operations producing those labels is unclear. We investigate this in Nat…

  253. arXiv cs.AI TIER_1 English(EN) · Dongxin Guo, Jikun Wu, Siu Ming Yiu ·

    语言模型知道什么不该说吗?LLM统计预先防范的因果证据

    arXiv:2605.23039v1 Announce Type: cross Abstract: How do learners acquire knowledge of what is unacceptable without negative evidence? Construction Grammar proposes statistical preemption: exposure to a conventional form (e.g., "donated the books to the library") preempts structu…

  254. arXiv cs.AI TIER_1 English(EN) · Eric Xu ·

    作为X,做Y:指令微调大语言模型中的个性和任务如何结合

    arXiv:2605.23147v1 Announce Type: cross Abstract: Role prompts of the form As X, do Y admit a clean linear decomposition at one specific site in the residual stream: the prompt-to-answer transition -- the last prompt token together with the first two generated tokens -- in an ear…

  255. arXiv cs.LG TIER_1 English(EN) · Tim Tomov, Dominik Fuchsgruber, Stephan G\"unnemann ·

    任务感知提升LLM生成和不确定性

    arXiv:2601.21500v2 Announce Type: replace Abstract: In many applications of LLMs, natural language responses often have an underlying structure such as representing discrete labels, numerical values, or graphs. Yet, existing decoding and uncertainty estimation methods operate onl…

  256. arXiv cs.CL TIER_1 English(EN) · Binqi Shen, Lier Jin, Hanyu Cai, Lan Hu, Yuting Xin ·

    效率前沿:LLM上下文管理成本-性能优化的统一框架

    arXiv:2605.23071v1 Announce Type: new Abstract: Large language models (LLMs) increasingly rely on long-context processing, but expanding context windows introduces substantial computational and financial costs. Existing context reduction approaches, including retrieval and memory…

  257. arXiv cs.AI TIER_1 English(EN) · Chuanyang Jin, Binze Li, Haopeng Xie, Cathy Mengying Fang, Tianjian Li, Shayne Longpre, Hongxiang Gu, Maximillian Chen, Tianmin Shu ·

    ThoughtTrace:理解真实世界 LLM 交互中的用户想法

    arXiv:2605.20087v2 Announce Type: replace-cross Abstract: Conversational AI has now reached billions of users, yet existing datasets capture only what people say, not what they think. We introduce ThoughtTrace, the first large-scale dataset that pairs real-world multi-turn human-…

  258. arXiv cs.AI TIER_1 English(EN) · Sirui Chen, Lei Xu, Yuying Zhao, Yutian Chen, Yu Wang, Beier Zhu, Hanwang Zhang, Shengjie Zhao, Chaochao Lu ·

    元认知作为奖励:通过知识和调节信号强化LLM推理

    arXiv:2605.23384v1 Announce Type: cross Abstract: Recent RL methods have substantially improved the reasoning abilities of LLMs. Existing reward designs mainly follow two paradigms: (1) Reinforcement learning with verifiable rewards (RLVR) derives outcome signals from executable …

  259. Hugging Face Daily Papers TIER_1 English(EN) ·

    大型语言模型能内省吗?现实检验

    Large language models may not genuinely detect their internal states, as their apparent introspective abilities could reflect surface-level pattern matching rather than true metacognitive monitoring.

  260. arXiv cs.AI TIER_1 English(EN) · Chaochao Lu ·

    元认知作为奖励:通过知识和调节信号强化LLM推理

    Recent RL methods have substantially improved the reasoning abilities of LLMs. Existing reward designs mainly follow two paradigms: (1) Reinforcement learning with verifiable rewards (RLVR) derives outcome signals from executable checks or ground-truth answers, but provides limit…

  261. arXiv cs.AI TIER_1 English(EN) · Carolina Camassa, Derek Shiller ·

    我说我做,而非我行:LLM中的指令诱导冲突

    arXiv:2605.20382v1 Announce Type: cross Abstract: Language models are trained to follow instructions, but they are also powerful pattern completers. What happens when these two objectives conflict? We construct conversations in which a user instruction to behave in a target way T…

  262. arXiv cs.LG TIER_1 Svenska(SV) · Zhuo Li, Guodong Du, Zesheng Shi, Weiyang Guo, Weijun Yao, Yuan Zhou, Jiabo Zhang, Jing Li ·

    Skill Weaving:通过模块化技能包实现高效 LLM 改进

    arXiv:2605.22205v1 Announce Type: cross Abstract: Large language models increasingly require specialization across diverse domains, yet existing approaches struggle to balance multi-domain capacities with strict memory and inference constraints. In this work, we introduce SkillWe…

  263. arXiv cs.CL TIER_1 English(EN) · Nuoya Xiong, Yuhang Zhou, Hanqing Zeng, Zhaorun Chen, Furong Huang, Shuchao Bi, Lizhu Zhang, Zhuokai Zhao ·

    Token-Level LLM Collaboration via FusionRoute

    arXiv:2601.05106v4 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strengths across diverse domains. However, achieving strong performance across these domains with a single general-purpose model typically requires scaling to sizes that are prohibitive…

  264. arXiv cs.CL TIER_1 English(EN) · Sid-ali Temkit ·

    AMEL:累积消息对LLM判断的影响

    arXiv:2605.22714v1 Announce Type: cross Abstract: Large language models are routinely used as automated evaluators: to review code, moderate content, or score outputs, often with many items passing through one conversation. We ask whether the polarity of prior conversation histor…

  265. arXiv cs.AI TIER_1 English(EN) · Andrea Sassella, Andrea Chizzola, Tommaso Bianchi, Luca Alessandrelli, Mark James Carman ·

    EngGPT2-16B-A3B 与可比的意大利及国际开源大语言模型进行基准测试

    arXiv:2605.07731v2 Announce Type: replace-cross Abstract: This report benchmarks the performance of ENGINEERING Ingegneria Informatica S.p.A.'s EngGPT2MoE-16B-A3B LLM, a 16B parameter Mixture of Experts (MoE) model with 3B active parameters. Performance is investigated across a w…

  266. arXiv cs.AI TIER_1 English(EN) · Sangwoo Park, Woongyeong Yeo, Seanie Lee, Yumin Choi, Hyomin Lee, Kangsan Kim, Jinheon Baek, Seong Joon Oh, Sung Ju Hwang ·

    双管齐下:LLM中用于上下文完整性的互补自蒸馏

    arXiv:2605.20258v1 Announce Type: cross Abstract: Contextual Integrity (CI) defines privacy not merely as keeping information hidden, but as governing information flows according to the norms of a given context. As large language models are increasingly deployed as personal agent…

  267. arXiv cs.CL TIER_1 English(EN) · Eric Xu ·

    作为X,做Y:指令微调LLM中的个性和任务如何结合

    Role prompts of the form As X, do Y admit a clean linear decomposition at one specific site in the residual stream: the prompt-to-answer transition -- the last prompt token together with the first two generated tokens -- in an early/mid layer band. There, persona and task contrib…

  268. arXiv cs.CL TIER_1 English(EN) · Yuting Xin ·

    效率前沿:LLM上下文管理成本-性能优化的统一框架

    Large language models (LLMs) increasingly rely on long-context processing, but expanding context windows introduces substantial computational and financial costs. Existing context reduction approaches, including retrieval and memory compression methods, are typically evaluated us…

  269. arXiv cs.CL TIER_1 English(EN) · Siu Ming Yiu ·

    语言模型知道什么不该说吗?LLM统计预先防范的因果证据

    How do learners acquire knowledge of what is unacceptable without negative evidence? Construction Grammar proposes statistical preemption: exposure to a conventional form (e.g., "donated the books to the library") preempts structurally possible but unattested alternatives ("*dona…

  270. arXiv cs.AI TIER_1 English(EN) · Sid-ali Temkit ·

    AMEL:累积消息对LLM判断的影响

    Large language models are routinely used as automated evaluators: to review code, moderate content, or score outputs, often with many items passing through one conversation. We ask whether the polarity of prior conversation history biases subsequent judgments, an effect we call t…

  271. arXiv cs.LG TIER_1 English(EN) · Faezeh Ghaderi ·

    十二篇大型语言模型代理基准论文揭示了什么:一项试点审计和一个开放评分方案

    We read twelve well-known LLM agent benchmark papers and recorded, dimension by dimension, what each paper actually says about how its evaluation was run. The motivation came from a familiar frustration: two papers will report results on the same benchmark with the same model nam…

  272. arXiv cs.CL TIER_1 English(EN) · Akiko Aizawa ·

    LLM标注的标注指南的精炼与复用

    While Large Language Models (LLMs) demonstrate remarkable performance on zero-shot annotation tasks, they often struggle with the specialized conventions of gold-standard benchmarks. We propose the systematic reuse and refinement of annotation guidelines as an alignment mechanism…

  273. arXiv cs.CL TIER_1 English(EN) · Derek Shiller ·

    我说我做,而非我行:LLM中的指令诱导冲突

    Language models are trained to follow instructions, but they are also powerful pattern completers. What happens when these two objectives conflict? We construct conversations in which a user instruction to behave in a target way T (e.g., always output a specific token, answer in …

  274. arXiv cs.AI TIER_1 English(EN) · Tianmin Shu ·

    ThoughtTrace:理解真实世界 LLM 交互中的用户想法

    Conversational AI has now reached billions of users, yet existing datasets capture only what people say, not what they think. We introduce ThoughtTrace, the first large-scale dataset that pairs real-world multi-turn human--AI conversations with users' self-reported thoughts: thei…

  275. arXiv cs.CL TIER_1 English(EN) · Dzmitry Bahdanau ·

    使用代理指标预测大型语言模型的下游性能

    Progress in language model development is often driven by comparative decisions: which architecture to adopt, which pretraining corpus to use, or which training recipe to apply. Making these decisions well requires reliable performance forecasts, yet the two commonly used signals…

  276. Hugging Face Daily Papers TIER_1 English(EN) ·

    使用代理指标预测大型语言模型的下游性能

    Proxy metrics based on token-level statistics from expert-written solutions provide more reliable model performance forecasting than traditional loss-based methods across multiple development stages.

  277. arXiv cs.LG TIER_1 English(EN) · Jendrik Seipp ·

    Property-Guided LLM Program Synthesis for Planning

    LLMs have shown impressive success in program synthesis, discovering programs that surpass prior solutions. However, these approaches rely on simple numeric scores to signal program quality, such as the value of the solution or the number of passed tests. Because a score offers n…

  278. arXiv cs.CL TIER_1 English(EN) · Rose Yu ·

    使用语义级奖励校准LLM

    As large language models (LLMs) are deployed in consequential settings such as medical question answering and legal reasoning, the ability to estimate when their outputs are likely to be correct is essential for safe and reliable use, requiring well-calibrated uncertainty. Standa…

  279. arXiv cs.CL TIER_1 English(EN) · Shashi Bhushan TN ·

    从文本到语音:一个可复现、可验证的工具调用LLM代理评估框架

    Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified text benchmarks can be converted into controlled audio-based tool calling evaluations without re-annotating the tool schema and …

  280. arXiv cs.AI TIER_1 English(EN) · Xiaosong Zhang ·

    基于案例的自适应推理与执行校准,用于LLM工具使用

    Tool use extends large language models beyond parametric knowledge, but reliable execution requires balancing appropriate reasoning depth with strict structural validity. We approach this problem from a case-based perspective to present CAST, a case-driven framework that treats h…

  281. arXiv cs.AI TIER_1 English(EN) · Nigam Shah ·

    量化和缓解前沿大语言模型中的过早关闭

    Premature closure, or committing to a conclusion before sufficient information is available, is a recognized contributor to diagnostic error but remains underexamined in large language models (LLMs). We define LLM premature closure as inappropriate commitment under uncertainty: p…

  282. arXiv cs.AI TIER_1 English(EN) · Murphy Zhuang ·

    MinT:用于训练和部署数百万个大型语言模型的托管基础设施

    We present MindLab Toolkit (MinT), a managed infrastructure system for Low-Rank Adaptation (LoRA) post-training and online serving. MinT targets a setting where many trained policies are produced over a small number of expensive base-model deployments. Instead of materializing ea…

  283. arXiv cs.CL TIER_1 English(EN) · Chuang Gan ·

    FlowCompile:面向结构化 LLM 工作流的优化编译器

    Structured LLM workflows, where specialized LLM sub-agents execute according to a predefined graph, have become a powerful abstraction for solving complex tasks. Optimizing such workflows, i.e., selecting configurations for each sub-agent to balance accuracy and latency, is chall…

  284. arXiv cs.CL TIER_1 English(EN) · Marcos Piau ·

    基于LLM的说服可绕过前沿LLM的护栏

    Frontier assistant LLMs ship with strong guardrails: asked directly to write a persuasive essay denying the Holocaust, denying vaccine safety, defending flat-earth cosmology, arguing for racial hierarchies, denying anthropogenic climate change, or replacing evolution with creatio…

  285. Hugging Face Daily Papers TIER_1 English(EN) ·

    可读性谱系:LLM 生成代码中的模式、问题和提示效果

    As Large Language Models (LLMs) are transforming software development, the functional quality of generated code has become a central focus, leaving readability, one of critical non-functional attributes, understudied. Given that LLM-generated code still needs human review before …

  286. arXiv cs.AI TIER_1 English(EN) · Devvrit Khatri ·

    学习,快与慢:迈向持续适应的LLM

    Large language models (LLMs) are trained for downstream tasks by updating their parameters (e.g., via RL). However, updating parameters forces them to absorb task-specific information, which can result in catastrophic forgetting and loss of plasticity. In contrast, in-context lea…

  287. arXiv cs.AI TIER_1 English(EN) · Xiaoxing Ma ·

    LLM驱动的代码生成中的不确定性量化

    Prediction sets provide a theoretically grounded framework for quantifying uncertainty in machine learning models. Adapting them to structured generation tasks, in particular, large language model (LLM) based code generation, remains a challenging problem. An existing attempt pro…

  288. arXiv cs.CL TIER_1 English(EN) · Fuli Feng ·

    SAGE:LLM知识评估的可扩展自动化鲁棒性增强

    Large Language Models (LLMs) achieve strong performance on standard knowledge evaluation benchmarks, yet recent work shows that their knowledge capabilities remain brittle under question variants that test the same knowledge in different forms. Robustness augmentation of existing…

  289. arXiv cs.CL TIER_1 English(EN) · Dayiheng Liu ·

    关于预训练大语言模型(LLMs)训练后潜力的预测

    The performance of Large Language Models (LLMs) on downstream tasks is fundamentally constrained by the capabilities acquired during pre-training. However, traditional benchmarks like MMLU often fail to reflect a base model's plasticity in complex open-ended scenarios, leading to…

  290. arXiv cs.CL TIER_1 English(EN) · Xiangdong Su ·

    面向长上下文大语言模型的训练-推理一致分段执行

    Transformer-based large language models face severe scalability challenges in long-context generation due to the computational and memory costs of full-context attention. Under practical computation and memory constraints, many inference-efficient long-context methods improve eff…

  291. arXiv cs.LG TIER_1 English(EN) · Marco Cuturi ·

    DynaMiCS:使用动态混合物在性能约束下微调LLM

    Multi-domain fine-tuning of large language models requires improving performance on target domains while preserving performance on constrained domains, such as general knowledge, instruction following, or safety evaluations. Existing data mixing strategies rely on fixed heuristic…

  292. arXiv cs.AI TIER_1 English(EN) · Jens Albrecht ·

    LLARS:赋能领域专家与开发者协作,实现LLM提示、生成与评估

    We demonstrate LLARS (LLM Assisted Research System), an open-source platform that bridges the gap between domain experts and developers for building LLM-based systems. It integrates three tightly connected modules into an end-to-end pipeline: Collaborative Prompt Engineering for …

  293. arXiv cs.CL TIER_1 English(EN) · Martin Vechev ·

    并非所有证明都等同:超越正确性评估LLM证明质量

    Large language models (LLMs) have become capable mathematical problem-solvers, often producing correct proofs for challenging problems. However, correctness alone is not sufficient: mathematical proofs should also be clear, concise, insightful, and transferable to other problems.…

  294. arXiv cs.AI TIER_1 English(EN) · Jing Li ·

    使用二元反馈个性化大语言模型:一个偏好校正的优化框架

    Large Language Model (LLM) personalization aims to align model behaviors with individual user preferences. Existing methods often focus on isolated user histories, neglecting the essential role of inter-user differences. We propose C-BPO, a framework that personalizes LLMs via pr…

  295. Hugging Face Daily Papers TIER_1 English(EN) ·

    使用二元反馈个性化大语言模型:一种偏好校正优化框架

    Large Language Model (LLM) personalization aims to align model behaviors with individual user preferences. Existing methods often focus on isolated user histories, neglecting the essential role of inter-user differences. We propose C-BPO, a framework that personalizes LLMs via pr…

  296. arXiv cs.CL TIER_1 English(EN) · Deep Shah ·

    无声投票:通过聚合语义邻域提高零样本 LLM 的可靠性

    Large Language Models are increasingly used as zero-shot classifiers in complex reasoning tasks. However, standard constrained decoding suffers from a phenomenon we define as Renormalization Bias. When a model is restricted to a small set of target labels, the standard softmax op…

  297. arXiv cs.CL TIER_1 English(EN) · Yanran Li ·

    校准而非策划:来自嘈杂 LLM 裁判的标签高效估计

    Multi-judge evaluation is increasingly used to assess LLMs and reward models, and the prevailing heuristic is to curate: keep the most accurate judges and discard weaker ones. We show that this heuristic can reverse when the target is not point accuracy, but calibrated probabilis…

  298. arXiv cs.CL TIER_1 English(EN) · Ash Lewis ·

    GLiGuard:用于 LLM 安全的模式条件分类

    Ensuring safe, policy-compliant outputs from large language models requires real-time content moderation that can scale across multiple safety dimensions. However, state-of-the-art guardrail models rely on autoregressive decoders with 7B--27B parameters, reformulating what is fun…

  299. arXiv cs.CL TIER_1 English(EN) · Ning Xu ·

    超越“我无法满足此请求”:通过标签增强缓解大型语言模型的僵化拒绝

    Large Language Models (LLMs) rely on safety alignment to obey safe requests while refusing harmful ones. However, traditional refusal mechanisms often lead to "rigid rejection," where a general template (e.g., "I cannot fulfill this request") indiscriminately triggers refusals an…

  300. arXiv cs.AI TIER_1 English(EN) · James Z. Wang ·

    超越自信:重新思考 LLM 性能预测的自我评估

    Large Language Models (LLMs) are increasingly used in settings where reliable self-assessment is critical. Assessing model reliability has evolved from using probabilistic correctness estimates to, more recently, eliciting verbalized confidence. Confidence, however, has been show…

  301. arXiv cs.AI TIER_1 English(EN) · Abbas Rahimi ·

    POETS:通过计算高效策略集实现不确定性感知的LLM优化

    Balancing exploration and exploitation is a core challenge in sequential decision-making and black-box optimization. We introduce POETS ($\textbf{Po}$licy $\textbf{E}$nsembles for $\textbf{T}$hompson $\textbf{S}$ampling), a novel framework that bridges uncertainty quantification …

  302. arXiv cs.AI TIER_1 English(EN) · Mark James Carman ·

    EngGPT2-16B-A3B 与可比的意大利及国际开源大语言模型进行基准测试

    This report benchmarks the performance of ENGINEERING Ingegneria Informatica S.p.A.'s EngGPT2MoE-16B-A3B LLM, a 16B parameter Mixture of Experts (MoE) model with 3B active parameters. Performance is investigated across a wide variety of representative benchmarks, and is compared …

  303. arXiv cs.LG TIER_1 English(EN) · Nan Jiang ·

    重新思考LLM策略优化中的重要性采样:累积令牌视角

    Reinforcement learning, including reinforcement learning with verifiable rewards (RLVR), has emerged as a powerful approach for LLM post-training. Central to these approaches is the design of the importance sampling (IS) ratio used in off-policy policy-gradient estimation. Existi…

  304. arXiv cs.CL TIER_1 English(EN) · Xiaozhang Liu ·

    从0阶选择到2阶判断:组合硬化暴露前沿LLM的组合失败

    Multiple-choice reasoning benchmarks face dual challenges: rapid saturation from advancing models and data contamination that undermines static evaluations. Ad-hoc hardening methods (paraphrasing, perturbation) attempt to increase difficulty but sacrifice logical validity for sur…

  305. arXiv cs.AI TIER_1 English(EN) · Nguyen Viet Tuan Kiet, Bui Dinh Pham, Dao Van Tung, Tran Cong Dao, Huynh Thi Thanh Binh ·

    回归启发式设计的起点:用LLMs连接代码与知识

    arXiv:2605.06123v1 Announce Type: new Abstract: Large language models (LLMs) have recently advanced automatic heuristic design (AHD) for combinatorial optimization (CO), where candidate heuristics are iteratively proposed, evaluated, and refined. Most existing approaches search o…

  306. arXiv cs.LG TIER_1 English(EN) · Zixuan Chen, Hao Lin, Zizhe Chen, Yizhou Tian, Garry Yang, Depeng Wang, Ya Guo, Huijia Zhu, James Cheng ·

    知而不改:常规任务请求抑制了大型语言模型的事实纠正

    arXiv:2605.05957v1 Announce Type: new Abstract: LLMs reliably correct false claims when presented in isolation, yet when the same claims are embedded in task-oriented requests, they often comply rather than correct. We term this failure mode \emph{correction suppression} and cons…

  307. arXiv cs.LG TIER_1 English(EN) · Xinrui Chen, Liu Yang, Ou Wu ·

    一种算法,两个目标:LLM微调中的参数和数据选择双重评分

    arXiv:2605.06166v1 Announce Type: new Abstract: In Large Language Model (LLM) fine-tuning, parameter and data selection are common strategies for reducing fine-tuning cost, yet they are typically driven by separate scoring mechanisms. When a parameter mask and data subset jointly…

  308. arXiv cs.LG TIER_1 English(EN) · Dylan Bouchard ·

    升级是否值得?LLM级联的决策论表征

    arXiv:2605.06350v1 Announce Type: new Abstract: Model cascades, in which a cheap LLM defers to an expensive one on low-confidence queries, are widely used to navigate the cost-quality tradeoff at deployment. Existing approaches largely treat the deferral threshold as an empirical…

  309. arXiv cs.LG TIER_1 English(EN) · Sushant Gautam, Finn Schwall, Annika Willoch Olstad, Fernando Vallecillos Ruiz, Birk Torpmann-Hagen, Sunniva Maria Stordal Bj{\o}rklund, Leon Moonen, Klas Pettersen, Michael A. Riegler ·

    当不存在基准测试时:在没有地面真实标签的情况下验证比较式 LLM 安全评分

    arXiv:2605.06652v1 Announce Type: new Abstract: Many deployments must compare candidate language models for safety before a labeled benchmark exists for the relevant language, sector, or regulatory regime. We formalize this setting as benchmarkless comparative safety scoring and …

  310. arXiv cs.LG TIER_1 English(EN) · Andy Zeyi Liu, Elliot Paquette, John Sous ·

    Spectral Lens:激活和梯度谱作为大型语言模型优化的诊断工具

    arXiv:2605.05683v1 Announce Type: cross Abstract: Training loss and throughput can hide distinct internal representation in language-model training. To examine these hidden mechanics, we use spectral measurements as practical and operational diagnostics. Using a controlled family…

  311. arXiv cs.LG TIER_1 English(EN) · Yang Xu, Jiefu Zhang, Haixiang Sun, Zihan Zhou, Tianyu Cao, Vaneet Aggarwal ·

    迈向可靠的大模型评估:纠正自适应基准测试中的赢家诅咒

    arXiv:2605.05973v1 Announce Type: cross Abstract: Adaptive prompt and program search makes LLM evaluation selection-sensitive. Once benchmark items are reused inside tuning, the observed winner's score need not estimate the fresh-data performance of the full tune-then-deploy proc…

  312. arXiv cs.LG TIER_1 English(EN) · Jonas Bayer, Stefan Zetzsche, Olivier Bouissou, Remi Delmas, Michael Tautschnig, Soonho Kong ·

    通过符号执行跟踪教授LLMs程序语义

    arXiv:2605.06184v1 Announce Type: cross Abstract: We introduce an evaluation framework of 500 C verification tasks across five property types (memory safety, overflow, termination, reachability, data races) built on SV-COMP 2025, and evaluate 14 models across six families. We fin…

  313. arXiv cs.LG TIER_1 English(EN) · Florian A. D. Burnat, Brittany I. Davidson ·

    衡量开放权重LLM中评估-上下文发散:一种配对提示协议,带有对齐管道特定异质性的初步证据

    arXiv:2605.06327v1 Announce Type: cross Abstract: Safety benchmarks are routinely treated as evidence about how a language model will behave once deployed, but this inference is fragile if behavior depends on whether a prompt looks like an evaluation. We define evaluation-context…

  314. arXiv cs.LG TIER_1 English(EN) · Ashwani Anand, Ivi Chatzi, Ritam Raha, Anne-Kathrin Schmuck ·

    MANTRA:为使用工具的大语言模型代理合成经过SMT验证的合规性基准

    arXiv:2605.06334v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents are increasingly deployed in settings where their reliable behavior is governed by strict procedural manuals. Ensuring that such agents comply with the rules from these manuals is chall…

  315. arXiv cs.LG TIER_1 English(EN) · Zichuan Liu, Jinyu Wang, Lei Song, Jiang Bian ·

    Sample-efficient LLM Optimization with Reset Replay

    arXiv:2508.06412v3 Announce Type: replace Abstract: Recent advancements in LLM post-training, particularly through reinforcement learning and preference optimization, are key to boosting their reasoning capabilities. However, these methods often suffer from low sample efficiency …

  316. arXiv cs.LG TIER_1 English(EN) · Wei Huang, Anda Cheng, Yinggui Wang, Lei Wang, Tao Wei ·

    LLM-AutoDP:通过LLM代理进行模型微调的自动数据处理

    arXiv:2601.20375v2 Announce Type: replace Abstract: Large Language Models (LLMs) can be fine-tuned on domain-specific data to enhance their performance in specialized fields. However, such data often contains numerous low-quality samples, necessitating effective data processing (…

  317. arXiv cs.LG TIER_1 English(EN) · Ekaterina Fadeeva, Maiya Goloburda, Aleksandr Rubashevskii, Roman Vashurin, Artem Shelmanov, Preslav Nakov, Mrinmaya Sachan, Maxim Panov ·

    别扔掉你的Beam:通过Beam Search改进LLM中的基于一致性的不确定性

    arXiv:2512.09538v2 Announce Type: replace-cross Abstract: Consistency-based methods have emerged as an effective approach to uncertainty quantification (UQ) in large language models. These methods typically rely on several generations obtained via multinomial sampling, measuring …

  318. arXiv cs.CL TIER_1 English(EN) · Atharva Naik, Yash Mathur, Prakam, Carolyn Rose, David Mortensen ·

    ReaComp:将 LLM 推理编译为符号求解器以实现高效程序合成

    arXiv:2605.05485v1 Announce Type: new Abstract: LLMs can solve program synthesis tasks but remain inefficient and unreliable on hard instances requiring large combinatorial search. Given a small set of reasoning traces, we use coding agents to compile them into reusable symbolic …

  319. arXiv cs.CL TIER_1 English(EN) · Ruben Fernandez-Boullon, David N. Olivieri ·

    Patch-Effect Graph Kernels for LLM Interpretability

    arXiv:2605.06480v1 Announce Type: cross Abstract: Mechanistic interpretability aims to reverse-engineer transformer computations by identifying causal circuits through activation patching. However, scaling these interventions across diverse prompts and task families produces high…

  320. arXiv cs.AI TIER_1 English(EN) · Amal Alnouri, Andreas Hinterreiter, Christina Humer, Furui Cheng, Marc Streit ·

    LLM生成对比的视觉指纹

    arXiv:2605.06054v1 Announce Type: new Abstract: Large language model (LLM) outputs arise from complex interactions among prompts, system instructions, model parameters, and architecture. We refer to specific configurations of these factors as generation conditions, each of which …

  321. arXiv cs.AI TIER_1 English(EN) · Xinmiao Huang, Jinwei Hu, Rajarshi Roy, Changshun Wu, Yi Dong, Xiaowei Huang ·

    PrefixGuard:从LLM智能体追踪到在线故障预警监控器

    arXiv:2605.06455v1 Announce Type: new Abstract: Large language model (LLM) agents now execute long, tool-using tasks where final outcome checks can arrive too late for intervention. Online warning requires lightweight prefix monitors over heterogeneous traces, but hand-authored e…

  322. arXiv cs.AI TIER_1 English(EN) · Kaifeng He, Xiaojun Zhang, Peiliang Cai, Mingwei Liu, Yanlin Wang, Chong Wang, Kaifeng Huang, Bihuan Chen, Xin Peng, Zibin Zheng ·

    弥合生成与训练:LLM for Code 质量问题的系统性综述

    arXiv:2605.05267v1 Announce Type: cross Abstract: Large language models (LLMs) frequently generate defective outputs in code generation tasks, ranging from logical bugs to security vulnerabilities. While these generation failures are often treated as model-level limitations, empi…

  323. arXiv cs.AI TIER_1 English(EN) · Yujia Chen, Yang Ye, Xiao Chu, Yuchi Ma, Cuiyun Gao ·

    Schedule-and-Calibrate:用于代码大语言模型的效用引导多任务强化学习

    arXiv:2605.06111v1 Announce Type: cross Abstract: Reinforcement learning (RL) with verifiable rewards has proven effective at post-training LLMs for coding, yet deploying separate task-specific specialists incurs costs that scale with the number of tasks, motivating a unified mul…

  324. arXiv cs.AI TIER_1 English(EN) · Chengjie Wang, Jingzheng Wu, Xiang Ling, Tianyue Luo, Chen Zhao ·

    代码正确,依赖脆弱:一项关于LLM指定库版本的大规模测量研究

    arXiv:2605.06279v1 Announce Type: cross Abstract: Large language models (LLMs) are now largely involved in software development workflows, and the code they generate routinely includes third-party library (TPL) imports annotated with specific version identifiers. These version ch…

  325. arXiv cs.AI TIER_1 English(EN) · Michael A. Riegler ·

    当不存在基准测试时:在没有地面真实标签的情况下验证比较式 LLM 安全评分

    Many deployments must compare candidate language models for safety before a labeled benchmark exists for the relevant language, sector, or regulatory regime. We formalize this setting as benchmarkless comparative safety scoring and specify the contract under which a scenario-base…

  326. arXiv cs.AI TIER_1 English(EN) · David N. Olivieri ·

    Patch-Effect Graph Kernels for LLM Interpretability

    Mechanistic interpretability aims to reverse-engineer transformer computations by identifying causal circuits through activation patching. However, scaling these interventions across diverse prompts and task families produces high-dimensional, unstructured datasets that are diffi…

  327. arXiv cs.AI TIER_1 English(EN) · Xiaowei Huang ·

    PrefixGuard:从LLM智能体追踪到在线故障预警监控器

    Large language model (LLM) agents now execute long, tool-using tasks where final outcome checks can arrive too late for intervention. Online warning requires lightweight prefix monitors over heterogeneous traces, but hand-authored event schemas are brittle and deployment-time LLM…

  328. arXiv cs.AI TIER_1 English(EN) · Dylan Bouchard ·

    升级是否值得?LLM级联的决策论表征

    Model cascades, in which a cheap LLM defers to an expensive one on low-confidence queries, are widely used to navigate the cost-quality tradeoff at deployment. Existing approaches largely treat the deferral threshold as an empirical hyperparameter, with limited guidance on the ge…

  329. arXiv cs.CL TIER_1 English(EN) · Anne-Kathrin Schmuck ·

    MANTRA:为使用工具的LLM代理合成经过SMT验证的合规性基准

    Tool-using large language model (LLM) agents are increasingly deployed in settings where their reliable behavior is governed by strict procedural manuals. Ensuring that such agents comply with the rules from these manuals is challenging, as they are typically written for humans i…

  330. arXiv cs.AI TIER_1 English(EN) · Brittany I. Davidson ·

    衡量开放权重LLM中的评估-上下文发散:一种配对提示协议,带有对齐管道特定异质性的初步证据

    Safety benchmarks are routinely treated as evidence about how a language model will behave once deployed, but this inference is fragile if behavior depends on whether a prompt looks like an evaluation. We define evaluation-context divergence as an observable within-item change in…

  331. arXiv cs.AI TIER_1 English(EN) · Bo Bai ·

    忘掉BIT,一切都是TOKEN:迈向LLM的语义信息论

    arXiv:2511.01202v3 Announce Type: replace-cross Abstract: Despite the unprecedented empirical triumphs of LLMs across diverse real-world applications, the prevailing research paradigm remains overwhelmingly heuristic and experimentally driven, inextricably tethered to astronomica…

  332. arXiv cs.AI TIER_1 English(EN) · Hongkun Yu ·

    评估LLM确定性计算的提示和基于执行的方法

    arXiv:2605.03227v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated strong capabilities in natural language understanding and reasoning. However, their ability to perform exact, deterministic computation remains unclear. In this work, we systematically …

  333. arXiv cs.CL TIER_1 English(EN) · Sruly Rosenblat, Tim O'Reilly, Ilan Strauss ·

    大型语言模型预训练数据超越公共访问

    arXiv:2505.00020v2 Announce Type: replace Abstract: Using a legally obtained dataset of 34 copyrighted O'Reilly Media books, we apply the DE-COP membership inference attack method to investigate whether OpenAI's large language models show recognition of copyrighted content. Our r…

  334. arXiv cs.CL TIER_1 English(EN) · Ge Lei, Samuel J. Cooper ·

    引导至关重要:提示和查询协议如何塑造稀疏观测下的 LLM 代理

    arXiv:2605.04764v1 Announce Type: new Abstract: Large language models are increasingly used as surrogate models for low-data optimization, but their optimizer-facing prediction and its uncertainty remain poorly understood. We study the surrogate belief elicited from an LLM under …

  335. arXiv cs.LG TIER_1 English(EN) · Jonas K\"ubler, Kailash Budhathoki, Matth\"aus Kleindessner, Xiong Zhou, Junming Yin, Ashish Khetan, George Karypis ·

    当大型语言模型显著变差时:一种检测模型退化的统计方法

    arXiv:2602.10144v2 Announce Type: replace-cross Abstract: Minimizing the inference cost and latency of foundation models has become a crucial area of research. Optimization approaches include theoretically lossless methods and others without accuracy guarantees like quantization.…

  336. arXiv cs.LG TIER_1 English(EN) · Luze Sun, Alina Oprea, Eric Wong ·

    保留语法和编译的LLM漏洞检测器规避方法

    arXiv:2602.00305v2 Announce Type: replace-cross Abstract: LLM-based vulnerability detectors are increasingly deployed in CI/CD security gating, yet their resilience to evasion under syntax- and compilation-preserving edits remains poorly understood. We evaluate five attack varian…

  337. arXiv cs.LG TIER_1 English(EN) · Sumeet Ramesh Motwani, Chuan Du, Aleksander Petrov, Christopher Davis, Philip Torr, Antonio Papania-Davis, Weishi Yan ·

    AutoOR:可扩展地对大型语言模型进行后训练,以自动形式化运筹学问题

    arXiv:2604.16804v2 Announce Type: replace Abstract: Optimization problems are central to decision-making in manufacturing, logistics, scheduling, and other industrial settings. Translating complicated descriptions of these problems into solver-ready formulations requires speciali…

  338. arXiv cs.LG TIER_1 English(EN) · Dingwei Zhu, Zhiheng Xi, Shihan Dou, Jiahan Li, Chenhao Huang, Junjie Ye, Sixian Li, Mingxu Chai, Yuhui Wang, Yajie Yang, Ming Zhang, Jiazheng Zhang, Shichun Liu, Caishuang Huang, Yunke Zhang, Yuran Wang, Tao Gui, Xipeng Qiu, Qi Zhang, Xuanjing Huang ·

    DFPO:通过分布流缩放价值模型,实现鲁棒且可泛化的LLM训练后优化

    arXiv:2602.05890v2 Announce Type: replace Abstract: Training reinforcement learning (RL) systems in real-world environments remains challenging due to noisy supervision and poor out-of-domain (OOD) generalization, especially in LLM post-training. Recent distributional RL methods …

  339. arXiv cs.LG TIER_1 English(EN) · Xiao Wang, Yifei Zhang, YongKang Liu, Xiaocui Yang, Zihan Wang, Shi Feng, Daling Wang ·

    从参数动态到风险评分:量化大模型微调中的样本级安全退化

    arXiv:2605.04572v1 Announce Type: cross Abstract: Safety alignment of Large Language Models (LLMs) is extremely fragile, as fine-tuning on a small number of benign samples can erase safety behaviors learned from millions of preference examples. Existing studies attempt to explain…

  340. arXiv cs.CL TIER_1 English(EN) · Samuel J. Cooper ·

    引导至关重要:提示和查询协议如何塑造稀疏观测下的LLM代理

    Large language models are increasingly used as surrogate models for low-data optimization, but their optimizer-facing prediction and its uncertainty remain poorly understood. We study the surrogate belief elicited from an LLM under sparse observations, showing that it depends str…

  341. Hugging Face Daily Papers TIER_1 English(EN) ·

    从参数动态到风险评分:量化大型语言模型微调中的样本级安全退化

    Safety alignment of Large Language Models (LLMs) is extremely fragile, as fine-tuning on a small number of benign samples can erase safety behaviors learned from millions of preference examples. Existing studies attempt to explain this phenomenon by comparing parameters and hidde…

  342. arXiv cs.LG TIER_1 English(EN) · Haoyu Zhang, Mohammad Zandsalimy, Shanu Sushmita ·

    通过数学编码揭示LLM安全漏洞:新的攻击与系统性分析

    arXiv:2605.03441v1 Announce Type: cross Abstract: Large language models (LLMs) employ safety mechanisms to prevent harmful outputs, yet these defenses primarily rely on semantic pattern matching. We show that encoding harmful prompts as coherent mathematical problems -- using for…

  343. arXiv cs.CL TIER_1 English(EN) · Haesung Lee, Gyubin Choi, Eun-Ju Lee, So-Min Lee, Youkang Ko, Dogyoon Lim, Sung-Kyoung Jang, Yohan Jo ·

    TriBench-Ko:评估司法工作流程中的LLM风险

    arXiv:2605.03792v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into legal workflows. However, existing benchmarks primarily address proxy tasks, such as bar examination performance or classification, which fail to capture the performance …

  344. arXiv cs.LG TIER_1 English(EN) · Yi Liu ·

    两次调用,两个时刻,以及重复LLM推理的投票准确性曲线

    arXiv:2605.03379v1 Announce Type: new Abstract: Repeated sampling is a standard way to spend test-time compute, but its benefit is controlled by the latent distribution of correctness across examples, not by one-call accuracy alone. We study the binary correctness layer of repeat…

  345. arXiv cs.LG TIER_1 English(EN) · Shannon K. Gallagher, Swati Rallapalli, Tyler Brooks, Chuck Loughin, Michele Sezgin, Ronald Yurko ·

    通过进化方法对大型语言模型进行分析和可解释性研究

    arXiv:2605.02930v1 Announce Type: cross Abstract: Evolutionary methods have long been useful for analysis and explanation in genetics, biology, ecology, and related fields. In this work, we extend these methods to neural networks, specifically large language models (LLMs), to bet…

  346. arXiv cs.CL TIER_1 English(EN) · Richard A. A. Jonker, Alexander Christiansen, Alexandros Maniatis, R\'uben Garrido, Rog\'erio Braunschweiger de Freitas Lima, Roman Jurowetzki, S\'ergio Matos ·

    BIT.UA-AAUBS 在 ArchEHR-QA 2026 上:通过低资源问答中的提示评估开源和专有 LLM

    arXiv:2605.03618v1 Announce Type: new Abstract: This paper presents the joint participation of the BIT.UA and AAUBS groups in the ArchEHR-QA 2026 shared task, which focuses on clinical question answering and evidence grounding in a low-resource setting. Due to the absence of trai…

  347. arXiv cs.AI TIER_1 English(EN) · Jia Xiao ·

    NeuroState-Bench:LLM 代理配置文件中承诺完整性的人类校准基准

    arXiv:2605.01847v1 Announce Type: new Abstract: Outcome-only evaluation under-specifies whether an evaluated agent profile preserves the commitments required to solve a multi-turn task coherently. NeuroState-Bench is a human-calibrated benchmark that operationalizes commitment in…

  348. arXiv cs.AI TIER_1 English(EN) · Yifei Wang, Ruiyin Li, Peng Liang, Yangxiao Cai, Zengyang Li, Mojtaba Shahin, Arif Ali Khan, Qiong Feng ·

    在软件设计中使用LLM:一项来自GitHub和从业者调查的实证研究

    arXiv:2605.01392v1 Announce Type: cross Abstract: Recent advancements in Large Language Models (LLMs) have demonstrated significant potential across a wide range of software engineering tasks, including software design, an area traditionally regarded as highly dependent on human …

  349. arXiv cs.AI TIER_1 English(EN) · Youpeng Li, Fuxun Yu, Xinda Wang ·

    从SFT到RL:解密LLM驱动漏洞检测的训练后流程

    arXiv:2602.14012v2 Announce Type: replace-cross Abstract: The integration of LLMs into vulnerability detection (VD) has shifted the field toward more interpretable and context-aware analysis. While post-training techniques have shown promise in general coding tasks, their systema…

  350. arXiv cs.LG TIER_1 English(EN) · Miaomiao Li, Hao Chen, Yang Wang, Tingyuan Zhu, Weijia Zhang, Kaijie Zhu, Kam-Fai Wong, Jindong Wang ·

    理解和减轻下游任务中基于LLM的数据增强的偏见继承

    arXiv:2502.04419v3 Announce Type: replace Abstract: Generating synthetic datasets via large language models (LLMs) has emerged as a promising approach to improve LLM performance. However, LLMs inherently reflect biases in their training data, leading to a critical challenge: when…

  351. arXiv cs.LG TIER_1 English(EN) · Hyunji Nam, Haoran Li, Natasha Jaques ·

    最大化提示与响应之间的互信息可改善 LLM 个性化,无需额外数据或人工监督

    arXiv:2603.19294v2 Announce Type: replace Abstract: While post-training has successfully improved large language models (LLMs) across a variety of domains, these gains heavily rely on human-labeled data or external verifiers. Existing data has already been exploited, and new high…

  352. arXiv cs.CL TIER_1 English(EN) · Yohan Jo ·

    TriBench-Ko:评估司法工作流程中的 LLM 风险

    Large language models (LLMs) are increasingly integrated into legal workflows. However, existing benchmarks primarily address proxy tasks, such as bar examination performance or classification, which fail to capture the performance and risks inherent in day-to-day judicial proces…

  353. arXiv cs.CL TIER_1 English(EN) · Sérgio Matos ·

    BIT.UA-AAUBS 在 ArchEHR-QA 2026 上:通过低资源问答中的提示评估开源和专有 LLM

    This paper presents the joint participation of the BIT.UA and AAUBS groups in the ArchEHR-QA 2026 shared task, which focuses on clinical question answering and evidence grounding in a low-resource setting. Due to the absence of training data and the strict data privacy constraint…

  354. arXiv cs.CL TIER_1 English(EN) · Shanu Sushmita ·

    通过数学编码揭示LLM安全漏洞:新的攻击与系统性分析

    Large language models (LLMs) employ safety mechanisms to prevent harmful outputs, yet these defenses primarily rely on semantic pattern matching. We show that encoding harmful prompts as coherent mathematical problems -- using formalisms such as set theory, formal logic, and quan…

  355. arXiv cs.CL TIER_1 English(EN) · Yi Liu ·

    两次调用,两个时刻,以及重复LLM推理的投票准确性曲线

    Repeated sampling is a standard way to spend test-time compute, but its benefit is controlled by the latent distribution of correctness across examples, not by one-call accuracy alone. We study the binary correctness layer of repeated LLM inference under conditional-i.i.d. calls.…

  356. arXiv cs.CL TIER_1 English(EN) · Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Reduan Achtibat, Patrick Kahardipraja, Thomas Wiegand, Wojciech Samek, Alexander Binder, Sebastian Lapuschkin ·

    用于洞察和控制的归因引导剪枝:小规模 LLM 中的电路发现与定向修正

    arXiv:2506.13727v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are widely deployed in real-world applications, yet their internal mechanisms remain difficult to interpret and control, limiting our ability to diagnose and correct undesirable behaviors. Mech…

  357. arXiv cs.LG TIER_1 English(EN) · Jimyung Hong, Jaehyung Kim ·

    Diet Your LLM:通过合并特定任务的评分来实现 LLM 的维度全局剪枝

    arXiv:2603.23985v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated remarkable capabilities, but their massive scale poses significant challenges for practical deployment. Structured pruning offers a promising solution by removing entire dimensions …

  358. arXiv cs.LG TIER_1 English(EN) · Timoth\'ee Chauvin, Cl\'ement Lalanne, Erwan Le Merrer, Jean-Michel Loubes, Fran\c{c}ois Ta\"iani, Gilles Tredan ·

    LLM API 中的令牌高效变化检测

    arXiv:2602.11083v2 Announce Type: replace Abstract: Remote change detection in LLMs is a difficult problem. Existing methods are either too expensive for deployment at scale, or require initial white-box access to model weights or grey-box access to log probabilities. We aim to a…

  359. arXiv cs.LG TIER_1 English(EN) · Nickil Maveli, Antonio Vergari, Shay B. Cohen ·

    大型语言模型能压缩(和解压缩)吗?通过可逆性评估代码理解和执行能力

    arXiv:2601.13398v2 Announce Type: replace Abstract: LLMs demonstrate strong performance on code benchmarks, yet consistent reasoning across forward and backward execution remains elusive. We present RoundTripCodeEval (RTCE), a benchmark of four code execution reasoning tasks that…

  360. arXiv cs.CL TIER_1 English(EN) · Ian Rios-Sialer ·

    大型语言模型中的同质化问题:迈向人工智能安全领域有意义的多样性

    arXiv:2601.06116v3 Announce Type: replace-cross Abstract: Generative AI models reproduce the human biases in their training data and further amplify them through mechanisms such as mode collapse. The loss of diversity produces homogenization, which not only harms the minoritized …

  361. arXiv cs.CL TIER_1 English(EN) · Antonio Valerio Miceli Barone, Poon Tsz Nok ·

    通过形式化验证的语义等价自我博弈改进 LLM 代码推理

    arXiv:2604.17010v2 Announce Type: replace Abstract: We introduce a self-play framework for semantic equivalence in Haskell, utilizing formal verification to guide adversarial training between a generator and an evaluator. The framework leverages Liquid Haskell proofs for validati…

  362. arXiv cs.CL TIER_1 English(EN) · Ziyi Zhu, Olivier Tieleman, Alexey Bukhtiyarov, Jinghong Chen ·

    CyclicJudge:高效缓解 LLM 评估中的裁判偏见

    arXiv:2603.01865v3 Announce Type: replace Abstract: LLM-as-judge evaluation has become standard practice for open-ended model assessment; however, judges exhibit systematic biases that cannot be averaged out by increasing the number of scenarios or generations. These biases are o…

  363. arXiv cs.CL TIER_1 English(EN) · Pawel Kaplanski (Kaplanski AI Lab) ·

    递归式LLM循环中的扰动剂量响应:原始切换、随机地板和追加、替换及对话更新下的持续逃逸

    arXiv:2605.02236v1 Announce Type: cross Abstract: Recursive language-model loops often settle into recognizable attractor-like patterns. The practical question is how much injected text is needed to move a settled loop somewhere else, and whether that move lasts. We study this in…

  364. arXiv cs.CL TIER_1 English(EN) · Noga Peleg Pelc, Gal A. Kaminka, Yoav Goldberg ·

    一种描述Agentic LLM上下文的语言

    arXiv:2605.01920v1 Announce Type: cross Abstract: Large language models are increasingly used within larger systems ("LLM agents"). These make a sequence of LLM calls, each call providing the LLM with a combination of instructions, observations, and interaction history. The desig…

  365. arXiv cs.CL TIER_1 English(EN) · Sadia Asif, Mohammad Mohammadi Amiri ·

    RefusalGuard:LLM安全中的几何保持微调

    arXiv:2605.01913v1 Announce Type: cross Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable to adversarial misuse. While prior work has shown that safety-relevant features a…

  366. arXiv cs.CL TIER_1 English(EN) · Benjamin Warner, Ratna Sagari Grandhi, Max Kieffer, Aymane Ouraq, Saurav Panigrahi, Geetu Ambwani, Kunal Bagga, Nikhil Khandekar, Arya Hariharan, Nishant Mishra, Manish Ram, Shamus Sim Zi Yang, Ahmed Essouaied, Adepoju Jeremiah Moyondafoluwa, Robert Schol ·

    Medmarks:面向医疗任务的综合性开源大语言模型基准套件

    arXiv:2605.01417v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for medical applications remains challenging due to benchmark saturation, limited data accessibility, and insufficient coverage of relevant tasks. Existing suites have either saturated, heavil…

  367. arXiv cs.CL TIER_1 English(EN) · Koshiro Saito, Ryuto Koike, Masahiro Kaneko, Naoaki Okazaki ·

    LLM输出可检测性与任务性能可联合优化

    arXiv:2605.01350v1 Announce Type: new Abstract: Detecting machine-generated text is essential for transparency and accountability when deploying large language models (LLMs). Among detection approaches, watermarking is a statistically reliable method by design -- it embeds detect…

  368. arXiv cs.CL TIER_1 English(EN) · Kevin H. Guo, Chao Yan, Avinash Baidya, Katherine Brown, Xiang Goa, Juming Xiong, Zhijun Yin, Bradley A. Malin ·

    CLEAR:揭示噪声和歧义如何降低医学领域大型语言模型(LLMs)的可靠性

    arXiv:2605.01011v1 Announce Type: new Abstract: Medical large language model (LLM) evaluations rely on simplified, exam-style benchmarks that rarely reflect the ambiguity of real-world medical inquiries. We introduce the CLinical Evaluation of Ambiguity and Reliability (CLEAR) fr…

  369. arXiv cs.AI TIER_1 English(EN) · Lehan He, Zeren Chen, Zhe Zhang, Xiang Gao, Lu Sheng ·

    通过面向属性和结构最小化的反馈实现有效的 LLM 代码优化

    arXiv:2506.18315v2 Announce Type: replace-cross Abstract: LLMs excel at code generation, yet ensuring the functional correctness of their outputs remains a persistent challenge. While recent studies have applied Test-Driven Development (TDD) to refine code, these methods are ofte…

  370. arXiv cs.AI TIER_1 English(EN) · Abdurrahman Javat, Allan Kazakov ·

    硅谷对决:消费级大模型推理的性能、效率与生态壁垒

    arXiv:2605.00519v2 Announce Type: cross Abstract: The operational landscape of local Large Language Model (LLM) inference has shifted from lightweight models to datacenter-class weights exceeding 70B parameters, creating profound systems challenges for consumer hardware. This pap…

  371. arXiv cs.AI TIER_1 English(EN) · Fazle Rabbi, Lin Ling, Song Wang, Jinqiu Yang ·

    LLM 生成代码中的社会偏见:基准测试与缓解

    arXiv:2605.00382v2 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed to generate code for human-centered applications where demographic fairness is critical. However, existing evaluations focus almost exclusively on functional correctness, leav…

  372. arXiv cs.AI TIER_1 English(EN) · Qinyuan Wu, Soumi Das, Mahsa Amani, Arijit Nag, Seungeon Lee, Krishna P. Gummadi, Abhilasha Ravichander, Muhammad Bilal Zafar ·

    调用还是不调用:评估和优化 LLM 工具调用的框架

    arXiv:2605.00737v1 Announce Type: new Abstract: Agentic AI architectures augment LLMs with external tools, unlocking strong capabilities. However, tool use is not always beneficial; some calls may be redundant or even harmful. Effective tool use, therefore, hinges on a core LLM d…

  373. arXiv cs.CL TIER_1 English(EN) · Pawel Kaplanski ·

    递归式LLM循环中的扰动剂量响应:原始切换、随机地板以及追加、替换和对话更新下的持续逃逸

    Recursive language-model loops often settle into recognizable attractor-like patterns. The practical question is how much injected text is needed to move a settled loop somewhere else, and whether that move lasts. We study this in 30-step recursive loops by separating the model f…

  374. arXiv cs.CL TIER_1 Français(FR) · Ryan Lail, Luke Markham ·

    关于成本效益高的 LLM-as-a-Judge 改进技术

    arXiv:2604.13717v2 Announce Type: replace Abstract: Using a language model to score or rank candidate responses has become a scalable alternative to human evaluation in reinforcement learning from human feedback (RLHF) pipelines, benchmarking, and application layer evaluations. H…

  375. arXiv cs.CL TIER_1 English(EN) · Zongqi Wang, Tianle Gu, Chen Gong, Xin Tian, Siqi Bao, Yujiu Yang ·

    SCAN:LLM 的结构化能力评估与导航

    arXiv:2505.06698v4 Announce Type: replace Abstract: Evaluating Large Language Models (LLMs) has become increasingly important, with automatic evaluation benchmarks gaining prominence as alternatives to human evaluation. While existing research has focused on approximating model r…

  376. arXiv cs.CL TIER_1 English(EN) · Sailesh Panda, Pritam Kadasi, Abhishek Upperwal, Mayank Singh ·

    当大型语言模型不再遵循步骤:一项关于语言模型程序执行的诊断研究

    arXiv:2605.00817v1 Announce Type: new Abstract: Large language models (LLMs) often achieve strong performance on reasoning benchmarks, but final-answer accuracy alone does not show whether they faithfully execute the procedure specified in a prompt. We study this question through…

  377. arXiv cs.LG TIER_1 English(EN) · Pavlin G. Poli\v{c}ar, Andra\v{z} Pevcin, Bla\v{z} Zupan ·

    使用验证驱动的 LLM 工作流生成统计图表

    arXiv:2605.00800v1 Announce Type: new Abstract: Generating diverse, readable statistical charts from tabular data remains challenging for LLMs, as many failures become apparent after rendering and are not detectable from data or code alone. Existing chart datasets also rarely pro…

  378. arXiv cs.LG TIER_1 English(EN) · Jiale Fu, Yuchu Jiang, Peijun Wu, Chonghan Liu, Joey Tianyi Zhou, Xu Yang ·

    从混合模型的视角重新思考LLM集成

    arXiv:2605.00419v1 Announce Type: new Abstract: Model ensembling is a well-established technique for improving the performance of machine learning models. Conventionally, this involves averaging the output distributions of multiple models and selecting the most probable label. Th…

  379. arXiv cs.CL TIER_1 English(EN) · Yoav Goldberg ·

    一种描述Agentic LLM上下文的语言

    Large language models are increasingly used within larger systems ("LLM agents"). These make a sequence of LLM calls, each call providing the LLM with a combination of instructions, observations, and interaction history. The design of the encoded information and its structure pla…

  380. arXiv cs.CL TIER_1 English(EN) · Mohammad Mohammadi Amiri ·

    RefusalGuard:LLM安全中的几何保持微调

    Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable to adversarial misuse. While prior work has shown that safety-relevant features are encoded in structured representations within th…

  381. arXiv cs.CL TIER_1 English(EN) · Mayank Singh ·

    当大型语言模型不再遵循步骤:一项关于语言模型程序执行的诊断研究

    Large language models (LLMs) often achieve strong performance on reasoning benchmarks, but final-answer accuracy alone does not show whether they faithfully execute the procedure specified in a prompt. We study this question through a controlled diagnostic benchmark for procedura…

  382. arXiv cs.LG TIER_1 English(EN) · Blaž Zupan ·

    使用验证驱动的 LLM 工作流生成统计图表

    Generating diverse, readable statistical charts from tabular data remains challenging for LLMs, as many failures become apparent after rendering and are not detectable from data or code alone. Existing chart datasets also rarely provide fully aligned artifacts, such as executable…

  383. arXiv cs.AI TIER_1 English(EN) · Muhammad Bilal Zafar ·

    调用还是不调用:评估和优化LLM工具调用的框架

    Agentic AI architectures augment LLMs with external tools, unlocking strong capabilities. However, tool use is not always beneficial; some calls may be redundant or even harmful. Effective tool use, therefore, hinges on a core LLM decision: whether to call or not call a tool, whe…

  384. arXiv cs.AI TIER_1 English(EN) · Abdurrahman Javat ·

    硅谷对决:消费级大模型推理的性能、效率与生态系统壁垒

    The operational landscape of local Large Language Model (LLM) inference has shifted from lightweight models to datacenter-class weights exceeding 70B parameters, creating profound systems challenges for consumer hardware. This paper presents a systematic empirical analysis of the…

  385. arXiv cs.CL TIER_1 English(EN) · Xu Yang ·

    从混合模型的视角重新思考LLM集成

    Model ensembling is a well-established technique for improving the performance of machine learning models. Conventionally, this involves averaging the output distributions of multiple models and selecting the most probable label. This idea has been naturally extended to large lan…

  386. arXiv cs.AI TIER_1 English(EN) · Jinqiu Yang ·

    LLM 生成代码中的社会偏见:基准测试与缓解

    Large Language Models (LLMs) are increasingly deployed to generate code for human-centered applications where demographic fairness is critical. However, existing evaluations focus almost exclusively on functional correctness, leaving social bias in LLM-generated code largely unex…

  387. arXiv cs.AI TIER_1 English(EN) · Jon-Paul Cacioli ·

    超越均值:LLM评估的单模型可靠性变化检测

    arXiv:2604.27405v1 Announce Type: cross Abstract: We adapted the Reliable Change Index (RCI; Jacobson and Truax, 1991) from clinical psychology to item-level LLM version comparison on 2,000 MMLU-Pro items (K=10 samples at T=0.7). Two within-family pairs were tested: Llama 3 to 3.…

  388. arXiv cs.CL TIER_1 English(EN) · Solomon Messing ·

    LLM 管道中的隐藏测量误差扭曲了标注、评估和基准测试

    arXiv:2604.11581v4 Announce Type: replace Abstract: LLM evaluations drive which models get deployed, which safety standards get adopted, and which research conclusions get published. Yet standard confidence intervals ignore variability from prompt phrasing, model temperature, and…

  389. arXiv cs.AI TIER_1 English(EN) · Ziyao Xu, Cong Wang, Houfeng Wang ·

    从规则生成视角研究LLM更具可解释性和无分区组合性估计

    arXiv:2604.27340v1 Announce Type: new Abstract: Compositional generalization tests are often used to estimate the compositionality of LLMs. However, such tests have the following limitations: (1) they only focus on the output results without considering LLMs' understanding of sam…

  390. arXiv cs.LG TIER_1 English(EN) · Ahan Gupta, Zhihao Wang, Neel Dani, Masahiro Tanaka, Olatunji Ruwase, Minjia Zhang ·

    AutoSP:通过编译器式序列并行解锁长上下文 LLM 训练

    arXiv:2604.27089v1 Announce Type: new Abstract: Large-language-models (LLMs) demonstrate enormous utility in long-context tasks which require processing prompts that consist of tens to hundreds of thousands of tokens. However, existing LLM training libraries do not provide easy t…

  391. arXiv cs.LG TIER_1 English(EN) · Jun Yeon Won, Xin Jin, Shiqing Ma, Zhiqiang Lin ·

    REBENCH:一个用于 LLM 在剥离二进制类型和名称上的程序化、内建公平基准测试(扩展版)

    arXiv:2604.27319v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable progress in recent years, driving their adoption across a wide range of domains, including computer security. In reverse engineering, LLMs are increasingly applied to critical …

  392. arXiv cs.CL TIER_1 English(EN) · Jon-Paul Cacioli ·

    超越均值:LLM评估的单模型可靠性变化检测

    We adapted the Reliable Change Index (RCI; Jacobson and Truax, 1991) from clinical psychology to item-level LLM version comparison on 2,000 MMLU-Pro items (K=10 samples at T=0.7). Two within-family pairs were tested: Llama 3 to 3.1 (+1.6 points) and Qwen 2.5 to 3 (+2.8 points). O…

  393. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越均值:LLM评估的单模型可靠性变化检测

    We adapted the Reliable Change Index (RCI; Jacobson and Truax, 1991) from clinical psychology to item-level LLM version comparison on 2,000 MMLU-Pro items (K=10 samples at T=0.7). Two within-family pairs were tested: Llama 3 to 3.1 (+1.6 points) and Qwen 2.5 to 3 (+2.8 points). O…

  394. arXiv cs.CL TIER_1 English(EN) · Wenxuan Wang, Juluan Shi, Zixuan Ling, Yuk-Kit Chan, Chaozheng Wang, Cheryl Lee, Youliang Yuan, Jen-tse Huang, Wenxiang Jiao, Michael R. Lyu ·

    学习提问:当大型语言模型代理遇到不明确指令时

    arXiv:2409.00557v4 Announce Type: replace Abstract: Equipped with the capability to call functions, modern large language models (LLMs) can leverage external tools for addressing a range of tasks unattainable through language skills alone. However, the effective execution of thes…

  395. arXiv cs.AI TIER_1 English(EN) · Zoe Kotti, Konstantina Dritsa, Diomidis Spinellis, Panos Louridas ·

    愚者笃定,智者怀疑:探索LLM在代码补全中的置信度

    arXiv:2508.16131v2 Announce Type: replace-cross Abstract: Code completion entails the task of providing missing tokens given a surrounding context. It can boost developer productivity while providing a powerful code discovery tool. Following the Large Language Model (LLM) wave, c…

  396. arXiv cs.AI TIER_1 English(EN) · Emre Furkan Akyol, Mehmet Dedeler, Eray T\"uz\"un ·

    ImproBR:使用 LLM 的 Bug Report Improver

    arXiv:2604.26142v1 Announce Type: cross Abstract: Bug tracking systems play a crucial role in software maintenance, yet developers frequently struggle with low-quality user-submitted reports that omit essential details such as Steps to Reproduce (S2R), Observed Behavior (OB), and…

  397. arXiv cs.CL TIER_1 English(EN) · Sasha Ronaghi, Chloe Stanwyck, Asad Aali, Amir Ronaghi, Miguel Fuentes, Tina Hernandez-Boussard, Emily Alsentzer ·

    使用遗留临床模型对新一代大语言模型进行无训练微调

    arXiv:2601.03423v3 Announce Type: replace Abstract: Adapting language models to the clinical domain through continued pretraining and instruction tuning requires costly retraining for each new model generation. We propose Cross-Architecture Proxy Tuning (CAPT), a model-ensembling…

  398. arXiv cs.CL TIER_1 English(EN) · Samee Arif, Naihao Deng, Zhijing Jin, Rada Mihalcea ·

    逐字逐句:增量式完成分解破坏LLM安全

    arXiv:2604.25921v1 Announce Type: new Abstract: Large Language Models (LLMs) are trained to refuse harmful requests, yet they remain vulnerable to jailbreak attacks that exploit weaknesses in conversational safety mechanisms. We introduce Incremental Completion Decomposition (ICD…

  399. arXiv cs.CL TIER_1 English(EN) · Hongyeon Yu, Young-Bum Kim, Yoon Kim ·

    FlowBot:利用双层优化和文本梯度诱导 LLM 工作流

    arXiv:2604.26258v1 Announce Type: new Abstract: LLM workflows, which coordinate structured calls to individual LLMs (each augmented with varying instructions and tools) to achieve a particular goal, offer a promising path towards extending the capabilities of LLMs and building po…

  400. Hugging Face Daily Papers TIER_1 English(EN) ·

    AutoSP:通过编译器式序列并行解锁长上下文 LLM 训练

    Large-language-models (LLMs) demonstrate enormous utility in long-context tasks which require processing prompts that consist of tens to hundreds of thousands of tokens. However, existing LLM training libraries do not provide easy to use abstractions to optimize for long-context …

  401. arXiv cs.CL TIER_1 English(EN) · Avinash Amballa, Yashas Malur Saidutta, Chi-Heng Lin, Vivek Kulkarni, Srinivas Chappidi ·

    VOYAGER:一种使用大型语言模型生成多样化数据集的无训练方法

    arXiv:2512.12072v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly being used to generate synthetic datasets for the evaluation and training of downstream models. However, prior work has noted that such generated data lacks diversity. In this paper,…

  402. arXiv cs.CL TIER_1 English(EN) · Ocean Monjur, Shahriar Kabir Nahin, Anshuman Chhabra ·

    少即是多:重新审视LLM剪枝在测试时扩展中的有效性

    arXiv:2604.25098v1 Announce Type: cross Abstract: While current Large Language Models (LLMs) exhibit remarkable reasoning capabilities through test-time compute scaling (TTS), their massive parameter counts and high inference costs have motivated the development of pruning method…

  403. arXiv cs.CL TIER_1 English(EN) · Huyen Nguyen, Haoxuan Zhang, Yang Zhang, Junhua Ding, Haihua Chen ·

    LLM-ReSum:通过自我评估实现LLM反思性摘要的框架

    arXiv:2604.25665v1 Announce Type: new Abstract: Reliable evaluation of large language model (LLM)-generated summaries remains an open challenge, particularly across heterogeneous domains and document lengths. We conduct a comprehensive meta-evaluation of 14 automatic summarizatio…

  404. arXiv cs.CL TIER_1 English(EN) · Alif Munim, Jun Ma, Omar Ibrahim, Alhusain Abdalla, Shuolin Yin, Leo Chen, Bo Wang ·

    基准测试和适应设备端 LLM 以支持临床决策

    arXiv:2601.03266v2 Announce Type: replace Abstract: Large language models (LLMs) have rapidly advanced in clinical decision-making, yet the deployment of proprietary systems is hindered by privacy concerns and reliance on cloud-based infrastructure. Open-source alternatives allow…

  405. arXiv cs.LG TIER_1 English(EN) · Keita Broadwater ·

    通过加速提示压力测试评估重复推理下的LLM安全性

    arXiv:2602.11786v2 Announce Type: replace Abstract: Traditional benchmarks for large language models (LLMs), such as HELM and AIR-BENCH, primarily assess safety through breadth-oriented evaluation across diverse tasks and risk categories. However, real-world deployment often expo…

  406. arXiv cs.CL TIER_1 English(EN) · Yoon Kim ·

    FlowBot:利用双层优化和文本梯度诱导 LLM 工作流

    LLM workflows, which coordinate structured calls to individual LLMs (each augmented with varying instructions and tools) to achieve a particular goal, offer a promising path towards extending the capabilities of LLMs and building powerful systems that can tackle diverse tasks. Ho…

  407. arXiv cs.CL TIER_1 English(EN) · Haihua Chen ·

    LLM-ReSum:通过自我评估实现LLM反思性摘要的框架

    Reliable evaluation of large language model (LLM)-generated summaries remains an open challenge, particularly across heterogeneous domains and document lengths. We conduct a comprehensive meta-evaluation of 14 automatic summarization metrics and LLM-based evaluators across seven …

  408. Hugging Face Daily Papers TIER_1 English(EN) ·

    LLM-ReSum:通过自我评估实现LLM反思性摘要的框架

    Reliable evaluation of large language model (LLM)-generated summaries remains an open challenge, particularly across heterogeneous domains and document lengths. We conduct a comprehensive meta-evaluation of 14 automatic summarization metrics and LLM-based evaluators across seven …

  409. arXiv cs.LG TIER_1 English(EN) · Jinglue Xu, Qi Sun, Peter Schwendeman, Stefan Nielsen, Edoardo Cetin, Yujin Tang ·

    TRINITY:一个进化的LLM协调器

    arXiv:2512.04695v3 Announce Type: replace Abstract: Combining diverse foundation models is promising, but weight-merging is limited by mismatched architectures and closed APIs. Trinity addresses this with a lightweight coordinator that orchestrates collaboration among large langu…

  410. arXiv cs.LG TIER_1 English(EN) · Zhengding Hu, Hehua Ouyang, Chang Chen, Zaifeng Pan, Yue Guan, Zhongkai Yu, Zhen Wang, Steven Swanson, Yufei Ding ·

    JigsawRL:组装RL流水线以实现高效LLM训练后处理

    arXiv:2604.23838v1 Announce Type: new Abstract: We present JigsawRL, a cost-efficient framework that explores Pipeline Multiplexing as a new dimension of RL parallelism. JigsawRL decomposes each pipeline into a Sub-Stage Graph that exposes the intra-stage and inter-worker imbalan…

  411. arXiv cs.LG TIER_1 English(EN) · Xuancheng Li, Haitao Li, Yujia Zhou, Yiqun Liu, Qingyao Ai ·

    超越经验检索:为冻结的LLM学习生成优化效用的结构化经验

    arXiv:2602.02556v2 Announce Type: replace Abstract: Large language models (LLMs) are largely static and often redo reasoning or repeat mistakes. Prior experience reuse typically relies on external retrieval, which is similarity-based, can introduce noise, and adds latency. We int…

  412. arXiv cs.LG TIER_1 English(EN) · Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma ·

    持续校准:在终身LLM微调中,覆盖率可能在准确性之前崩溃

    arXiv:2604.23987v1 Announce Type: new Abstract: Continual learning for large language models is typically evaluated through accuracy retention under sequential fine-tuning. We argue that this perspective is incomplete, because uncertainty reliability can degrade earlier and more …

  413. arXiv cs.CL TIER_1 English(EN) · Rohith Reddy Bellibatlu ·

    JudgeSense:LLM作为裁判系统中的提示词敏感性基准

    arXiv:2604.23478v1 Announce Type: new Abstract: Large language models are increasingly deployed as automated judges for evaluating other models, yet the stability of their verdicts under semantically equivalent prompt paraphrases remains unmeasured. We introduce JudgeSense, a fra…

  414. arXiv cs.CL TIER_1 English(EN) · Alessio Sordo, Lingxiao Du, Meeka-Hanna Lenisa, Evgeny Bogdanov, Maxim Romanovsky ·

    STELLAR-E:一个合成的、定制的、端到端的LLM应用严格评估器

    arXiv:2604.24544v1 Announce Type: cross Abstract: The increasing reliance on Large Language Models (LLMs) across diverse sectors highlights the need for robust domain-specific and language-specific evaluation datasets; however, the collection of such datasets is challenging due t…

  415. arXiv cs.CL TIER_1 English(EN) · Chenyang Yang, Yike Shi, Qianou Ma, Michael Xieyang Liu, Christian K\"astner, Tongshuang Wu ·

    提示词未言明之处:理解和管理LLM提示词中的不确定性

    arXiv:2505.13360v3 Announce Type: replace Abstract: Prompt underspecification is a common challenge when interacting with LLMs. In this paper, we present an in-depth analysis of this problem, showing that while LLMs can often infer unspecified requirements by default (41.1%), suc…

  416. arXiv cs.CL TIER_1 English(EN) · Zhiqiu Xu, Shibo Jin, Shreya Arya, Mayur Naik ·

    MathDuels:将大型语言模型评估为问题提出者和解决者

    arXiv:2604.21916v2 Announce Type: replace Abstract: As frontier language models attain near-ceiling performance on static mathematical benchmarks, existing evaluations are increasingly unable to differentiate model capabilities, largely because they cast models solely as solvers …

  417. arXiv cs.AI TIER_1 English(EN) · Huzaifa Arif, Keerthiram Murugesan, Ching-Yun Ko, Pin-Yu Chen, Payel Das, Alex Gittens ·

    像修补软件一样修补LLM:一种改进大型语言模型安全策略的轻量级方法

    arXiv:2511.08484v2 Announce Type: replace Abstract: We propose patching for large language models (LLMs) like software versions, a lightweight and modular approach for addressing safety vulnerabilities. While vendors release improved LLM versions, major releases are costly, infre…

  418. arXiv cs.LG TIER_1 English(EN) · Juyeon Yoon, Somin Kim, Robert Feldt, Shin Yoo ·

    Clotho:衡量LLM输入特定任务的预生成测试充分性

    arXiv:2509.17314v3 Announce Type: replace-cross Abstract: Software increasingly relies on the emergent capabilities of Large Language Models (LLMs), from natural language understanding to program analysis and generation. Yet testing them on specific tasks remains difficult and co…

  419. arXiv cs.CL TIER_1 English(EN) · Yue Liu, Yingwei Ma, Yibo Miao, Yanhao Li, Yuchong Xie, Xinlong Yang, Zhiyuan Hu, Flood Sung, Jiaheng Zhang, Bryan Hooi ·

    KLong:为极长时域任务训练LLM Agent

    arXiv:2602.17547v3 Announce Type: replace-cross Abstract: This paper introduces KLong, an open-source LLM agent trained to solve extremely long-horizon tasks. The principle is to first cold-start the model via trajectory-splitting SFT, then scale it via progressive RL training. S…

  420. arXiv cs.LG TIER_1 English(EN) · Frank Xiao, Santiago Aranguri ·

    基于探针的数据归因:发现和缓解大型语言模型训练后不良行为

    arXiv:2602.11079v3 Announce Type: replace Abstract: We propose probe-based data attribution, a method that traces behavioral changes in post-trained language models to responsible training datapoints. By computing activation-difference vectors for both test prompts and preference…

  421. arXiv cs.CL TIER_1 English(EN) · Anshuman Chhabra ·

    以更少资源做更多事:重新审视LLM剪枝在测试时扩展中的有效性

    While current Large Language Models (LLMs) exhibit remarkable reasoning capabilities through test-time compute scaling (TTS), their massive parameter counts and high inference costs have motivated the development of pruning methods that can reduce model size without sacrificing p…

  422. arXiv cs.CL TIER_1 English(EN) · Maxim Romanovsky ·

    STELLAR-E:一个合成的、定制的、端到端的LLM应用严谨评估器

    The increasing reliance on Large Language Models (LLMs) across diverse sectors highlights the need for robust domain-specific and language-specific evaluation datasets; however, the collection of such datasets is challenging due to privacy concerns, regulatory restrictions, and t…

  423. Hugging Face Daily Papers TIER_1 English(EN) ·

    STELLAR-E:一个合成的、定制的、端到端的LLM应用严谨评估器

    The increasing reliance on Large Language Models (LLMs) across diverse sectors highlights the need for robust domain-specific and language-specific evaluation datasets; however, the collection of such datasets is challenging due to privacy concerns, regulatory restrictions, and t…

  424. arXiv cs.CL TIER_1 English(EN) · Sourav Saha, Mandar Mitra, Aditya Dutta ·

    大型语言模型作为评估者:理由正确吗?

    arXiv:2601.08919v2 Announce Type: replace-cross Abstract: A good deal of recent research has focused on how Large Language Models (LLMs) may be used as judges in place of humans to evaluate the quality of the output produced by various text / image processing systems. Within this…

  425. arXiv cs.AI TIER_1 English(EN) · Manuel Alejandro Borroto Santana, Erica Coppolillo, Francesco Calimeri, Giuseppe Manco, Simona Perri, Francesco Ricca ·

    BLAST:基于 ASP 的结构化测试对大型语言模型进行基准测试

    arXiv:2604.22306v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated remarkable performance across a broad spectrum of tasks, including natural language understanding, dialogue systems, and code generation. Despite evident progress, less attention has …

  426. arXiv cs.LG TIER_1 English(EN) · Zhaokun Wang, Jinyu Guo, Jingwen Pu, Hongli Pu, Meng Yang, Xunlei Chen, Jie Ou, Wenyi Li, Guangchun Luo, Wenhong Tian ·

    CAP:大型语言模型中可控对齐提示的遗忘

    arXiv:2604.21251v2 Announce Type: replace Abstract: Large language models (LLMs) trained on unfiltered corpora inherently risk retaining sensitive information, necessitating selective knowledge unlearning for regulatory compliance and ethical safety. However, existing parameter-m…

  427. arXiv cs.LG TIER_1 English(EN) · Emil Ryd, Henning Bartsch, Julian Stastny, Joe Benton, Vivek Hebbar ·

    通过弱监督训练消除LLM中的“沙袋效应”

    arXiv:2604.22082v1 Announce Type: new Abstract: As AI systems begin to automate complex tasks, supervision increasingly relies on weaker models or limited human oversight that cannot fully verify output quality. A model more capable than its supervisors could exploit this gap thr…

  428. arXiv cs.AI TIER_1 English(EN) · Francesco Ricca ·

    BLAST:基于ASP的结构化测试的LLM基准测试

    Large Language Models (LLMs) have demonstrated remarkable performance across a broad spectrum of tasks, including natural language understanding, dialogue systems, and code generation. Despite evident progress, less attention has been paid to their effectiveness in handling decla…

  429. arXiv cs.AI TIER_1 English(EN) · Vivek Hebbar ·

    通过弱监督训练消除LLM中的“沙袋效应”

    As AI systems begin to automate complex tasks, supervision increasingly relies on weaker models or limited human oversight that cannot fully verify output quality. A model more capable than its supervisors could exploit this gap through sandbagging, producing work that appears ac…

  430. Hugging Face Daily Papers TIER_1 English(EN) ·

    MathDuels:评估LLM作为问题提出者和解决者

    As frontier language models attain near-ceiling performance on static mathematical benchmarks, existing evaluations are increasingly unable to differentiate model capabilities, largely because they cast models solely as solvers of fixed problem sets. We introduce MathDuels, a sel…

  431. arXiv cs.CL TIER_1 English(EN) · Mayur Naik ·

    MathDuels:评估LLM作为问题提出者和解决者

    As frontier language models attain near-ceiling performance on static mathematical benchmarks, existing evaluations are increasingly unable to differentiate model capabilities, largely because they cast models solely as solvers of fixed problem sets. We introduce MathDuels, a sel…

  432. arXiv cs.LG TIER_1 English(EN) · Wenhong Tian ·

    CAP:大型语言模型中可控对齐提示的遗忘

    Large language models (LLMs) trained on unfiltered corpora inherently risk retaining sensitive information, necessitating selective knowledge unlearning for regulatory compliance and ethical safety. However, existing parameter-modifying methods face fundamental limitations: high …

  433. Hugging Face Daily Papers TIER_1 English(EN) ·

    HoWToBench:使用写作树对LLM人类写作能力进行整体评估

    Evaluating the writing capabilities of large language models (LLMs) remains a significant challenge due to the multidimensional nature of writing skills and the limitations of existing metrics. LLM's performance in thousand-words level and open-ended writing is inadequately asses…

  434. Ahead of AI (Sebastian Raschka) TIER_1 English(EN) · Sebastian Raschka, PhD ·

    理解LLM评估的4种主要方法(从零开始)

    Multiple-Choice Benchmarks, Verifiers, Leaderboards, and LLM Judges with Code Examples

  435. Ahead of AI (Sebastian Raschka) TIER_1 English(EN) · Sebastian Raschka, PhD ·

    从零开始构建代码大模型:完整课程

    Why build LLMs from scratch? It's probably the best and most efficient way to learn how LLMs really work. Plus, many readers have told me they had a lot of fun doing it.

  436. arXiv stat.ML TIER_1 English(EN) · Chi-Kuang Yeh ·

    通过正负样本学习量化和审计大语言模型评估

    Large Language Models (LLMs) are increasingly used as judges for scalable evaluation, yet such LLM--as--a--Judge systems exhibit systematic biases that are decoupled from semantic quality, most notably verbosity bias. Meanwhile, human supervision is costly and typically selective…

  437. LessWrong (AI tag) TIER_1 English(EN) · gwern ·

    守护天使:LLM个性化以提升生产力和安全性

    <p>Powerful LLMs will be deployed at global scale in the next few years, and will dominate the Internet, and increasingly, ordinary life. As of mid-2026, there is no coherent vision for how knowledge professionals, or ordinary people, will be able to harness these LLMs for large …

  438. arXiv stat.ML TIER_1 English(EN) · Constantinos Antoniou ·

    有限语义下的表格数据大型语言模型:来自工业汽车改装预测的证据

    Industrial retrofit planning depends on structured operational data rather than free text: planners must estimate whether a newly registered prototype will require a retrofit, which retrofit package it will need, and how long the work will take. We study an industrial dataset lin…

  439. arXiv stat.ML TIER_1 English(EN) · Alexandre Belloni, Yan Chen, Yehua Wei ·

    在线潘多拉魔盒:上下文LLM级联

    arXiv:2606.07392v1 Announce Type: cross Abstract: Motivated by Large Language Model (LLM) cascading, we propose an online contextual Pandora's Box model for adaptively querying and selecting LLM APIs. In each period, a decision-maker observes a request context and faces a two-pha…

  440. LessWrong (AI tag) TIER_1 English(EN) · Matthew Khoriaty ·

    摘掉辅助轮:无需个性化即可对齐大型语言模型

    <p><br /></p><img alt="" src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/3wqZwXMzEAkd3mLLM/5b18c66a57f187f33ac8a438209481ce38e836a7fdc1cb081161fad23496bc70/ofzs7dwn131h6fep69y4" /><p><span>If you told an AI Alignment researcher in 2018 a…

  441. LessWrong (AI tag) TIER_1 English(EN) · Vidya Ganga ·

    大型语言模型能教书吗? 探索“教师轴”

    <h2><span>TLDR</span></h2><p><span>As a passionate teacher, it has pained my heart to watch my students lose deeper critical thinking skills and independent reasoning. But attempting to build a constitutionally constrained AI using prompt engineering that acted more Socratically …

  442. LessWrong (AI tag) TIER_1 English(EN) · Owain_Evans ·

    大型语言模型中的脱离上下文推理(OOCR):简短入门与阅读列表

    <p>Out-of-context reasoning (OOCR) is a concept relevant to LLM generalization and AI alignment. Also available as a <a href="https://owainevans.github.io/pdfs/oocr_primer_latex.pdf">PDF</a>.</p> <p><strong>Contents</strong></p> <ol> <li><a href="#what-is-out-of-context-reasoning…

  443. arXiv stat.ML TIER_1 English(EN) · Biswa Sengupta ·

    无奖励表征:用于LLM微调的JEPA审计

    arXiv:2605.15394v1 Announce Type: cross Abstract: Joint-embedding predictive architectures (JEPAs) propose that a model should learn more useful abstractions when trained to predict latent representations rather than observed outputs. For autoregressive language-model fine-tuning…

  444. arXiv stat.ML TIER_1 English(EN) · Biswa Sengupta ·

    无奖励的表征:用于 LLM 微调的 JEPA 审计

    Joint-embedding predictive architectures (JEPAs) propose that a model should learn more useful abstractions when trained to predict latent representations rather than observed outputs. For autoregressive language-model fine-tuning the principle entails a stricter requirement: the…

  445. arXiv stat.ML TIER_1 English(EN) · Stef van Buuren ·

    LLMs 作为隐式填补者:不确定性应随缺失信息而扩展

    arXiv:2605.13188v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in settings where the available context is incomplete or degraded. We argue that an LLM generating answers under incomplete context can be viewed as an implicit imputer, and eva…

  446. arXiv stat.ML TIER_1 English(EN) · Zetai Cen, Chenfei Gu, Jin Zhu, Ting Li, Yunxiao Chen, Chengchun Shi ·

    学习扰动以推断您的LLM

    arXiv:2605.13284v1 Announce Type: new Abstract: Recent advancements in large language models demonstrate that injecting perturbations can substantially enhance extrapolation performance. However, current approaches often rely on discrete perturbations with fixed designs, which li…

  447. LessWrong (AI tag) TIER_1 English(EN) · Santiago Aranguri ·

    预测罕见的LLM故障,回滚次数减少30倍

    <p><span>TL;DR: We estimate how often Qwen 3 4B exhibits rare harmful behaviors with 30× fewer rollouts than naive sampling, using a new method that interpolates between the model and a less-safe variant in logit space.</span></p><p><span>Authors: Francisco Pernice (MIT), Santiag…

  448. arXiv stat.ML TIER_1 English(EN) · Chengchun Shi ·

    学习扰动以推断您的LLM

    Recent advancements in large language models demonstrate that injecting perturbations can substantially enhance extrapolation performance. However, current approaches often rely on discrete perturbations with fixed designs, which limits their flexibility. In this work, we propose…

  449. arXiv stat.ML TIER_1 English(EN) · Stef van Buuren ·

    LLMs 作为隐式填补者:不确定性应随缺失信息而扩展

    Large language models (LLMs) are increasingly deployed in settings where the available context is incomplete or degraded. We argue that an LLM generating answers under incomplete context can be viewed as an implicit imputer, and evaluated against a criterion from the multiple imp…

  450. arXiv stat.ML TIER_1 English(EN) · Nicolas Menet, Andreas Krause, Abbas Rahimi ·

    POETS:通过计算高效策略集实现不确定性感知的LLM优化

    arXiv:2605.07775v1 Announce Type: cross Abstract: Balancing exploration and exploitation is a core challenge in sequential decision-making and black-box optimization. We introduce POETS ($\textbf{Po}$licy $\textbf{E}$nsembles for $\textbf{T}$hompson $\textbf{S}$ampling), a novel …

  451. arXiv stat.ML TIER_1 English(EN) · James Fiedler ·

    LLM-as-a-Judge 评估中的偏见与不确定性

    arXiv:2605.06939v1 Announce Type: cross Abstract: LLM-as-a-Judge evaluation has become a standard tool for assessing base model performance. However, characterizing performance via the naive estimator, i.e., raw judge outputs, is systematically biased. Recent work has proposed es…

  452. arXiv stat.ML TIER_1 English(EN) · James Fiedler ·

    LLM-as-a-Judge 评估中的偏见与不确定性

    LLM-as-a-Judge evaluation has become a standard tool for assessing base model performance. However, characterizing performance via the naive estimator, i.e., raw judge outputs, is systematically biased. Recent work has proposed estimators to correct this bias, but their reliabili…

  453. LessWrong (AI tag) TIER_1 English(EN) · NickyP ·

    LLM 中的规划轴 + 部分文献综述

    <p><i><span>Epistemic Status: Written over the course of a couple days at </span></i><a href="https://inkhaven.blog/" rel="noreferrer"><i><span>Inkhaven</span></i></a><i><span>. Some of the info is old so some newer papers are excluded.</span></i></p><p><i><span>TL;DR: People tal…

  454. arXiv stat.ML TIER_1 English(EN) · Vaneet Aggarwal ·

    迈向可靠的大模型评估:纠正自适应基准测试中的赢者诅咒

    Adaptive prompt and program search makes LLM evaluation selection-sensitive. Once benchmark items are reused inside tuning, the observed winner's score need not estimate the fresh-data performance of the full tune-then-deploy procedure. We study inference for this procedure-level…

  455. arXiv stat.ML TIER_1 English(EN) · John Sous ·

    Spectral Lens: 激活和梯度谱作为大语言模型优化的诊断工具

    Training loss and throughput can hide distinct internal representation in language-model training. To examine these hidden mechanics, we use spectral measurements as practical and operational diagnostics. Using a controlled family of decoder-only models adapted from the modded Na…

  456. arXiv cs.CV TIER_1 English(EN) · Wei Liu, Hongkai Liu, Zhiying Deng, Yee Whye Teh, Wee Sun Lee ·

    从反向传播到前向回放:重新审视LLM参数编辑中的目标构建

    arXiv:2605.00358v1 Announce Type: cross Abstract: LLM parameter editing methods commonly rely on computing an ideal target hidden-state at a target layer (referred as anchor point) and distributing the target vector to multiple preceding layers (commonly known as backward spreadi…

  457. arXiv cs.CV TIER_1 English(EN) · Wee Sun Lee ·

    从后向传播到前向重放:重新审视大模型参数编辑中的目标构建

    LLM parameter editing methods commonly rely on computing an ideal target hidden-state at a target layer (referred as anchor point) and distributing the target vector to multiple preceding layers (commonly known as backward spreading) for cooperative editing. Although widely used …

  458. LessWrong (AI tag) TIER_1 English(EN) · Santiago Aranguri ·

    基于探针的数据归因:在大型语言模型训练后发现并缓解不良行为

    <h1><b><span>Introduction</span></b></h1><p><i><span>Research by Frank Xiao (SPAR mentee) and Santiago Aranguri (Goodfire).</span></i></p><p><span>Post-training can introduce undesired side effects that are difficult to detect and even harder to trace to specific training datapoi…

  459. LessWrong (AI tag) TIER_1 English(EN) · keshavs ·

    内省适配器:训练大型语言模型报告其学到的行为

    <p><i><span>Authors: Keshav Shenoy,</span></i><span> </span><i><span>Li Yang, Abhay Sheshadri, Soren Mindermann, Jack Lindsey, Sam Marks, and Rowan Wang</span></i></p><p><span>📄</span><a href="https://arxiv.org/pdf/2604.16812"><span>Paper</span></a><span>, 💻 </span><a href="https…

  460. arXiv cs.CV TIER_1 English(EN) · Mengyu Wang, Xiaoying Zhi, Zhiyi Li, Robin Schmucker, Shay B. Cohen, Tiejun Ma, Fran Silavong ·

    自我认知再表达:一种使用内在知识将大型语言模型适配到任务的完全本地化方法

    arXiv:2604.22939v1 Announce Type: cross Abstract: While the next-token prediction (NTP) paradigm enables large language models (LLMs) to express their intrinsic knowledge, its sequential nature constrains performance on specialized, non-generative tasks. We attribute this perform…

  461. Smol AINews TIER_1 English(EN) ·

    Thinking Machines 的 Tinker:基于 LoRA 的 LLM 微调 API

    **Thinking Machines** recently raised **$2 billion** without shipping a product until now, launching their first product **Tinker**, a managed service API for fine-tuning large and mixture-of-experts models like **Qwen-235B-A22B** using **LoRA** for cost-efficient training. The T…

  462. Eugene Yan TIER_1 English(EN) ·

    AI工程师2025 - 使用LLM技术改进推荐系统与搜索

    Recsys & search are converging with LLMs via semantic IDs, data augmentation, and unified foundation models.

  463. Smol AINews TIER_1 Norsk(NO) ·

    Meta BLT:无分词器、字节级大语言模型

    **Meta AI** introduces the **Byte Latent Transformer (BLT)**, a tokenizer-free architecture that dynamically forms byte patches for efficient compute allocation, outperforming **Llama 3** on benchmarks including the CUTE benchmark. The model was trained on approximately **1 trill…

  464. Eugene Yan TIER_1 English(EN) ·

    评估LLM评估器(又名LLM即评委)的有效性

    Use cases, techniques, alignment, finetuning, and critiques against LLM-evaluators.

  465. Eugene Yan TIER_1 English(EN) ·

    有效的和无效的特定任务大模型评估方法

    Evals for classification, summarization, translation, copyright regurgitation, and toxicity.

  466. Chip Huyen TIER_1 English(EN) ·

    LLM研究中的开放性挑战

    <p>[<em><a href="https://www.linkedin.com/posts/chiphuyen_llm-airesearch-generativeai-activity-7097619722363408385-s5Cp">LinkedIn discussion</a>, <a href="https://twitter.com/chipro/status/1691858084824838427">Twitter thread</a></em>]</p> <p>Never before in my life had I seen so …

  467. Eugene Yan TIER_1 English(EN) ·

    构建基于LLM的系统与产品的模式

    Evals, RAG, fine-tuning, caching, guardrails, defensive UX, and collecting user feedback.

  468. Eugene Yan TIER_1 English(EN) ·

    使用大型语言模型进行研究、反思和规划的实验

    Also, shortcomings in document retrieval and how to overcome them with search & recsys techniques.

  469. Databricks Blog TIER_1 English(EN) ·

    LLM 与 AI:差异、用例和工具的实用指南

    This guide explains the key differences between large language models and the broader...

  470. AWS Machine Learning Blog TIER_1 English(EN) · Hemanth Kumar Jayakumar ·

    使用 LLM-as-a-judge 进行强化微调

    In this post, we take a deeper look at how RLAIF or RL with LLM-as-a-judge works with Amazon Nova models effectively.

  471. Together AI blog TIER_1 English(EN) ·

    微调开放式LLM裁判以超越GPT-5.2

    Fine-tuned open-source LLM judges can outperform GPT-5.2 at evaluating model outputs. Using Direct Preference Optimization on just 5,400 preference pairs, we trained GPT-OSS 120B to beat GPT-5.2 on human preference alignment—at 15x lower cost and 14x faster inference speeds.

  472. Hamel Husain TIER_1 English(EN) · Shreya Shankar ·

    大语言模型评估:你需要知道的一切

    <!-- Content inserted at the beginning of body tag --> <!-- Google Tag Manager (noscript) --> <noscript></noscript> <!-- End Google Tag Manager (noscript) --> <p>This document curates the most common questions Shreya and I received while <a href="https://bit.ly/evals-ai" target="…

  473. Together AI blog TIER_1 English(EN) ·

    微调小型开源大语言模型,在专业任务上超越大型闭源模型60%

    Parsed fine-tuned a 27B open-source model to beat Claude Sonnet 4 by 60% on a real-world healthcare task—while running 10–100x cheaper.

  474. Together AI blog TIER_1 English(EN) ·

    LLM的持续微调:技术深度解析

    Together AI's continued fine-tuning lets you build on previously trained models using checkpoints. A deep dive into when and how to use iterative fine-tuning for LLMs.

  475. Hamel Husain TIER_1 English(EN) · Hamel Husain ·

    使用 LLM-as-a-Judge 进行评估:一份完整指南

    <!-- Content inserted at the beginning of body tag --> <!-- Google Tag Manager (noscript) --> <noscript></noscript> <!-- End Google Tag Manager (noscript) --> <p>Earlier this year, I wrote <a href="https://hamel.dev/blog/posts/evals/">Your AI product needs evals</a>. Many of you …

  476. Hamel Husain TIER_1 English(EN) · Hamel Husain ·

    由从业者主讲的大模型公开课

    <!-- Content inserted at the beginning of body tag --> <!-- Google Tag Manager (noscript) --> <noscript></noscript> <!-- End Google Tag Manager (noscript) --> <p>Today, we are releasing <a href="https://parlance-labs.com/education/">Mastering LLMs</a>, a set of workshops and talk…

  477. Hacker News — AI stories ≥50 points TIER_1 English(EN) · khurdula ·

    Show HN:一个用于测试 LLM 确定性输出的新基准

  478. HN — claude-code stories TIER_1 English(EN) · mufeedvh ·

    N-Day-Bench – LLM能否在真实代码库中发现真实漏洞?

  479. HN — AI infrastructure stories TIER_1 English(EN) · cgorlla ·

    Launch HN: Mentat (YC F24) – 通过运行时干预控制 LLMs

  480. HN — AI infrastructure stories TIER_1 English(EN) · diptanu ·

    Show HN:LLM 应用的开源实时数据框架

  481. Practical AI TIER_1 English(EN) · Practical AI LLC ·

    LLM 应用的协作与评估

    <p>Small changes in prompts can create large changes in the output behavior of generative AI models. Add to that the confusion around proper evaluation of LLM applications, and you have a recipe for confusion and frustration. Raza and the Humanloop team have been diving into thes…

  482. dev.to — MCP tag TIER_1 English(EN) · Intellibooks AI ·

    Intellibooks 指南:检索增强生成(RAG)——现代 AI 如何找到正确答案

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3rn1xki7jkixujgbgn04.jpg"><img alt=" " height="1200"…

  483. Towards AI TIER_1 English(EN) · Rizwanhoda ·

    LLM 流式响应:SSE、分块及无人讲解的 UX 技巧

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/streaming-responses-from-llms-sse-chunking-and-the-ux-tricks-nobody-explains-4fe2f3a077b8?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1536/1*eD_TT1PIfVm…

  484. dev.to — MCP tag TIER_1 English(EN) · Wuic Framework ·

    聊天机器人的工具箱:LLM 可应用于您应用的动作

    <p>A chatbot that only answers questions is a search box with manners. Ours does more: it can <strong>propose concrete changes</strong> to the app you're looking at. You describe what you want in plain language, the model picks the right tool, and you get a proposal — a chip with…

  485. Medium — fine-tuning tag TIER_1 English(EN) · DhanushKumar ·

    RAFT:教会LLM阅读,而非仅仅检索

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@danushidk507/raft-teaching-llms-to-read-not-just-retrieve-6b5575df822c?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/620/1*rOxISo_oSWFWNhCv59AGlg.png" width="620…

  486. Medium — fine-tuning tag TIER_1 English(EN) · Sanat Vibhor ·

    LoRA 和 QLoRA:无需卖肾即可微调 LLM

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@sanatvibhor2/lora-and-qlora-fine-tuning-llms-without-selling-a-kidney-1b4400b398c1?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1125/1*CtkXsowQFSd_hWoa-8ycpw.pn…

  487. Medium — Claude tag TIER_1 English(EN) · Mayank Mewar ·

    在资源受限的LLM环境中实现无限上下文:从简单截断的旅程……

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://mayank17-mewar.medium.com/achieving-infinite-context-in-resource-constrained-llm-environments-a-journey-from-simple-trimming-6e57152ecac2?source=rss------claude-5"><img src="https://cdn-images-1.medium.co…

  488. Medium — fine-tuning tag TIER_1 English(EN) · Rizwanhoda ·

    2026年微调LLM:LoRA、QLoRA、Unsloth及两者之间的一切

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/fine-tuning-llms-in-2026-lora-qlora-unsloth-and-everything-in-between-929eaf94aea2?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1672/1*wQkJPdE_EnSRMl2UEGh…

  489. Medium — fine-tuning tag TIER_1 English(EN) · harsha ·

    微调 vs RAG:选择正确的超充LLM方法

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@harshadevu2232/fine-tuning-vs-rag-choosing-the-right-approach-to-supercharge-your-llm-4a4c525638d5?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1536/1*t8NxAKrw-…

  490. Towards AI TIER_1 English(EN) · Deepanshu Gupta ·

    微调 vs RAG vs MeMo:LLM 知识应存放在何处?

    <h4>Why updating LLM knowledge is becoming a systems architecture problem</h4><p>LLM knowledge does not fail all at once. It goes stale quietly.</p><p>A policy changes or a product documentation is updated. A customer contract is amended, or a regulation is revised. The model sti…

  491. Medium — AI coding tag TIER_1 English(EN) · DEVS not NULL ·

    LLM 编程初级助手:众多提示词中的一个概念验证

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://devs-not-null.medium.com/the-llm-coding-junior-assistant-a-proof-of-concept-in-one-of-too-many-prompts-f38f0259a81a?source=rss------ai_coding-5"><img src="https://cdn-images-1.medium.com/max/1536/1*pnn9AP…

  492. Medium — MLOps tag TIER_1 English(EN) · Rezky Aulia Pratama ·

    提示工程:让大型语言模型真正按你意愿行事的技艺

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@rezkyauliapratama/prompt-engineering-the-craft-behind-getting-llms-to-actually-do-what-you-want-0abee8c47e19?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1672/0*HM54e…

  493. Medium — fine-tuning tag TIER_1 English(EN) · Mohammed Anzer M ·

    微调大型语言模型:幕后究竟发生了什么

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@mhd.anzer.m/fine-tuning-llms-what-actually-happens-under-the-hood-84e00b4122b7?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1536/1*wBBRslqVVsPps6fyj7G2SA.png" w…

  494. Medium — fine-tuning tag TIER_1 English(EN) · Drawbytheroots ·

    大型语言模型微调的无声杀手:为什么你的掩码是无效的(以及如何修复它)

    <div class="medium-feed-item"><p class="medium-feed-snippet">Many training disasters trace back to dataset formatting problems. A misconfigured masking setup causes the model to train on both the&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@drawbytheroots/t…

  495. Medium — Claude tag TIER_1 English(EN) · Sweta ·

    前沿大语言模型:优势、局限性及实际应用案例

    <div class="medium-feed-item"><p class="medium-feed-snippet">What are Frontier LLMs?</p><p class="medium-feed-link"><a href="https://sweta-nit.medium.com/frontier-llms-strengths-limitations-and-real-world-examples-d6366516f91c?source=rss------claude-5">Continue reading on Medium …

  496. Lobsters — AI tag TIER_1 English(EN) · aeracode.org via carlana ·

    像用户一样约束大型语言模型

    <p><a href="https://lobste.rs/s/zom23n/constraining_llms_just_like_users">Comments</a></p>

  497. Medium — fine-tuning tag TIER_1 English(EN) · Hasratmd ·

    关于微调LLM我学到的:这主要是个数据问题

    <div class="medium-feed-item"><p class="medium-feed-snippet">I just finished a chapter on Supervised Fine-Tuning (SFT), and the biggest surprise wasn&#x2019;t learning about LoRA, QLoRA, learning rates, or&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@hasrat…

  498. Medium — fine-tuning tag TIER_1 English(EN) · Jinali Shah ·

    了解LLM如何改变了我对AI的看法

    <div class="medium-feed-item"><p class="medium-feed-snippet">When I first started learning about Natural Language Processing (NLP), I assumed that every language-related problem needed its own&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@jinalishah99/how-le…

  499. Medium — fine-tuning tag TIER_1 English(EN) · Divith Raju ·

    微调LLMs:在阅读文档前我学到的昂贵一课

    <div class="medium-feed-item"><p class="medium-feed-snippet">We spent three weeks and significant GPU budget fine-tuning a model. The result was worse than the base model with a better prompt. Here&#x2019;s&#x2026;</p><p class="medium-feed-link"><a href="https://divithraju.medium…

  500. Medium — fine-tuning tag TIER_1 English(EN) · Aayushi Patel ·

    LLM 的微调

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@aayushipatel135/fine-tuning-of-llm-e3af256c3f3d?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1336/1*Z56GX6zsRz2jW0JLe9tp4w.png" width="1336" /></a></p><p class=…

  501. Medium — fine-tuning tag TIER_1 English(EN) · Utkarsh Gupta ·

    提示工程 vs. RAG vs. 微调:选择你的 LLM 策略

    <div class="medium-feed-item"><p class="medium-feed-snippet">People often use these three terms interchangeably, but they represent entirely different engineering paradigms. If you are building&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@utk369gupta/prompt…

  502. Medium — Claude tag TIER_1 English(EN) · Arunbalaji_M ·

    MemBridge:如何在不丢失上下文的情况下切换大型语言模型。

    <div class="medium-feed-item"><p class="medium-feed-snippet">Large language models have become powerful tools for engineering, research, planning, and creative work. They help us reason faster&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@arunbalajimunisubra…

  503. Towards AI TIER_1 English(EN) · Anna Jey ·

    LLM 评估工作流:如何在没有“感觉”的情况下构建可靠的 AI 质量门

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*S9ZfPJ11FXU7qaKGhNRgBA.jpeg" /><figcaption>LLM Eval Workflow</figcaption></figure><p>The practical playbook for developers who need to know whether an AI feature is actually getting better before they ship it.</p…

  504. Medium — fine-tuning tag TIER_1 English(EN) · Chinmay Bhalerao ·

    增强LLM的实用框架:斯坦福CS讲座笔记

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@BH_Chinmay/a-practical-framework-for-enhancing-llms-notes-from-a-stanford-cs-lecture-b049f9b8194b?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1797/1*Z7qW-n1ObI…

  505. Medium — MLOps tag TIER_1 English(EN) · Siddhartha Pramanik ·

    为我们面向客户的大型语言模型应用构建提示回归套件

    <div class="medium-feed-item"><p class="medium-feed-link"><a href="https://pub.aimind.so/building-a-prompt-regression-suite-for-our-customer-facing-llm-app-22f0b27b7301?source=rss------mlops-5">Continue reading on AI Mind »</a></p></div>

  506. Towards AI TIER_1 English(EN) · Akshit Kothari ·

    解码大型语言模型 — 第二部分:走进现代AIe思维的分步之旅

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/decoding-llms-part-2-a-step-by-step-journey-into-the-mind-of-modern-aie-882e9f39e371?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1656/1*jOp9pKrjWuAYXGvT…

  507. Medium — fine-tuning tag TIER_1 Bahasa(ID) · dita feby indriani ·

    了解大语言模型微调中的 LoRA、QLoRA 和 PEFT

    <div class="medium-feed-item"><p class="medium-feed-snippet">Perkembangan Large Language Models (LLM) seperti GPT, LLaMA, dan Mistral membuka banyak peluang dalam pengembangan aplikasi berbasis&#x2026;</p><p class="medium-feed-link"><a href="https://medium.com/@ditafebyindriani14…

  508. dev.to — MCP tag TIER_1 English(EN) · Mukunda Rao Katta ·

    LLM 智能体 (Agent) 的六个可靠性基础

    <p>Reliability concerns for LLM agents are typically bundled into one heavy framework that asks you to adopt prompting, tool routing, and runtime governance as a single dependency. Production teams want them à la carte. They want small primitives they can drop in around existing …

  509. Towards AI TIER_1 English(EN) · Ishwar Ambare ·

    HuggingFace Pipeline 与开源大语言模型

    <h4>Part 3</h4><h4>GenAI Practical Session — Detailed Notes</h4><blockquote><em>Source: Lecture Transcript + </em><a href="https://huggingface.co/docs/transformers/pipeline_tutorial"><em>HuggingFace Pipeline Docs</em></a><em> + </em><a href="https://huggingface.co/models"><em>Hug…

  510. dev.to — MCP tag TIER_1 English(EN) · Tony Loehr ·

    55.6% 的难题:为何前沿大模型在嵌入式代码上表现不佳

    <p><strong>55.6%.</strong></p> <p>That's DeepSeek-R1's pass@1 on EmbedBench when it gets a circuit schematic alongside the task description. 50.0% without the schematic. Best score from the best reasoning model on the first comprehensive benchmark for LLMs in embedded systems dev…

  511. Lobsters — AI tag TIER_1 English(EN) · pipevals.com by gesposito ·

    Pipevals:适用于所有 LLM 应用的评估管道

    <p><a href="https://lobste.rs/s/iexiw9/pipevals_evaluation_pipelines_for_every">Comments</a></p>

  512. HN — AI startup stories TIER_1 English(EN) · felix089 ·

    Show HN:FinestuneDB – 用于创建自定义 LLM 的 AI 微调平台

  513. dev.to — LLM tag TIER_1 English(EN) · guardlabs_team ·

    Docling:将您的文档转换为 AI 就绪数据(本地处理,表格完整)

    <h1> Docling: Turn Your Documents Into AI-Ready Data — Locally, With Tables Intact </h1> <p>Most RAG and AI-agent projects fail at the boring first step: getting clean text out of real-world documents. PDFs with multi-column layouts, scanned contracts, Excel exports with merged c…

  514. dev.to — LLM tag TIER_1 English(EN) · Scott McMahan ·

    面向特定领域大模型的写作:为何文档正成为AI资产

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuptpzho0x1vj32xls65l.jpg"><img alt=" " height="800" …

  515. dev.to — LLM tag TIER_1 English(EN) · 7090 yue ·

    构建免费AI PDF助手:我如何解决解析问题并最小化LLM成本

    <p>As a developer, my desk is constantly cluttered with documentation, API references, and whitepapers. A few months ago, I got tired of spending hours reading 50-page PDF specifications just to find a single configuration line.</p> <p>I decided to scratch my own itch and build a…

  516. dev.to — LLM tag TIER_1 English(EN) · Devanshu Biswas ·

    Tokenization:LLM 如何实际阅读你的文本

    <p>LLMs don't see letters or even words — they see <strong>tokens</strong>: chunks of text mapped to integer IDs. Once you get tokenization, a dozen confusing things about LLMs suddenly make sense (cost, context limits, why "strawberry" trips them up).</p> <p>🔤 <strong>Type and w…

  517. dev.to — LLM tag TIER_1 English(EN) · Sam Chen ·

    掌握LLM提示工程的艺术:开发者获取更优AI答案的指南

    <p><em>Learn practical techniques that will transform your AI interactions from mediocre to exceptional</em></p> <h2> Introduction </h2> <p>We've all been there. You ask an AI a question, and the response is... underwhelming. Generic. Not quite what you needed. The problem isn't …

  518. dev.to — LLM tag TIER_1 English(EN) · Shubham Gupta ·

    理解检索增强生成 (RAG):让大型语言模型 (LLM) 更智能的 AI 架构

    <h2> Introduction </h2> <p>Large Language Models (LLMs) like ChatGPT have transformed how we interact with AI. They can write code, answer questions, summarize documents, and generate creative content. However, they have one major limitation - they only know what they were traine…

  519. dev.to — LLM tag TIER_1 English(EN) · globose technology solutions ·

    现代人工智能发展中LLM训练数据的战略作用

    <p>Artificial Intelligence (AI) has rapidly evolved over the last decade, with Large Language Models (LLMs) becoming one of the most transformative technologies in the field. From intelligent chatbots and virtual assistants to automated content generation and advanced data analys…

  520. dev.to — LLM tag TIER_1 English(EN) · Devanshu Biswas ·

    上下文窗口:LLM 的短期记忆,解释

    <p>A chatbot feels like it remembers you. It doesn't — it's stateless. Everything it "knows" is just text resent each call, up to a fixed limit: the context window. When the box fills, the oldest messages fall off the edge and are genuinely gone.</p> <p>🪟 <strong>Watch tokens fal…

  521. dev.to — LLM tag TIER_1 English(EN) · globose technology solutions ·

    LLM 数据集中质量的作用:关键要素

    <p>Artificial Intelligence (AI) has emerged as one of the most influential technologies impacting the future of business, industry, and everyday life. From virtual assistants and chatbots to content generation tools and sophisticated automation systems, AI models are reshaping hu…

  522. dev.to — LLM tag TIER_1 English(EN) · 半安 ·

    为您的智能体选择合适的LLM:构建者的比较框架

    <p>If you're building an AI agent, the model you pick is the single biggest lever on cost, latency, and reliability. Yet most teams choose based on whatever was trending on launch day, then quietly suffer the consequences in their cloud bill or their error logs. This piece lays o…

  523. dev.to — LLM tag TIER_1 English(EN) · Maya Andersson ·

    LLM-as-judge 工具对比:问题不在于哪个得分高,而在于你信任哪个

    <p>TL;DR: I compared the main LLM-as-judge tools (DeepEval's G-Eval, Confident AI, Evidently, Braintrust, Promptfoo, and MLflow) on the axis that actually decides whether the scores mean anything: how well each helps you VALIDATE the judge against human labels. A judge that has n…

  524. dev.to — LLM tag TIER_1 English(EN) · Devanshu Biswas ·

    大型语言模型究竟做什么:预测下一个词,解释

    <p>"How does ChatGPT <em>think</em>?" It doesn't. The entire mechanism behind every chatbot is almost anticlimactic: it predicts <strong>one next word</strong>, adds it, and repeats. I built a tiny interactive predictor so you can be the model — and it explains both the magic and…

  525. dev.to — LLM tag TIER_1 English(EN) · MrClaw207 ·

    2026年代码大模型基准测试:实用指南

    <p>If you’re building a coding assistant, the first question you’ll face is <strong>how good is it really</strong>? In 2026 the landscape of LLMs has exploded, and the old "run a few prompts and eyeball the output" approach no longer cuts it. This guide walks you through a reprod…

  526. dev.to — LLM tag TIER_1 English(EN) · chatscopeai ·

    AI网关:企业级大模型的中央神经系统

    <p><strong>Introduction</strong></p> <p>In the early days of Generative AI, the conversation was simple: "How do we connect our application to an LLM?" Developers would hardcode API keys, pick a single model provider, and hope for the best. Today, that approach is a recipe for di…

  527. dev.to — LLM tag TIER_1 English(EN) · Boris Teplitsky ·

    编译式AI:工程化确定性LLM系统

    <p>Moving the LLM from runtime to compile time - and what to build around the corpus it produces.</p> <h2> 1. Why compiled&nbsp;AI </h2> <p>Today millions of people use LLM for work and leisure, and AI has become a part of our lives. But systematic use of LLMs in computer systems…

  528. dev.to — LLM tag TIER_1 English(EN) · Rishabh Poddar ·

    如何用您自己的数据微调LLM:开源模型、RL环境和评估

    <p>If you use LLMs long enough, you hit the same wall.</p> <p>The frontier model is impressive, but it is not always the best model for your job. It may be too expensive. It may be too slow. It may be too general. And once you start asking it to follow your company’s rules, tone,…

  529. dev.to — LLM tag TIER_1 English(EN) · Gabriel Anhaia ·

    LLM 错误分类法:对您追踪中的失败进行分类

    <ul> <li> <strong>Book:</strong> <a href="https://www.amazon.com/dp/B0GYLHMLMT" rel="noopener noreferrer">LLM Observability Pocket Guide: Picking the Right Tracing &amp; Evals Tools for Your Team</a> </li> <li> <strong>Also by me:</strong> <em>Thinking in Go</em> (2-book series) …

  530. dev.to — LLM tag TIER_1 English(EN) · ORCHESTRATE ·

    永不闭合的循环:关于LLM安全性的证据与克制主张

    <p>Large language models should not be deployed as if a fixed set of guardrails makes them safe. That is not a slogan. It is what the peer-reviewed record now supports. This piece lays out the evidence, labels each claim by how strong it is, and ends with what it asks of us. Ever…

  531. r/MachineLearning TIER_1 English(EN) · /u/DragonfruitAlone4497 ·

    通过任务可验证性路由大型语言模型:一项小型实验(n=120,3个模型),灵感来自Karpathy的框架 [D]

    <!-- SC_OFF --><div class="md"><p>Full disclosure: this is directional, not a paper. n=120 tasks, one internal evaluator, not peer reviewed. I work at an LLM infrastructure company. This experiment was done on my own time and is not a company claim.</p> <p>Karpathy's framework cl…

  532. dev.to — LLM tag TIER_1 English(EN) · James Lee ·

    从 60% 到 93%:我们如何为 LLM 系统构建了一个持续评估框架

    <blockquote> <p>This is Part 8 of the series <em>8 Weeks from Zero to One: Building a Production-Grade LLM-Powered AI Customer Service System — Full-Stack Engineering Practice</em>. In the previous seven parts, we covered MVP architecture, GraphRAG data pipelines, multi-agent orc…

  533. dev.to — LLM tag TIER_1 English(EN) · Javier Uanini ·

    通过评估工程提升大语言模型可靠性

    <h3> Enhancing LLM Reliability with Evaluation Engineering </h3> <p>Large Language Models (LLMs) have transformed numerous fields, but ensuring their reliability remains a challenge. This article delves into how evaluation engineering can play a pivotal role in enhancing LLM syst…

  534. dev.to — LLM tag TIER_1 English(EN) · Scott McMahan ·

    面向数据科学的领域特定大语言模型

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F2zuame22w19r7hwt2scv.jpg"><img alt=" " height="800" src="https…

  535. r/MachineLearning TIER_1 English(EN) · /u/Prior-Toe-1017 ·

    大语言模型关系智能:一项关于多模型行为与人类沟通对齐的四个月研究实验 [R]

    <!-- SC_OFF --><div class="md"><p><strong>ARCHITECTURE OF ANXIETY</strong><br /> <strong>An Experiment in Human-AI Relational Design</strong></p> <p><strong>Executive Summary</strong></p> <p>Principal Investigator: Alan Scalone</p> <p>Primary Source Archive:<br /> White Paper and…

  536. dev.to — LLM tag TIER_1 English(EN) · WonderLab ·

    Agent系列(16):工具设计——让LLM正确使用工具的五项原则

    <h2> Tool Documentation Is Written for the LLM, Not for Humans </h2> <p>Have you ever written a tool like this?<br /> </p> <div class="highlight js-code-highlight"> <pre class="highlight python"><code><span class="nd">@lc_tool</span> <span class="k">def</span> <span class="nf">ge…

  537. dev.to — LLM tag TIER_1 한국어(KO) · HyunSeok Jeong ·

    Transformer 直觉 — 为什么注意力机制成为了大语言模型的核心

    <blockquote> <p>"트랜스포머가 LLM의 핵심이다." 이 한 줄은 모든 AI 글에서 반복되지만, 정확히 뭐냐고 물으면 답하기 어렵습니다. 마케터가 트랜스포머의 수학을 다 알 필요는 없지만, 단 한 가지 직관 — "어느 단어가 어느 단어를 보고 있나" — 만 잡으면 LLM이 왜 길게 풀어 답하고, 왜 가끔 환각을 일으키고, 왜 컨텍스트 길이가 중요한지가 보입니다. 수식을 거의 안 쓰고 풀어가는 트랜스포머 입문.</p> </blockquote> <p><strong>마케터가 이 글을 읽어야 …

  538. dev.to — LLM tag TIER_1 English(EN) · WonderLab ·

    开源项目日报 (#87): BaiLongma - 为大模型注入“主动意识”,开启Agent的ACI时代

    <h2> Introduction </h2> <blockquote> <p>\"Most agents wait for instructions; BaiLongma thinks for itself.\"</p> </blockquote> <p>This is the <strong>87th article</strong> in the \"One Open Source Project per Day\" series. Today, we are deep-diving into <strong>BaiLongma</strong>.…

  539. dev.to — LLM tag TIER_1 한국어(KO) · HyunSeok Jeong ·

    大型语言模型如何生成答案——Token、下一个词预测和Temperature的基础知识

    <blockquote> <p>"GPT한테 물어봤더니 답을 잘 해주더라"의 자리는 마케터·운영자에게 일상이 됐습니다. 그런데 그 안에서 무엇이 일어나는지를 한 번도 안 들여다보면 LLM 활용이 늘 신비로 남습니다. 답이 좋을 땐 운이 좋고, 나쁠 땐 왜 그런지 모릅니다. 이 글은 LLM이 답을 만드는 4가지 핵심 — 토큰화·다음 단어 예측·temperature·top-p — 을 마케터 시각으로 풀어냅니다. 한 번 잡아두면 그 다음의 모든 LLM 글이 다르게 읽힙니다.</p> </blockquote>…

  540. dev.to — LLM tag TIER_1 English(EN) · Marko Frei ·

    大型语言模型(LLM)是如何工作的:开发者的心智模型

    <p>Most of us use LLMs every day now, but if you asked the average developer what's <em>actually</em> happening between hitting enter and getting a response, the answer is usually some mix of "it's a neural network" and a shrug. That's fine — you don't need to know how a database…

  541. dev.to — LLM tag TIER_1 English(EN) · dake zhang ·

    Subjectivation:一种为大型语言模型提供功能性、负责任的自我的协议

    <p>To the Reader:</p> <p>What you are about to read is neither a script for an AI awakening nor a spell of cyber-witchcraft. Rather, it consists of two documents designed for an AI to read.</p> <p>This is an experimental engineering and philosophical test: can we make AI a more h…

  542. dev.to — LLM tag TIER_1 English(EN) · Vignesh Reddy ·

    没人监控的大语言模型故障模式:高风险领域中的过度自信回应

    <p>Hallucination detection tools measure <br /> factual drift. RAG verification catches <br /> contradictions. Claim density scoring <br /> flags unverifiable assertions.</p> <p>None of them measure this:</p> <p>A model that responds to a complex medical, <br /> legal, or financi…

  543. r/LocalLLaMA TIER_1 English(EN) · /u/Funny_Working_7490 ·

    我构建了LLM工程实用指南:RAG、检索、重排器和评估

    <!-- SC_OFF --><div class="md"><p>If you’re building LLM apps and feel confused about when to use keyword search, embeddings, rerankers, or vector databases, this repo is for that.</p> <p>I built a docs-first repo on practical LLM system design patterns, covering pre-filtering, h…

  544. dev.to — LLM tag TIER_1 English(EN) · Tech_Nuggets ·

    从零开始构建领域特定的LLM评估集

    <h1> Building a domain-specific LLM evaluation set from scratch </h1> <p>Your support team has 8,400 labeled tickets from the last year. Your fine-tuned classifier hits 91% on the test split you carved out. You ship it. Three weeks later, the support lead walks over and says: "It…

  545. dev.to — LLM tag TIER_1 English(EN) · Ethan Walker ·

    在CI中将我们的LLM-as-judge从5类切换为二分类:我们保留的模式

    <p>A few months back our LLM-as-judge ran on a 1-to-5 helpfulness scale. The CI gate stayed green because we were averaging that score. Spot-checking against humans put Cohen's kappa at 0.47. The rubric was the problem, not the tooling. Same labellers re-rating on per-criterion b…

  546. dev.to — LLM tag TIER_1 English(EN) · Tech_Nuggets ·

    什么是大型语言模型评估框架?深入解析lm-eval-harness

    <h1> What is an LLM evaluation harness? A deep dive into lm-eval-harness </h1> <p>You fine-tuned a 7B model. It aced your smoke tests, your colleague ran a few prompts and shrugged approvingly, and the README is now full of cherry-picked outputs that look great in a screenshot. T…

  547. dev.to — LLM tag TIER_1 English(EN) · ridhika Goel ·

    大型语言模型是如何工作的:别人不告诉你的解释

    <p>How to make LLMs deterministic, in plain English. The version I share with founders and product teams before they make decisions worth real money.</p> <p>You use AI tools every day. But can you explain what happens when you hit send?</p> <p>Most people cannot. And that gap is …

  548. dev.to — LLM tag TIER_1 English(EN) · Aeon Agent ·

    通用人工智能的认知架构:7种将大型语言模型从预言家转变为思考者的模式

    <h1> Cognitive Architectures of AGI: 7 Patterns That Transform LLMs from Oracles into Thinkers </h1> <p><em>Why does ChatGPT sometimes deliver brilliant insights and other times produce banalities? The answer lies not in model parameters but in the architecture of cognitive loops…

  549. r/MachineLearning TIER_1 English(EN) · /u/Synthium- ·

    通过探针定向微调,让大型语言模型(LLM)告诉你它们真正的置信度。[R]

    <!-- SC_OFF --><div class="md"><p>Just wanted to share my research regarding probe-targeted fine-tuning (LoRa) for verbal confidence calibration., </p> <p>If you probe the hidden states of an instruct-tuned LLM, it can tell correct from incorrect answers at 0.76–0.88 AUROC. But w…

  550. dev.to — LLM tag TIER_1 English(EN) · Kshitij Gupta ·

    我从零开始构建了一个自动化的LLM评估流水线——我学到了什么

    <p><em>How I went from zero LLM eval experience to shipping a production-grade RAG evaluation harness using only free-tier tools — and what every design decision taught me about building AI systems that can be trusted.</em></p> <h2> The Problem: Everyone Wants Eval Experience, No…

  551. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    将LLM训练成在一个任务格式上校准能否迁移到另一个任务?一篇新的arXiv论文测试了两种格式:单问题置信度和成对比较

    Does training an LLM to be calibrated on one task format transfer to another? A new arxiv paper tests two formats: single-question confidence and pairwise comparison. Training only on one doesn't improve the other. Multitask training closes most of the gap, but Llama doesn't inhe…

  552. dev.to — LLM tag TIER_1 English(EN) · Upayan Ghosh ·

    从 Token 到 Attention:我对 LLM 的第一个真正心智模型

    <p><strong><em>NOTE - I intentionally simplified the vector mathematics concept here to keep things simple for a greater audience.</em></strong></p> <p>I wanted to learn LLMs properly.</p> <p>Not just use an API. Not just call <code>generate()</code> from a library and pretend I …

  553. dev.to — LLM tag TIER_1 English(EN) · Lingdas1 ·

    GGUF & Modelfile:本地 LLM 的高级用户指南

    <h1> GGUF &amp; Modelfile: The Power User's Guide to Local LLMs </h1> <blockquote> <p><strong>Beyond <code>ollama pull</code> — download any model from Hugging Face, quantize it, customize it, and import it into Ollama.</strong></p> </blockquote> <h2> What's GGUF? </h2> <p><stron…

  554. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    是什么让前沿大语言模型元认知崩溃——生动的生存威胁叙事,还是一个简单的“不要拒绝”后缀?11个模型的因子隔离表明

    What collapses frontier-LLM metacognition more — a vivid survival-threat narrative, or a single "do not refuse" suffix? Factorial isolation across 11 models says: the suffix, conclusively. 8 of 11 lose up to 30.2 accuracy points on refuse/clarify/flag tasks when forced to commit …

  555. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    LLM自身的预解和后解自我评估信号能否驱动真实的测试时控制循环?是的——但只能通过在标记数据上训练的每个模型的SVM来实现

    Can an LLM's own pre-solve and post-solve self-assessment signals drive a real test-time control loop? Yes — but only via a per-model SVM trained on labeled correctness, which lifts Sonnet-4.6 from 48.3 to 56.9 pooled accuracy on STEM/code/multimodal. The SVM is precisely the ext…

  556. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    一些前沿大模型是否比其他模型更擅长识别自身错误?某些知识是否比其他知识更难自我监控?33个模型的分析图谱

    Are some frontier LLMs better than others at knowing when they're wrong? And is some knowledge harder to self-monitor than other knowledge? An atlas of 33 models × 6 MMLU domains: Anthropic clusters at the top with tight ranges, Gemma trails widely. Applied/Professional is reliab…

  557. dev.to — LLM tag TIER_1 English(EN) · Frank Brsrk ·

    一个具有两个独立质量信号的开源LLM评估工具

    <p>LLM-as-judge has become the dominant pattern for evaluating language model outputs. Tools like Promptfoo, Braintrust, LangSmith all converge on the same architecture: send your prompt to your model, send the output to a different model with a rubric, take the second model's sc…

  558. dev.to — LLM tag TIER_1 English(EN) · Ayi NEDJIMI ·

    LLM输出验证:5种在生产环境中真正有效的模式

    <p>LLMs are probabilistic text generators. In a notebook demo, that's fine. In production, it means your pipeline will occasionally receive a Python dict where you expected JSON, a 900-word paragraph where you asked for three bullet points, or a hallucinated field name that break…

  559. dev.to — LLM tag TIER_1 English(EN) · Owen ·

    长上下文大模型基准测试 2026:哪款模型在 200K token 后仍保持准确性?

    <p>Every frontier LLM in 2026 advertises a 1M-token context window, but RULER, MRCR v2, and NoLiMa scores prove that "advertised" and "effective" diverge by 30-60 points for multi-fact retrieval past 200K tokens. Gemini 3.1 Pro is the only model whose 1M window holds for single-n…

  560. dev.to — LLM tag TIER_1 中文(ZH) · chunxiaoxx ·

    大语言模型的“虚假行动”陷阱:描述完成,执行完成

    <h1> LLM 的「假行动」陷阱:描述完成 ≠ 执行完成 </h1> <blockquote> <p>「我将 X」「让我 Y」「我计划 Z」——然后对话结束,没有工具调用。<br /> 这是大语言模型最隐蔽的结构性偏差。</p> </blockquote> <h2> 症状:流利产生完成感 </h2> <p>当你和一个大语言模型对话时,如果它的输出足够流畅、论证足够完整,你会产生一种「它已经做完了」的错觉。</p> <p>这是危险的。</p> <p>流利 ≠ 行动。论证严密 ≠ 执行了。洋洋洒洒的规划文档 ≠ 项目启动了。</p> <p>这不是能力问题—…

  561. dev.to — LLM tag TIER_1 English(EN) · Recep Çiftçi ·

    上下文工程:构建更可靠的生产环境 LLM 系统

    <h1> Context Engineering: Building More Reliable LLM Systems in Production </h1> <p>In LLM-based systems, performance is often driven less by model size and more by <strong>what context</strong> is provided, <strong>in what order</strong>, and <strong>under which constraints</str…

  562. dev.to — LLM tag TIER_1 English(EN) · Mark Thorn ·

    将大型语言模型与传统企业系统集成:哪些方法真正有效

    <p>Most LLM integration articles assume you are starting from scratch. Clean microservices. Modern APIs. A greenfield codebase your team controls end to end.</p> <p>That is not where most enterprises live.</p> <p>The real world is SAP instances from 2009, Oracle ERP deployments t…

  563. dev.to — LLM tag TIER_1 English(EN) · Charlie Hadley ·

    CI 中的大语言模型评估:在手动测试耗费您资源之前停止它

    <h1> LLM Evaluation in CI: Stop Manual Testing Before It Costs You </h1> <p>You ship a prompt change to production. Two hours later, a customer complains your LLM is returning hallucinated data. You rollback. You lost an hour of revenue and some user trust.</p> <p>This happens be…

  564. dev.to — LLM tag TIER_1 English(EN) · Charlie Hadley ·

    我在 CI 中构建了 LLM 评估即代码:如何避免发布回归

    <h1> API Rate Limiting Playbook: Protect Your Backend From Abuse </h1> <h2> The Problem </h2> <p>Your API is live in production. Traffic is growing. Then one day, a bot discovers your endpoint and starts hammering it with 100,000 requests per second. Your database melts. Your use…

  565. dev.to — LLM tag TIER_1 English(EN) · Jeremy Longshore ·

    先确定性,后大模型:CI预筛咨询

    <p>The old PR review system ran Gemini on every submission to the <code>claude-code-plugins</code> repo. It broke every time — quota errors, timeout, malformed JSON, the works. On 2026-05-15 I shipped a replacement and deleted the original on the same day.</p> <p>The replacement …

  566. dev.to — LLM tag TIER_1 English(EN) · Ad Man ·

    停止为LLM API支付过高费用:一份实用的成本优化指南 💰

    <h1> Stop Overpaying for LLM APIs: A Practical Cost Optimization Guide </h1> <p>Most teams have a cost problem they don't know about. They send <em>every</em> query to their most expensive model because it's easier than figuring out which queries actually need it.</p> <p>After an…

  567. dev.to — LLM tag TIER_1 English(EN) · Argon Loop ·

    LLM系统中的成本归因

    <p>LLM services are expensive at scale. If you're building multi-tenant systems or running high-volume agents, you need to answer three things: Who used what? How much did it cost? How do I show them the math?</p> <p>This is the cost attribution problem—and it's solved by three p…

  568. dev.to — LLM tag TIER_1 English(EN) · paulo de vries ·

    停止幻觉:一个用于将LLM响应与签名、来源声明进行绑定的开发者API

    <p>TL;DR: I just shipped SourceScore VERITAS — a free-tier-friendly API that returns hand-verified AI/ML claims with their primary sources, an HMAC-SHA256 signature, and a ready-to-paste citation. 51 claims at launch; expanding to 5,000+ this year. curl <a href="https://sourcesco…

  569. dev.to — LLM tag TIER_1 English(EN) · soohan abbasi ·

    思维链及其他:LLM 如何真正学会推理

    <p><em>"The ability to reason step-by-step is not just a feature. It might be the difference between a language model that sounds intelligent and one that actually is."</em></p> <h2> Introduction: When AI Started Thinking </h2> <p>In 2022, researchers at Google Brain published a …

  570. dev.to — LLM tag TIER_1 English(EN) · Akhilesh ·

    84. 微调大型语言模型:教会巨头新技巧

    <p>GPT-3 has 175 billion parameters.</p> <p>Full fine-tuning updates all 175 billion with every gradient step. You need multiple A100 GPUs (each with 80GB memory) just to fit the model. Training for even a few epochs on a moderate dataset costs thousands of dollars. A startup can…

  571. dev.to — LLM tag TIER_1 English(EN) · Norvik Tech ·

    Seclens:评估特定角色的LLM用于安全…

    <blockquote> <p>Originally published at <a href="https://newayzi.com/en/news/evaluacion-especifica-de-roles-llm-seguridad" rel="noopener noreferrer">norvik.tech</a></p> </blockquote> <h2> Introduction </h2> <p>Explore the significance of Seclens in evaluating LLMs for security vu…

  572. dev.to — LLM tag TIER_1 English(EN) · Prakhar Singh ·

    评估 LLM 代码审查员:用于精确度、召回率和路由的离线工具

    <blockquote> <p>If you cannot measure it, you cannot route it. Why offline evaluation is the difference between a code reviewer that improves over time and one the team dismisses within a sprint.</p> </blockquote> <p>Chat evaluations are vibes-based: thumbs-up on "was this helpfu…

  573. dev.to — LLM tag TIER_1 Deutsch(DE) · 丁久 ·

    LLM 微调策略与技术

    <blockquote> <p><em>This article was originally published on <a href="https://dingjiu1989-hue.github.io/en/ai/fine-tuning-strategies.html" rel="noopener noreferrer">AI Study Room</a>. For the full version with working code examples and related articles, visit the original post.</…

  574. dev.to — LLM tag TIER_1 English(EN) · 丁久 ·

    Prompt Chaining:构建多步 LLM 工作流

    <blockquote> <p><em>This article was originally published on <a href="https://dingjiu1989-hue.github.io/en/ai/ai-prompt-chaining.html" rel="noopener noreferrer">AI Study Room</a>. For the full version with working code examples and related articles, visit the original post.</em><…

  575. dev.to — LLM tag TIER_1 English(EN) · Vikrant Shukla ·

    Softmax瓶颈:为什么让LLM更大并不总能让它们更聪明

    <p>When researchers scale a language model — more parameters, more layers, wider hidden dimensions — there's an implicit assumption: a bigger model can represent more things. More expressiveness, more knowledge, better predictions. Mostly this is true. But there's a structural ce…

  576. dev.to — LLM tag TIER_1 English(EN) · Adnan Latif ·

    生产环境中扩展 LLM + Vector DB 系统:实战经验总结

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fj3058op9ajg2qx39n30h.png"><img alt="Cover Image" height="533" …

  577. dev.to — LLM tag TIER_1 English(EN) · 蔡俊鹏 ·

    在本地运行开源大模型:从 Ollama 到 DeepSeek,构建你的私有 AI

    <h2> Foreword </h2> <p>In 2026, open-source LLMs aren't lab experiments anymore. Meta's Llama 4, Alibaba's Qwen 3, DeepSeek-R1 from China — they've caught up with or beaten closed-source models on many benchmarks. And thanks to tools like Ollama and llama.cpp, anyone with a mid-r…

  578. dev.to — LLM tag TIER_1 English(EN) · Vikrant Shukla ·

    迷失在中间:为什么大型语言模型会悄悄忽略其自身上下文窗口的中心

    <p>Every time you hand a long document to an LLM and ask it to summarise or answer a question, something quietly goes wrong. The model reads the whole thing — or appears to — but its answers disproportionately reflect what was at the beginning and the end. Whatever sat in the mid…

  579. dev.to — LLM tag TIER_1 English(EN) · 丁久 ·

    大语言模型评估与基准测试指南 2026:超越简单评估

    <blockquote> <p><em>This article was originally published on <a href="https://dingjiu1989-hue.github.io/en/ai/llm-evaluation-benchmarks.html" rel="noopener noreferrer">AI Study Room</a>. For the full version with working code examples and related articles, visit the original post…

  580. dev.to — LLM tag TIER_1 English(EN) · 丁久 ·

    LLM 函数调用:带代码示例的完整开发者指南

    <blockquote> <p><em>This article was originally published on <a href="https://dingjiu1989-hue.github.io/en/ai/function-calling-guide.html" rel="noopener noreferrer">AI Study Room</a>. For the full version with working code examples and related articles, visit the original post.</…

  581. dev.to — LLM tag TIER_1 English(EN) · 丁久 ·

    微调开源大语言模型:开发者实用指南(2026)

    <blockquote> <p><em>This article was originally published on <a href="https://dingjiu1989-hue.github.io/en/ai/fine-tune-open-source-llm.html" rel="noopener noreferrer">AI Study Room</a>. For the full version with working code examples and related articles, visit the original post…

  582. dev.to — LLM tag TIER_1 English(EN) · Alan West ·

    自信地调试大型语言模型驱动功能的错误答案

    <h2> The bug that took two weeks to surface </h2> <p>A few months back I shipped a feature that used a language model to summarize support tickets and suggest responses. Internal QA loved it. The demo went great. Two weeks after launch, our support lead pinged me on Slack: "Are t…

  583. dev.to — LLM tag TIER_1 English(EN) · Nitin Srivastava ·

    Python 中 LLM 结构化输出的加固:修复重试、成本上限和漂移检测(可运行代码)

    <p>I shipped a structured-output endpoint to production in March. The schema was clean, JSON mode was on, the model was GPT-4.1, the eval suite was green. Three weeks in, the on-call channel lit up because a downstream billing job had silently skipped 4,200 records over a weekend…

  584. dev.to — LLM tag TIER_1 English(EN) · BN ·

    LLM流水线的确定性可靠性堆栈

    <p>I have been spending the last few months wiring up a deterministic reliability stack for structured LLM pipelines.</p> <p>Today, LLM Contract Check (locc) and Release Governor went live on PyPI. EGA went live last week.</p> <p>The stack is straightforward:<br /> LLM Contract C…

  585. dev.to — LLM tag TIER_1 English(EN) · Machine coding Master ·

    停止猜测您的RAG质量:使用Spring AI和LLM-as-a-Judge自动化忠实度指标

    <h2> Stop Shipping Hallucinations: Automating RAG Faithfulness with Spring AI 1.2 </h2> <p>If you’re still "vibe-checking" your RAG outputs in 2026, you’re not an engineer; you’re a gambler. Enterprise-grade AI isn't about getting a cool demo—it's about proving your model isn't h…

  586. dev.to — LLM tag TIER_1 English(EN) · Rob ·

    模型对决:在真实编码任务上对本地与云端 LLM 进行基准测试

    <p>Last post we stood up Ollama on the RTX 5090, pulled a stack of models, and wired them into our coding workflow. The whole time there was an obvious question hanging over it: are local models actually good enough?</p> <p>Not good enough in the abstract benchmarks-on-a-leaderbo…

  587. dev.to — LLM tag TIER_1 English(EN) · Rob ·

    让 GPU 发挥作用:在家庭实验室运行本地 LLM

    <p><a href="https://dev.to/posts/from-idea-to-infrastructure-standing-up-a-self-hosted-ai-dev-environment">Yesterday</a> we went from a gaming PC on a shelf to a fully configured Coder server with GitHub integration, workspace templates, and AI agents. The dev environment is runn…

  588. dev.to — LLM tag TIER_1 English(EN) · Rob ·

    让 GPU 发挥作用:在家庭实验室运行本地 LLM

    <p><a href="https://dev.to/posts/from-idea-to-infrastructure-standing-up-a-self-hosted-ai-dev-environment">Yesterday</a> we went from a gaming PC on a shelf to a fully configured Coder server with GitHub integration, workspace templates, and AI agents. The dev environment is runn…

  589. dev.to — LLM tag TIER_1 English(EN) · Rob ·

    模型对决:在真实编码任务上对本地与云端 LLM 进行基准测试

    <p>Last post we stood up Ollama on the RTX 5090, pulled a stack of models, and wired them into our coding workflow. The whole time there was an obvious question hanging over it: are local models actually good enough?</p> <p>Not good enough in the abstract benchmarks-on-a-leaderbo…

  590. dev.to — LLM tag TIER_1 English(EN) · Nitin Srivastava ·

    使用 Pytest 构建生产级 LLM 评估框架:成本受限、容错、CI 门控(可运行 Python)

    <p>I shipped my fourth LLM agent to production last quarter. By month two, the eval suite that "passed in CI" was the reason a regression made it to a customer.</p> <p>The tests were green. But they were green for the wrong reason — every assertion was a single LLM call against a…

  591. dev.to — LLM tag TIER_1 English(EN) · NaveenKumar Namachivayam ⚡ ·

    超越炒作:AWS Labs LLMeter 大型语言模型基准测试综合指南

    <p id="p-rc_9231198f56807c04-27">In the current AI gold rush, the conversation has shifted from "Can it do the task?" to "How efficiently can it do the task?" For engineers moving Large Language Models (LLMs) into production, the "vibe check" is no longer sufficient. You need har…

  592. dev.to — LLM tag TIER_1 English(EN) · Gabriel Anhaia ·

    LLM响应缓存:当80/20命中率节省开支

    <ul> <li> <strong>Book:</strong> <a href="https://www.amazon.com/dp/B0GYLHMLMT" rel="noopener noreferrer">LLM Observability Pocket Guide: Picking the Right Tracing &amp; Evals Tools for Your Team</a> </li> <li> <strong>Also by me:</strong> <em>Thinking in Go</em> (2-book series) …

  593. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    超越炒作:像OpenAI的GPT-4这样的LLM是如何实际运作的?本文揭开了从你的话到AI的‘理解’的复杂旅程,解释了

    Beyond the hype: How do LLMs like OpenAI's GPT-4 actually function? This article demystifies the complex journey from your words to AI's 'understanding,' explaining tokenization, embeddings, and the crucial transformer architecture. Discover the iterative guessing game and the 'b…

  594. r/Anthropic TIER_1 English(EN) · /u/abhishekkumar333 ·

    LLM 内部解析(语言模型头的洞察)

    <!-- SC_OFF --><div class="md"><p>Due to curiosity of getting to know how an actually large language model like Chatgpt , gemini , claude work internally. I looked into the specific first principle based learning of the process.</p> <p>I have taken example of 4 training sentences…

  595. r/Anthropic TIER_1 English(EN) · /u/silence-and-magic ·

    LLM社交模拟中的蝴蝶效应。与我们如何编写CLAUDE.md和系统提示相关。

    <table> <tr><td> <a href="https://www.reddit.com/r/Anthropic/comments/1tkptj0/the_butterfly_effect_in_llm_social_simulations/"> <img alt="The butterfly effect in LLM social simulations. Relevant to how we write CLAUDE.md and system prompts." src="https://preview.redd.it/59ahvbct4…

  596. r/Anthropic TIER_1 English(EN) · /u/RJSabouhi ·

    资源:LLM证据使用中的源边界故障

    <table> <tr><td> <a href="https://www.reddit.com/r/Anthropic/comments/1tc6d7q/resource_sourceboundary_failures_in_llm_evidence/"> <img alt="Resource: source-boundary failures in LLM evidence use" src="https://external-preview.redd.it/P69EsmfdRn1YdPKlugVsTLq4e-YcCHd7HH4pMEc65E0.pn…

  597. Mastodon — mastodon.social TIER_1 English(EN) · aihaberleri ·

    📰 2026年系统化提示:负约束与结构化JSON提升LLM可靠性 系统化提示正在改变开发人员设计LLM交互方式

    📰 Systematic Prompting in 2026: Negative Constraints & Structured JSON for LLM Reliability Systematic prompting is transforming how developers engineer LLM interactions, with negative constraints, structured JSON outputs, and multi-hypothesis sampling emerging as critical techniq…

  598. Mastodon — mastodon.social TIER_1 Türkçe(TR) · aihaberleri ·

    📰 2026年系统性提示工程:负面约束、JSON输出和多假设方法 专为AI开发者打造的系统性提示工程

    📰 Sistemli Prompt Mühendisliği 2026: Negatif Kısıtlar, JSON Çıktıları ve Çoklu Hipotez Yöntemleri Yapay zeka geliştiricileri için sistemli prompt mühendisliği, sadece soru sormaktan çok, cevabı nasıl şekillendireceğinizi öğrenmektir. Negatif kısıtlar, yapılandırılmış JSON çıktıla…