研究人员正在探索评估和理解大型语言模型(LLM)的新方法。一项研究引入了SAST-IR框架,以测试LLM在面对说服性攻击时的事实鲁棒性,并揭示了简单策略的高成功率。另一篇论文研究了反事实自我解释,发现模型规模对这些解释的质量和忠实度有显著影响。此外,一项研究提出了一个结合线性聊天和空间画布的新颖界面,以改善复杂LLM对话历史的导航和探索,尽管存在采用挑战。最后,研究检查了LLM如何检索和使用内部知识,以及外部工具的可用性如何意外地阻碍它们回答自身知识库中问题的能力。
AI
arXiv:2609.19155v1 Announce Type: new Abstract: In Human-LLM dialogue, follow-up user utterances may implicitly conflict with earlier intents, leading the LLM to misinterpret user needs and generate inappropriate responses. A reliable dialogue system should proactively detect use…
arXiv cs.AI
TIER_1English(EN)·Mahsa Amani, Seungeon Lee, Abhisek Dash, Asmaa El Fraihi, Yunah Jang, Elisabeth Kirsten, Qinyuan Wu, Krishna P. Gummadi, Manish Gupta, Abhilasha Ravichander, Muhammad Bilal Zafar, Soumi Das·
arXiv:2609.19244v1 Announce Type: new Abstract: Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claud…
arXiv:2609.18649v1 Announce Type: new Abstract: Warning: This paper contains examples of stereotypes and social bias. LLMs are increasingly used in interactive settings by the general public, making the evaluation of model behavior in multi-turn conversational scenarios important…
arXiv cs.CL
TIER_1English(EN)·Claudiu Creanga, Liviu P. Dinu·
arXiv:2609.19070v1 Announce Type: new Abstract: This paper introduces ``question archaeology'', a specific evaluation task focused on inferring the single, authentic "genesis question" that motivated the creation of a complete text. Distinct from question generation, which target…
arXiv:2609.17119v1 Announce Type: new Abstract: Large language models can easily generate explanations for their own outputs, but such self-explanations are not necessarily faithful to the model's behavior. We study this issue through counterfactual self-explanations, where a mod…
arXiv cs.CL
TIER_1English(EN)·Rifat Mehreen Amin, Alperen Adatepe, Daniela Fernandes, Daniel Buschek, Andreas Butz·
arXiv:2605.15848v2 Announce Type: replace-cross Abstract: As LLM conversations grow, their histories capture alternative directions, decisions, and evolving lines of thought that can be difficult to navigate through chat alone. We investigate an interaction concept that represent…
arXiv:2609.16777v1 Announce Type: new Abstract: As Large Language Models (LLMs) increasingly serve as primary knowledge retrieval interfaces, their robustness against \textit{persuasion attacks}---attempts to inject misinformation or enforce counterfactuals---has become a critica…
arXiv:2609.14157v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed with external tools that extend what they can do beyond their own knowledge. Tools help on tasks that need external information, but their availability may also change how a mod…
arXiv cs.AI
TIER_1English(EN)·Myra Cheng, Lujain Ibrahim, Grace Liu, Michelle S. Lam, Vishakh Padmakumar, Nick Madibekov, Diyi Yang, Dan Jurafsky·
arXiv:2609.14849v1 Announce Type: cross Abstract: We characterize how people are turning to LLMs as oracles: all-knowing authorities on subjective personal questions. Motivated by risks to users' autonomy and well-being, we develop a typology and LLM-based methods to measure this…
arXiv:2609.11859v1 Announce Type: new Abstract: How does a language model's dependence on query-routing information and target knowledge change as it answers a question? We study this question through layerwise interventions on the hidden state at the end of the question. Across …
Medium — fine-tuning tag
TIER_1English(EN)·Zeynep Kara·
Определитель повадок: пять способов, которыми разговор с LLM имитирует доказательство Парный опыт над двумя фронтирными моделями, разбор двух разных механизмов отказа и проверка собственного вывода по литературе, которая его обрушила. https:// habr.com/ru/articles/1083472/ # иску…
<!-- SC_OFF --><div class="md"><p>The OpenAI-HuggingFace incident has raised critical security concerns. Much of the surrounding discussion has focused on increasingly autonomous or “rogue” AI behavior and AI capability outpacing human governance. This white paper proposes a diff…
<!-- SC_OFF --><div class="md"><p>I got tired of LLMs writing essays when I just need a quick answer. Instead of fighting it or typing "keep it short" every time, I made a tiny tag system (<code>.</code> and <code>!</code>) for my prompts. It's super fast to type on bot…
<!-- SC_OFF --><div class="md"><p>The OpenAI-HuggingFace incident has raised critical security concerns. Much of the surrounding discussion has focused on increasingly autonomous or “rogue” AI behavior and AI capability outpacing human governance. This white paper proposes a diff…