PulseAugur
实时 11:16:42
English(EN) My Workflow for Understanding LLM Architectures

OpenAI 训练 LLM 以改进指令层级;新研究聚焦优化与验证

OpenAI 推出了 IH-Challenge 数据集,用于训练大型语言模型更好地优先处理来自不同来源(如系统消息、开发者和用户)的指令。此训练旨在通过教会模型遵循一个系统指令最受信任的层级结构,来提高安全可控性和对抗提示注入攻击的鲁棒性。该数据集旨在克服指令层级强化学习中的常见陷阱,确保模型即使在面对冲突的用户或工具生成的提示时,也能可靠地遵守安全策略。 AI

影响 通过提高 LLM 遵循优先指令的能力,增强了 LLM 的安全性和可靠性,降低了提示注入和策略违规的风险。

排序理由 OpenAI 发布了一个新的训练数据集和方法论,以改进 LLM 的安全性和指令遵循能力。

在 Ahead of AI (Sebastian Raschka) 阅读 →

AI 生成摘要 · Google Gemini · 来自 30 个来源。 我们如何撰写摘要 →

OpenAI 训练 LLM 以改进指令层级;新研究聚焦优化与验证

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
OpenAI 发布了一个新的训练数据集和方法论,以改进 LLM 的安全性和指令遵循能力。
Source corroboration
30 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
safety, paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
878 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+9 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [30]

  1. OpenAI News TIER_1 English(EN) ·

    改进前沿大型语言模型中的指令层级

    IH-Challenge trains models to prioritize trusted instructions, improving instruction hierarchy, safety steerability, and resistance to prompt injection attacks.

  2. OpenAI News TIER_1 English(EN) ·

    指令层级:训练LLM优先处理特权指令

    Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with their own malicious prompts.

  3. arXiv cs.LG TIER_1 English(EN) · Katelyn Crumpacker, Dimitrios Nikolopoulos ·

    LLM驱动的运行时参数优化,实现能效模型推理

    arXiv:2604.27032v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become an integral part of many real-world workflows. However, LLMs consume a lot of energy, which becomes a large concern in the scale of the demand for these tools. As LLMs become integrated int…

  4. arXiv cs.CL TIER_1 English(EN) · Ting-Wei Li, Sirui Chen, Jiaru Zou, Yingbing Huang, Tianxin Wei, Jingrui He, Hanghang Tong ·

    EvoSelect:数据高效的LLM进化以实现目标任务适应

    arXiv:2604.26170v1 Announce Type: new Abstract: Adapting large language models (LLMs) to a targeted task efficiently and effectively remains a fundamental challenge. Such adaptation often requires iteratively improving the model toward a targeted task, yet collecting high-quality…

  5. arXiv cs.AI TIER_1 English(EN) · Junbo Jacob Lian, Yujun Sun, Huiling Chen, Chaoyu Zhang, Hanzhang Qin, Chung-Piaw Teo ·

    ReLoop:用于可靠的基于LLM的优化的结构化建模和行为验证

    arXiv:2602.15983v2 Announce Type: replace-cross Abstract: Large language models (LLMs) can translate natural language into optimization code, but silent failures pose a critical risk: code that executes and returns solver-feasible solutions may encode semantically incorrect formu…

  6. arXiv cs.CL TIER_1 English(EN) · Mukai Li, Qingcheng Zeng, Tianqing Fang, Zhenwen Liang, Linfeng Song, Qi Liu, Haitao Mi, Dong Yu ·

    LLM智能体已验证的关键步骤优化

    arXiv:2602.03412v2 Announce Type: replace Abstract: As large language model agents tackle increasingly complex long-horizon tasks, effective post-training becomes critical. Prior work faces fundamental challenges: outcome-only rewards fail to precisely attribute credit to interme…

  7. arXiv cs.LG TIER_1 English(EN) · Jianghao Lin, Zi Ling, Chenyu Zhou, Tianyi Xu, Ruoqing Jiang, Zizhuo Wang, Dongdong Ge ·

    从自言自语到广场:具有去中心化辩论的增强记忆大语言模型代理用于优化建模

    arXiv:2604.25847v1 Announce Type: cross Abstract: Optimization modeling underpins real-world decision-making in logistics, manufacturing, energy, and public services, but reliably solving such problems from natural-language requirements remains challenging for current large langu…

  8. arXiv cs.CL TIER_1 English(EN) · Alex Bogdan, Adrian de Valois-Franklin ·

    大型语言模型输出的惊人普适性:一种实时验证原语

    arXiv:2604.25634v1 Announce Type: cross Abstract: We report a striking statistical regularity in frontier LLM outputs that enables a CPU-only scoring primitive running at 2.6 microseconds per token, with estimated latency up to 100,000$\times$ (five orders of magnitude) below exi…

  9. arXiv cs.CL TIER_1 English(EN) · Hanghang Tong ·

    EvoSelect:数据高效的LLM进化以实现目标任务适应

    Adapting large language models (LLMs) to a targeted task efficiently and effectively remains a fundamental challenge. Such adaptation often requires iteratively improving the model toward a targeted task, yet collecting high-quality human-labeled data to support this process is c…

  10. arXiv cs.AI TIER_1 English(EN) · Dongdong Ge ·

    从独白到广场:具有去中心化辩论的增强记忆大型语言模型代理用于优化建模

    Optimization modeling underpins real-world decision-making in logistics, manufacturing, energy, and public services, but reliably solving such problems from natural-language requirements remains challenging for current large language models (LLMs). In this paper, we propose \emph…

  11. arXiv cs.CL TIER_1 English(EN) · Adrian de Valois-Franklin ·

    LLM输出令人惊讶的普遍性:一种实时验证原语

    We report a striking statistical regularity in frontier LLM outputs that enables a CPU-only scoring primitive running at 2.6 microseconds per token, with estimated latency up to 100,000$\times$ (five orders of magnitude) below existing sampling-based detectors. Across six contemp…

  12. arXiv cs.CL TIER_1 English(EN) · Tao Feng, Haozhen Zhang, Zijie Lei, Peixuan Han, Jiaxuan You ·

    GraphPlanner:用于多智能体 LLM 的图记忆增强型智能体路由

    arXiv:2604.23626v1 Announce Type: new Abstract: LLM routing has achieved promising results in integrating the strengths of diverse models while balancing efficiency and performance. However, to support more realistic and challenging applications, routing must extend into agentic …

  13. arXiv cs.CL TIER_1 English(EN) · Yating Wu, Yuhao Zhang, Sayan Ghosh, Sourya Basu, Anoop Deoras, Jun Huan, Gaurav Gupta ·

    ContextWeaver:为LLM智能体构建选择性且依赖结构化的记忆

    arXiv:2604.23069v1 Announce Type: new Abstract: Large language model (LLM) agents often struggle in long-context interactions. As the agent accumulates more interaction history, context management approaches such as sliding window and prompt compression may omit earlier structure…

  14. arXiv cs.CL TIER_1 English(EN) · Parsa Ashrafi Fashi, Utkarsh Saxena, Mehdi Rezagholizadeh, Aref Jafari, Akash Haridas, Mingyu Yang, Vansh Bhatia, Guihong Li, Vikram Appia, Emad Barsoum ·

    长上下文感知升级再造:混合大模型扩展的新前沿

    arXiv:2604.24715v1 Announce Type: new Abstract: Hybrid sequence models that combine efficient Transformer components with linear sequence modeling blocks are a promising alternative to pure Transformers, but most are still pretrained from scratch and therefore fail to reuse exist…

  15. arXiv cs.CL TIER_1 English(EN) · Max Schaffelder, Albert Gatt ·

    多个篮子里的合成鸡蛋:合成数据多样性对LLM微调的影响

    arXiv:2511.01490v2 Announce Type: replace Abstract: As synthetic data becomes widely used in language model development, understanding its impact on model behavior is crucial. This paper investigates the impact of the diversity of sources of synthetic data on fine-tuned large lan…

  16. arXiv cs.CL TIER_1 English(EN) · Hao Ban, Kaiyi Ji ·

    多 LoRA 联合微调 LLM 的参数共享新思路

    arXiv:2509.25414v2 Announce Type: replace-cross Abstract: Large language models are often adapted using parameter-efficient techniques such as Low-Rank Adaptation (LoRA), formulated as $y = W_0x + BAx$, where $W_0$ is the pre-trained parameters and $x$ is the input to the adapted…

  17. arXiv cs.LG TIER_1 English(EN) · Irene Tenison, Stella Ahn, Miriam Kim, Ebtisam Alshehri, Lalana Kagal ·

    参数效率不等于内存效率:重新思考用于设备端 LLM 适配的微调

    arXiv:2604.22783v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) has become the standard for adapting large language models (LLMs). In this work we challenge the wide-spread assumption that parameter efficiency equates memory efficiency and on-device adaptab…

  18. arXiv cs.AI TIER_1 English(EN) · Shiju Wang, Yujie Wang, Ao Sun, Fangcheng Fu, Zijian Zhu, Bin Cui, Xu Han, Kaisheng Ma ·

    InfiniPipe:用于高效可变长度长上下文大模型训练的弹性流水线并行

    arXiv:2509.21275v4 Announce Type: replace-cross Abstract: Long context training is crucial for LLM's context extension. Existing schemes, such as sequence parallelism, incur substantial communication overhead. Pipeline parallelism (PP) reduces this cost, but its effectiveness hin…

  19. arXiv cs.CL TIER_1 English(EN) · Emad Barsoum ·

    长上下文感知升级再造:混合大模型扩展的新前沿

    Hybrid sequence models that combine efficient Transformer components with linear sequence modeling blocks are a promising alternative to pure Transformers, but most are still pretrained from scratch and therefore fail to reuse existing Transformer checkpoints. We study upcycling …

  20. arXiv cs.LG TIER_1 English(EN) · Rajinder Sandhu, Di Mu, Cheng Chang, Md Shahriar Tasjid, Himanshu Rai, Maksims Volkovs, Ga Wu ·

    通过蒸馏使密集检索器与LLM效用对齐

    arXiv:2604.22722v1 Announce Type: cross Abstract: Dense vector retrieval is the practical backbone of Retrieval- Augmented Generation (RAG), but similarity search can suffer from precision limitations. Conversely, utility-based approaches leveraging LLM re-ranking often achieve s…

  21. arXiv cs.LG TIER_1 English(EN) · Xiucheng Xu, Bingbing Xu, Xueyun Tian, Zihe Huang, Rongxin Chen, Yunfan Li, Huawei Shen ·

    Chain-of-Memory:LLM智能体轻量级记忆构建与动态演化

    arXiv:2601.14287v2 Announce Type: replace Abstract: External memory systems are pivotal for enabling Large Language Model (LLM) agents to maintain persistent knowledge and perform long-horizon decision-making. Existing paradigms typically follow a two-stage process: computational…

  22. arXiv cs.AI TIER_1 English(EN) · Ga Wu ·

    通过蒸馏使密集检索器与LLM效用对齐

    Dense vector retrieval is the practical backbone of Retrieval- Augmented Generation (RAG), but similarity search can suffer from precision limitations. Conversely, utility-based approaches leveraging LLM re-ranking often achieve superior performance but are computationally prohib…

  23. arXiv cs.CL TIER_1 English(EN) · Shumin Deng ·

    StructMem:LLM长时行为的结构化记忆

    Long-term conversational agents need memory systems that capture relationships between events, not merely isolated facts, to support temporal reasoning and multi-hop question answering. Current approaches face a fundamental trade-off: flat memory is efficient but fails to model r…

  24. Hugging Face Daily Papers TIER_1 English(EN) ·

    StructMem:LLM长时行为的结构化记忆

    Long-term conversational agents need memory systems that capture relationships between events, not merely isolated facts, to support temporal reasoning and multi-hop question answering. Current approaches face a fundamental trade-off: flat memory is efficient but fails to model r…

  25. arXiv cs.CL TIER_1 English(EN) · Emily Chen ·

    面向自动化大语言模型红队测试的自适应指令组合

    Many approaches to LLM red-teaming leverage an attacker LLM to discover jailbreaks against a target. Several of them task the attacker with identifying effective strategies through trial and error, resulting in a semantically limited range of successes. Another approach discovers…

  26. Hugging Face Daily Papers TIER_1 English(EN) ·

    面向低成本LLM服务的连续语义缓存

    As Large Language Models (LLMs) become increasingly popular, caching responses so that they can be reused by users with semantically similar queries has become a vital strategy for reducing inference costs and latency. Existing caching frameworks have proposed to decide which que…

  27. Ahead of AI (Sebastian Raschka) TIER_1 English(EN) · Sebastian Raschka, PhD ·

    我理解LLM架构的工作流程

    A learning-oriented workflow for understanding new open-weight model releases

  28. Ahead of AI (Sebastian Raschka) TIER_1 English(EN) · Sebastian Raschka, PhD ·

    大型语言模型架构大比拼

    From DeepSeek-V3 to Kimi K2: A Look At Modern LLM Architecture Design

  29. arXiv stat.ML TIER_1 English(EN) · Jonathan Geuter, Youssef Mroueh, David Alvarez-Melis ·

    Guided Speculative Inference for Efficient Test-Time Alignment of LLMs

    arXiv:2506.04118v3 Announce Type: replace-cross Abstract: We propose Guided Speculative Inference (GSI), a novel algorithm for efficient reward-guided decoding in large language models. GSI combines soft best-of-$n$ test-time scaling with a reward model $r(x,y)$ and speculative s…

  30. Smol AINews TIER_1 English(EN) ·

    OpenAI 为大语言模型操作系统设计的指令层级

    **OpenAI** published a paper introducing the concept of privilege levels for LLMs to address prompt injection vulnerabilities, improving defenses by 20-30%. **Microsoft** released the lightweight **Phi-3-mini** model with 4K and 128K context lengths. **Apple** open-sourced the **…