Grok 4.6
PulseAugur coverage of Grok 4.6 — every cluster mentioning Grok 4.6 across labs, papers, and developer communities, ranked by signal.
- 2026-09-01 research_milestone LatchBio's independent analysis shows Grok 4.6 excels at refusing disguised biosecurity hazards while performing routine biological tasks. 来源
- 2026-08-26 product_launch xAI's Grok 4.6 model has been made available on Microsoft Foundry. 来源
- 2026-08-21 research_milestone Grok 4.6 achieved the #1 score on the CursorBench benchmark. 来源
- 2026-08-21 product_launch xAI's Grok 4.6 model has been made available on Google Cloud's Vertex AI platform. 来源
- 2026-08-20 product_launch xAI's Grok 4.6 model was launched on AWS Bedrock, offering a 500K token context window. 来源
- 2026-08-20 product_launch xAI released Grok 4.6 on AWS Bedrock, featuring enhanced context handling and reasoning. 来源
- 2026-08-19 product_launch xAI's Grok 4.6 model has been launched and is now available on Amazon Bedrock. 来源
- 2026-08-14 product_launch SpaceXAI has reportedly released an updated version of its Grok model, Grok 4.6. 来源
- 2026-08-14 product_launch xAI has released Grok 4.6, a new model focused on long-running agents and complex interactive/visual tasks. 来源
- 2026-08-14 product_launch xAI released Grok 4.6, a new model focused on long-running agents and complex interactive tasks. 来源
- 2026-08-14 product_launch xAI has begun distributing usage reset tokens for its new Grok 4.6 model. 来源
- 2026-08-14 product_launch Elon Musk announced that Grok 4.6 is optimized to work with the Grok Build harness. 来源
- 2026-08-14 product_launch xAI released Grok 4.6, integrating it into GitHub Copilot and making it available via the SpaceXAI console. 来源
- 2026-08-13 product_launch xAI released Grok 4.6, featuring a significant improvement in its non-hallucination rate. 来源
- 2026-08-13 product_launch xAI released Grok 4.6, featuring a significant improvement in its non-hallucination rate. 来源
27 天有情绪数据
Grok 4.6 pricing and performance details emerge
Grok 4.6 has been released, matching top-tier OpenAI models on benchmarks and undercutting them on price. It demonstrates superior efficiency in agentic tasks, completing complex workflows in fewer steps and at a lower cost than Claude Opus 5. This suggests xAI is focusing on value and efficiency for its users.
Grok 4.6 shows specific weaknesses in real terminal work despite strong benchmark performance
While Grok 4.6 performs well on benchmarks and agentic tasks, it reportedly shows weaknesses in areas like real terminal work. This suggests that despite its overall strong performance, it may not be suitable for all types of command-line or system interaction tasks without further refinement.
xAI to release a faster, more expensive version of Grok 4.6 within 30 days
The release notes for Grok 4.6 mention a faster version available at double the cost. This indicates a tiered product strategy, with a premium offering likely to be rolled out soon to capture users who require maximum performance.
Grok 4.6 shows specific weaknesses in real terminal work despite strong benchmark performance
While Grok 4.6 matches top-tier models on benchmarks and excels in agentic tasks, reports indicate it shows weaknesses in 'real terminal work'. This suggests a potential gap in its practical application for certain command-line interface or system administration tasks, despite its overall advanced capabilities.
xAI to release a faster, more expensive version of Grok 4.6 within 30 days
The mention of a 'faster version available at double the cost' for Grok 4.6 indicates xAI's intention to offer tiered performance options. Given the recent release, it's plausible they will formally launch this enhanced version soon to cater to users requiring maximum speed and throughput for demanding applications.
-
SpaceXAI 发布 Grok Build 编码代理,面向开发者
SpaceXAI 推出了 Grok Build,这是一款新的基于终端的编码代理,旨在直接集成到软件开发环境中。该代理由 Grok 4.6 大语言模型驱动,能够读取和写入文件、执行命令以及评估代码存储库以进行更改、修复错误或添加新功能。可在 Windows、macOS 和 Linux 上安装,访问权限根据 Grok 订阅计划分级,免费用户可能会遇到令牌限制。
-
Grok 4.6 基准测试得分因报告方法不同而差异巨大
xAI 的 Grok 4.6 在同一基准测试中,根据测量方式的不同,表现得分差异巨大。根据 xAI 自家的 Terminal-Bench 3.0 模型卡,该模型得分为 26%,但 Artificial Analysis 的一项独立分析报告称,在基准测试的一个略有不同的版本上得分为 88.4%。这种差异凸显了在报告基准测试结果时限定条件的重要性。
-
LLM API价格飙升,DeepSeek费率翻三倍,OpenAI削减成本
此前五个月一直稳定的Frontier大型语言模型(LLM)API价格在8月份出现显著变动。DeepSeek通过引入基于一天中不同时段的阶梯定价方案,将其V4 Pro模型的峰值费率提高了两倍。OpenAI将其GPT-5.6 Sol模型的价格降低了29%,但此促销定价仅保证到2026年11月。尽管发生了这些变化,但由于DeepSeek的价格上涨被OpenAI的降价以及其他模型保持先前费率所抵消,旗舰LLM价格的整体指数在本月下降了5.7%。
-
Cursor 用户讨论为不同任务配置特定 AI 模型
Cursor subreddit 的用户正在讨论在 Cursor IDE 中为不同任务或模式配置特定 AI 模型的可能性。对话探讨了为规划、代理操作和协调器角色设置不同的模型,这表明希望在开发环境中对 AI 模型的部署进行精细控制。
-
前沿 AI 模型面临定价考验,token 使用量激增 25 倍 · 已追踪 2 个来源
在过去一年中,token 使用量激增了 25 倍,仅上个月就翻了一番,前沿 AI 模型正面临巨大的定价压力。这种激增是由 AI 服务成本的降低推动的,导致了一种类似于杰文斯悖论的现象,即可负担性的提高导致消费量大幅增加。虽然 OpenAI 和 Anthropic 的顶级模型仍然非常强大,但成本高昂,促使人们更加关注中端模型,这些模型以六分之一的价格提供旗舰级智能约 90% 的能力。
-
面向LLM的成本感知影子测试:实用指南
一位开发者概述了一种方法,通过在生产流水线中进行“影子测试”来评估新的大型语言模型。这种方法将候选模型与现有模型进行比较,使用真实世界的提示和失败案例,而不是仅仅依赖公开基准。目标是在完全集成新模型之前,评估其在特定工作负载上的性能,包括延迟、令牌使用量和输出正确性。作者建议使用免费的OpenAI兼容端点,例如MonkeyCode提供的端点,以促进这些成本感知的测试。
-
Gemini 3.8 Flash 在基准测试中表现媲美高端 LLM,成本却仅为其一小部分 · 跟踪 2 个来源
对三款新的大型语言模型——谷歌的 Gemini 3.8 Flash、Anthropic 的 Claude Fable 5.1 和 OpenAI 的 GPT-5.6 Sol——的比较显示,在独立基准测试中,它们的性能相当,但价格差异显著。Gemini 3.8 Flash 的价格仅为另外两款模型的一小部分,但在人工智能分析智能指数(Artificial Analysis Intelligence Index)上得分相同。然而,Gemini…
-
Google DeepMind发布Gemini 3.8 Flash;Anthropic维持Claude Sonnet 3.5定价
Google DeepMind发布了两款新模型Gemini 3.8 Flash和Gemini 3.8 Flash Cyber,前者在保持低成本的同时增强了智能体编码和多步推理能力。与此同时,Anthropic决定将Claude Sonnet 3.5的定价固定为每百万token输入$2,输出$10,尽管此前有涨价意向。这种定价的稳定性,加上Grok 4.6的持续价格,凸显了在API层面使高性能模型更实惠的竞争格局。
-
Meta的Muse Spark 1.3以具有竞争力的价格挑战顶级AI模型
Meta发布了Muse Spark 1.3,这是一款专为长周期代理和编码任务设计的AI模型。与前代Muse Spark 1.2相比,该模型在效率上有所提高,使用的工具调用和token更少。Muse Spark 1.3可通过Muse Code和Meta Model API获得,其定价策略尤为引人注目,为选择允许其数据用于未来模型训练的用户提供大幅折扣。
-
谷歌发布 Gemini 3.8 Flash,提升 AI 性能和成本效益
谷歌发布了其最新的 AI 模型 Gemini 3.8 Flash,旨在重塑其在竞争激烈的前沿模型领域的地位。这一新版本在智能得分方面取得了显著进步,可与 GPT-5.6 Sol 和 Grok 4.6 等顶级模型相媲美,同时还提供了更快的速度和更高的成本效益。该模型旨在更勤奋地处理复杂任务,执行额外的推理步骤并迭代调用工具,这可能会导致更高的 token 使用量,但最终能为企业知识工作带来更好的性能。
-
Meta AI 发布 Muse Spark 1.3,增强了代理和编码能力 · 跟踪 6 个来源
Meta AI 发布了 Muse Spark 1.3,这是其大型语言模型的更新版本,专注于增强代理工作流和编码任务。新版本在性能、通过澄清性问题和确认来改善用户协作以及减少幻觉方面都有所提高。基准测试表明,与前代 Muse Spark 1.2 相比,Muse Spark 1.3 使用的工具调用减少了约 20%,令牌减少了 25%,从而提高了效率和成本效益。
-
Multiverse Computing推出Quasar 438B,欧洲领先的AI模型
Multiverse Computing推出了Quasar 438B,这是一款专为企业级代理和编码任务设计的新型大型语言模型。该模型在人工智能分析指数(Artificial Analysis Intelligence Index)上表现强劲,得分43,使其成为该基准测试中欧洲领先的AI模型。Quasar 438B还展现出具有竞争力的速度,在15.3秒内生成500个token,并在长上下文推理和代理编码评估方面表现出色。
-
用户报告 Grok 4.6 延迟问题,比 Claude Opus 慢
用户报告称,Grok 4.6 自最初发布以来速度明显变慢。一位用户指出,其响应速度现在比 Anthropic 的 Claude Opus 慢。这种减速可能是由于容量限制或开发人员的调整。
-
用户报告:Sol 在 AI 辅助设计任务中表现优于 Grok 4.6
Reddit r/cursor 版块的一位用户分享了他们使用 draw.io 进行科学图表创建任务时,比较 Sol 和 Grok 4.6 这两个 AI 模型的使用体验。用户发现,当 Sol 设置到最高思考力时,它能高效地一次性完成任务。相比之下,Grok 4.6 即使在最高思考力设置下,也需要十几次迭代,仍未能产生令人满意的结果。该用户正在寻求关于如何编写模型无关的 AI 技能,或如何管理特定于 Sol 等模型的技能的见解。
-
用户质疑 Grok 4.6 与 4.5 的 API 使用量和成本效益
一位 Reddit r/cursor 版块的用户正在询问 Grok 4.6 和 Grok 4.5 在 API 使用量和速度方面的差异。他们担心目前在 20 美元套餐上的使用情况,并希望就如何优化这些模型的“性价比”征求建议,同时考虑“Effort: High”和“Speed: Fast”等设置。用户还提到考虑切换回 Composer。
-
xAI的Grok 4.6在生物安全评估中领先,表现优于同类模型
xAI的Grok 4.6在LatchBio进行的生物安全评估中表现出卓越的性能。该模型在LatchBio的BioSecBench-Refusal套件上表现出色,显示出区分合法生物研究和危险请求的强大能力,同时在常规生物任务上保持高性能。尽管Grok 4.6在生物监测任务上的表现与其他前沿模型相当,但其对伪装的危险查询的拒绝率却显著更高。
-
SpaceXAI 的 Grok 4.6 模型在 Microsoft Foundry 上发布,用于代理任务
SpaceXAI 的 Grok 4.6 模型现已在 Microsoft Foundry 上公开预览,并作为 Azure Direct Model 集成。该模型专为长周期、代理式任务而设计,专注于可靠的多步规划、工具利用和错误恢复,以生成完整的工作产品。它以具有竞争力的价格提供前沿推理能力,并具有可选的推理深度、多模态输入和 200K 令牌上下文窗口等功能,特别适合 .NET 开发人员处理复杂的代理式应用程序。
-
AI模型在文案质量方面引发讨论:Claude优于Grok
在Cursor子论坛上,用户正在讨论各种AI模型在文案任务中的有效性。虽然Grok 4.6等模型在编码能力方面受到关注,但它们在创意写作中被批评生成过于简化和机械化的语言。参与者建议,Anthropic的Claude模型,尤其是在提供具体指令和规则时,能为生成类人文案提供更有希望的结果。
-
OpenAI 发布 Jalapeño 芯片结果,在泄露事件后加强安全;SpaceXAI 发布 Grok 4.6
OpenAI 分享了其 Jalapeño 推理芯片的初步结果,声称其每瓦性能和延迟优于现有系统,并计划在年底前内部部署。该公司还宣布了在 Hugging Face 发生由 AI 驱动的泄露事件后加强安全措施,该事件导致一次重要的训练运行暂时中断并审查内部安全协议。另外,SpaceXAI 发布了 Grok 4.6,一个具有 500K 上下文窗口的模型,专为长期运行的代理和编码任务设计,可能受益于 Cursor 收购带来的训练数据和分发。
-
Cursor IDE 用户可以通过将努力程度设置为超高来提升 Grok 4.6 的性能
一位 Reddit 用户分享了在 Cursor IDE 中优化 Grok 4.6 模型的一个技巧。通过将模型的努力程度从默认的“高”调整为“超高”,用户发现这显著减少了错误,尽管并未改变其整体使用模式。这一设置调整被作为一项有用的技巧提供给其他使用 Grok 4.6 的 Cursor 用户。