PulseAugur
中
实时 18:01:40
English(EN) 📰 Open-Source AI Agent Scores 65.2% on TerminalBench 2.0 in 2026, Beating Gemini and Junie CLI An open-source AI agent has achieved a record 65.2% success rate

开源AI代理在TerminalBench 2.0上超越Gemini和GPT-4

一个在土耳其开发的名为OSS Agent I的开源AI代理在TerminalBench 2.0基准测试中取得了65.2%的成功率。这一表现超越了Google的Gemini-3-flash-preview、GPT-4和Anthropic的Claude 3等成熟模型。开发者已确认未采用任何欺骗性手段,凸显了该代理在处理复杂终端任务方面的真实能力。 AI

影响 展示了开源AI代理在自主完成复杂现实世界任务方面的显著进步。

排序理由 开源模型发布取得了显著的基准测试结果。

在 Mastodon — mastodon.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

开源AI代理在TerminalBench 2.0上超越Gemini和GPT-4

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
开源模型发布取得了显著的基准测试结果。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
164 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. Mastodon — mastodon.social TIER_1 English(EN) · aihaberleri ·

    📰 开源AI代理在2026年TerminalBench 2.0上得分65.2%,超越Gemini和Junie CLI 一个开源AI代理取得了创纪录的65.2%成功率

    📰 Open-Source AI Agent Scores 65.2% on TerminalBench 2.0 in 2026, Beating Gemini and Junie CLI An open-source AI agent has achieved a record 65.2% success rate on TerminalBench 2.0, surpassing Google's Gemini-3-flash-preview and Junie CLI. The developer confirms no cheating mecha…

  2. Mastodon — mastodon.social TIER_1 Türkçe(TR) · aihaberleri ·

    📰 OSS Agent I 在 2026 年 TerminalBench 排行第一:土耳其 AI 超越 GPT-4 和 Claude 3。OSS Agent I 由土耳其开发,...

    📰 OSS Agent I 2026'da TerminalBench'de Birinci Oldu: Türkiye Yapay Zekâsı GPT-4 ve Claude 3'ü Geçti Türkiye'de geliştirilen OSS Agent I, TerminalBench adlı dünyanın en zorlu terminal ortamı testinde ilk sıraya yükseldi. Bu başarı, yapay zekânın gerçek dünya görevlerini bağımsız t…