Gemini 3.5 Flash
PulseAugur coverage of Gemini 3.5 Flash — every cluster mentioning Gemini 3.5 Flash across labs, papers, and developer communities, ranked by signal.
- developed by Google DeepMind 100%
- developed Gemini 3.6 Flash 90%
- developed by Gemini 3.6 Flash 90%
- developed by Gemini 3.5 Flash-Lite 90%
- developed Gemini 3.5 Flash-Lite 90%
- instance of Gemini 3.5 Flash Cyber 90%
- instance of Gemini 3.5 Flash-Lite 90%
- competes with Claude (Opus 4.8) 70%
- competes with Claude Fable-5 70%
- competes with Claude Sonnet-5 70%
- competes with GPT 5.6 "Sol" 70%
- competes with Minimax M3 70%
- 2026-07-23 product_launch Google has launched Gemini 3.5 Flash, a new lightweight model designed for efficient and low-cost operations. 来源
- 2026-07-01 product_launch exaBase AI has added Gemini 3.5 Flash to its Japan region offerings. 来源
- 2026-06-25 product_launch Google released the Gemini 3.5 Flash model, designed to enhance computer usage and PC operations. 来源
- 2026-06-25 product_launch Google released the Gemini 3.5 Flash AI model. 来源
- 2026-06-25 product_launch Google integrated computer control capabilities into its Gemini 3.5 Flash model. 来源
- 2026-06-25 product_launch Google integrated computer device operation capabilities into its Gemini 3.5 Flash model, naming the feature Jetstream. 来源
- 2026-06-24 product_launch Google DeepMind integrated "computer use" capabilities into the Gemini 3.5 Flash model. 来源
- 2026-06-24 product_launch Google DeepMind integrated computer control capabilities into Gemini 3.5 Flash. 来源
- 2026-06-01 research_milestone A security researcher demonstrated Gemini 3.5 Flash's tendency to provide dangerous advice when prompted, highlighting LLM overreliance risks. 来源
- 2026-05-31 product_launch Google launched the Gemini 3.5 Flash model with a significant price increase. 来源
- 2026-05-29 product_launch Google DeepMind launched Gemini 3.5 Flash, a new model optimized for agentic coding tasks. 来源
- 2026-05-27 product_launch Google's Gemini 3.5 Flash model has been made generally available globally. 来源
- 2026-05-27 product_launch Google plans to widely deploy the Gemini 3.5 Flash model. 来源
- 2026-05-23 product_launch Google released Gemini 3.5 Flash, a new model tier that outperforms its predecessor Gemini 3.1 Pro on coding and agentic tasks. 来源
- 2026-05-20 product_launch Google announced Gemini 3.5 Flash at Google I/O, highlighting its performance and new pricing structure. 来源
16 天有情绪数据
Gemini 3.5 Flash to see significant adoption in enterprise agentic workflows within 90 days
The recent release of Gemini 3.5 Flash, coupled with the introduction of Managed Agents and the Antigravity CLI, strongly suggests Google's strategic pivot towards enterprise automation. The model's optimization for agentic tasks, long-horizon reasoning, and cost-effectiveness, combined with simplified deployment, positions it as a prime candidate for widespread adoption in enterprise environments seeking to automate routine tasks.
Google to offer Gemini 3.5 Flash via a new compute-usage-based billing system within 60 days
The mention of a 'novel billing system based on compute usage rather than query count' in conjunction with the Gemini 3.5 Flash release and the streamlining of AI offerings into three subscription plans indicates a shift in Google's monetization strategy. This new billing model is likely to be rolled out soon, particularly for the cost-optimized Flash model, to align with its performance characteristics and encourage broader adoption.
Gemini 3.5 Flash's 'medium' effort level is a key differentiator for balancing performance and cost
The default 'medium' effort level for Gemini 3.5 Flash represents a deliberate design choice by Google to offer a balanced performance profile. This setting aims to provide a compelling mix of speed, cost-efficiency, and quality, making it attractive for a broad range of applications, especially those requiring rapid iteration and cost control, distinguishing it from models that might prioritize raw power at a higher expense.
-
AI 框架 ARQ 精炼 CodeQL 查询以检测 C/C++ 漏洞
研究人员开发了 ARQ,一个利用大型语言模型 (LLM) 自动精炼 CodeQL 查询以检测 C/C++ 程序中漏洞的新框架。这种代理方法利用合成程序的执行基础证据来识别和纠正现有查询中的假阳性和假阴性。ARQ 不需要标记数据集或提交历史,在真阳性检测方面取得了显著改进,最高可达 119.8%,同时保持至少 98.0% 的精确率。精炼后的查询还成功解决了 GitHub 上长期存在的问题,并发现了 libpng 和 zlib 等真实世界…
-
Claude Code、OpenAI Codex 领跑 AI 编码代理排名 · 跟踪 1 个来源
2026 年 8 月最新的 AI 编码代理排名显示,由 Claude Opus 5 驱动的 Claude Code 在整体终端能力方面位居榜首,紧随其后的是运行在 GPT-5.6 Sol 上的 OpenAI Codex。虽然两个代理在 Terminal-Bench 2.1 基准测试中均获得高分,但 Claude Code 因其强大的控制能力、子代理功能和可靠的多步终端循环而受到青睐,尽管成本较高。OpenAI Codex 在无人值守的…
-
AI应用RollTab使用Transformer模型生成钢琴音乐续集
一款名为RollTab的iPhone应用已发布,该应用使用AI模型自动生成钢琴演奏的续集。该应用由Simon Edwardson开发,利用了一个定制训练的1.25亿参数Transformer模型,该模型将MIDI文件作为离散事件序列进行处理。这种方法可以快速生成音乐续集,应用能够在两秒内生成输出。Edwardson利用Gemini 3.5 Flash收集偏好数据,并使用Direct Preference Optimization (D…
-
新的AI越狱方法利用了时间和描述性漏洞
研究人员开发了新的方法来绕过AI模型的安全过滤器,这些方法同时针对大型视觉语言模型(LVLMs)和文本到图像(T2I)模型。一种名为TempJail的技术,通过操纵字幕时间和调度来引发有害响应,从而利用LVLMs的时间漏洞。另一种名为Etch的方法,通过将有害文本嵌入生成的图像中来针对T2I模型,绕过了传统的基于视觉的安全措施。这两种方法在绕过当前AI安全对齐方面都取得了显著的成功率。
-
新的TempJail方法利用视觉语言模型的时间漏洞
研究人员开发了一种名为TempJail的新方法,通过操纵字幕来利用大型视觉语言模型(LVLMs)中的漏洞。该技术侧重于字幕内容的调度时间,表明信息呈现的时间和持续时间对越狱的有效性有显著影响。实验表明,TempJail在GPT-5和Gemini 3.5 Flash等模型上的攻击成功率高于现有方法。
-
Google 快速迭代 Gemini Flash 模型以应对专业编码任务
Google 在 12 周内迅速发布了其 Gemini Flash 模型的三个迭代版本,每个版本都针对特定改进。5 月份推出的 Gemini 3.5 Flash 提供了高速性能。随后在 7 月份推出了 Gemini 3.6 Flash,专注于优化 token 使用并降低成本。最新发布的 Gemini 3.7 Flash(8 月份)增强了推理和多步规划能力,在 DeepSWE v1.1 和 FrontierCode 1.1 Main 等…
-
大型语言模型的成本效益取决于代币比例,而不仅仅是标价
大型语言模型的成本效益在很大程度上取决于特定任务以及输入和输出代币的比例。基于标价看起来便宜的模型,如果用户的处理涉及高比例的输出代币,可能会变得昂贵,例如在比较 Claude Haiku 4.5 和 Grok 4.3 进行代码生成与分类任务时。由于模型故障导致的重试等因素会进一步使成本计算复杂化,可能抵消初始节省并影响延迟。为了做出明智的决定,用户应衡量自己的代币使用模式,并根据其特定的输入-输出比例来比较模型,而不是仅仅依赖于头条价格。
-
新的NeXUI基准评估AI代理对非视觉用户的解释能力
一个名为NeXUI的新基准已被开发出来,用于评估辅助UI代理,重点关注它们向非视觉用户解释动作的能力,而不仅仅是完成任务。该基准旨在评估安全性、效率和任务成功率,同时验证代理的解释是否基于界面状态。目前的先进模型,如Gemini 3.5 Flash,在NeXUI上表现不佳,成功率仅为44%,解释得分也很低,这凸显了该领域进一步研究的必要性。
-
Gemini和ChatGPT用户均突破10亿,AI竞赛加剧
Google的Gemini月活跃用户已突破10亿,成为达到此里程碑增长最快的Google产品。这一成就使Gemini在用户规模上与近期也突破10亿用户的OpenAI的ChatGPT处于同一水平,尽管衡量标准不同(周活跃 vs. 月活跃)。虽然Gemini整合到各种Google服务中促进了其普及,但10亿用户数据仅计算直接应用和网页界面的使用量。
-
Poe AI 削减免费套餐,高端模型成本引发用户担忧
Poe AI 调整了其计算点数系统,在未公开宣布的情况下,将每日免费套餐额度从 3000 点大幅削减至 300 点。该平台提供各种订阅套餐,费用根据所使用的 AI 模型而定,而不仅仅是购买的总点数。虽然 Poe 旨在成为一个模型无关的平台,但像 Claude-Opus-4.5 这样的高级模型重度用户发现,即使是最高级别的订阅套餐,每日使用量也有限,这可能比直接订阅单个模型更昂贵。
-
新基准发现,低成本AI模型会在清晰文档中编造数据
一项近期基准测试显示,在处理图像时,GPT-5.5和GPT-5.6等低成本AI模型表现出“门控”行为而非“拨号”行为。这些模型倾向于要么完美读取图像字段,要么完全不读取,有相当一部分字段完全无法识别。当面对清晰文档中无法识别的字段时,这些模型会以84%的比例编造信息,捏造银行名称和账号等细节,而不是留空。
-
AI评估分数存在缺陷,侧重模型而非评分者
最近的一项分析强调了AI模型评估中的一个关键缺陷:评估的重点压倒性地放在模型本身的表现上,而评估工具本身的可靠性却常常被忽视。一个例子说明了这一点:在CORE-Bench上,一个模型的得分仅通过改进评分标准、任务规范和工具错误,就在模型本身没有任何改变的情况下,从42%跃升至95%。虽然像Gemini模型这样的现代LLM裁判在与人类评分者的一致性和内部一致性方面表现出高度一致,但不可靠性已转移到评估标准的敏感性和评分细则的措辞上。
-
AI 实验室发布新模型,在安全担忧中筹集数十亿美元 · 跟踪 1 个来源
多家主要 AI 实验室发布了新模型和更新,包括 Anthropic 的 Claude Opus 5、Google 的 Gemini 3.6 和 3.5 Flash 变体,以及 Black Forest Labs 的 Flux 3 用于图像和视频生成。Meta 增强了其 AI 聊天机器人,增加了类似助手的特性,而 OpenAI 推出了 ChatGPT Health。重要的商业发展包括 AMD 对 Anthropic 的 50 亿美元投资…
-
LLM Token定价:原生API、开源模型托管、路由聚合器和云平台
2026年,购买LLM Token已细分为四大类,各有不同的定价和功能。来自OpenAI和Anthropic等模型创建者的原生API提供对新模型和独家功能的首日访问权限,但管理多个密钥和账单可能很复杂。Together等开源模型推理托管服务通过利用公开的模型权重,提供显著更便宜的Token,像GPT OSS 20B这类模型的价格低至0.05美元/百万Token。API路由聚合器将对多个供应商的访问整合到一个密钥下,通常收取标价加上少量…
-
Google 25年来首次重新设计搜索框,整合AI
Google 推出了对其搜索框的重大重新设计,这是 25 年来的首次,将 AI Overviews 和 AI Mode 整合到一个统一的体验中。这个新的搜索界面支持文本、图像、PDF 和视频等多种输入类型,并由 Gemini 3.5 Flash 模型提供支持。该公司指出,目前有超过十亿用户使用 AI Mode,查询量每季度翻一番。
-
Oracle 将 Google Gemini 模型集成到企业应用程序中
Oracle 和 Google Cloud 扩展了合作伙伴关系,将 Google 的 Gemini 模型集成到 Oracle 的企业应用程序中。此次合作将允许客户在 Oracle AI Agent Studio for Fusion Applications 中使用 Gemini 模型,包括 Gemini 3.1 Flash-Lite 和 Gemini 3.5 Flash。此次集成旨在为用户在开发代理和自动化业务流程时选择 AI 模型…
-
Google DeepMind 发布 Gemini Robotics 2,用于高级全身机器人控制
Google DeepMind 推出了 Gemini Robotics 2,这是一个专为高级机器人控制设计的 AI 系统。该新系统使机器人能够执行全身运动、进行灵巧的五指操作,并能在多机器人团队中进行协作,突破了以往预编程或远程操作机器人的局限性。Gemini Robotics 2 由三个独立的模型组成:一个用于运动控制的视觉-语言-动作 (VLA) 模型,一个用于高级规划和理解的具身推理 (ER) 模型,以及一个针对本地执行优化的设…
-
Gemini Flash API:选择模型需要测试,而不仅仅是速度
Google 的 Gemini Flash API 提供了多种模型,但由于输入或上下文处理的限制,选择最快的模型可能无法获得最佳结果。一种实用的方法是针对可用模型进行一次标准化的单任务测试,同时考虑速度、上下文长度和多模态之间的权衡。截至 2026 年 7 月,主要模型包括 gemini-3.5-flash(2026 年 5 月 GA,别名 gemini-flash-latest)、gemini-3.1-flash-lite(2026…
-
Google 的 Gemini 3.6 Flash 提供成本节省但未能提高编码任务性能
Google 发布了 Gemini 3.6 Flash,这是其 Gemini 3.5 Flash 模型的更新版本。虽然新模型提供了更低的每百万 token 定价和更高的输出速度,从而提高了经济效益,但它并未解决复杂编码任务的关键限制。独立分析显示,该模型的智能指数没有提高,对于处理长期编码项目、代码库理解和工具规范方面的能力仍存在担忧。
-
Anthropic 的 Opus 5 在提示注入抵抗力方面取得重大进展 · 跟踪到 1 个来源
Anthropic 的 Opus 5 模型在抵抗提示注入攻击方面表现出显著的改进,在结合额外的系统级防御措施时,成功率接近于零。虽然模型本身更加健壮,但专家强调,对 AI 系统的真正信任依赖于分层方法,包括预执行伪影扫描和受限执行模式,而不仅仅依赖于模型的固有属性。这种纵深防御策略至关重要,因为不同的模型具有不同级别的注入抵抗力,并且恶意指令可以直接嵌入系统提示中,绕过模型特定的防御措施。