Together AI
PulseAugur coverage of Together AI — every cluster mentioning Together AI across labs, papers, and developer communities, ranked by signal.
- employed by Dan Fu 95%
- instance of Provisioned Throughput 95%
- founded Together 90%
- uses Minimax M3 90%
- uses llama 90%
- used by cubic metre 90%
- used by Nvidia Blackwell B200 90%
- developed GLM-5.2 90%
- invested by Aramco Ventures 90%
- used by IBM Cloud 90%
- developed Artificial Analysis 90%
- invested in Baseten 90%
- 2026-08-19 product_launch Together AI launched 1kpapers.com, a platform that summarizes and visualizes the top 1,000 research papers from the past year. 来源
- 2026-08-19 partnership Together AI has partnered with Higgsfield.ai to provide Dedicated Container Inference services. 来源
- 2026-08-17 partnership Together AI announced a partnership with Relace AI to train specialized small models for code generation. 来源
- 2026-08-13 product_launch Together AI and OpenRouter are co-hosting a meetup in New York City on August 20th. 来源
- 2026-08-12 product_launch Together AI has made the Qwen3.8-2.4T-A95B model available via its Serverless Inference service. 来源
- 2026-08-11 partnership Together AI entered into a $240 million deal with IBM Cloud to fund infrastructure deployments. 来源
- 2026-08-11 partnership Together AI and IBM Cloud have entered into a $240 million infrastructure deal to deploy AI workloads. 来源
- 2026-08-11 partnership Together AI, IBM, and NVIDIA have entered into a multi-year agreement to scale AI inference on IBM Cloud. 来源
- 2026-08-06 partnership Together AI partnered with Roomote to integrate Together AI as an inference provider within Roomote. 来源
- 2026-08-05 research_milestone Together AI achieved a 10,000x increase in token serving capacity, scaling from 30 billion to 400 trillion tokens per month. 来源
- 2026-07-29 product_launch Together AI launched a system connecting LLMs with robot policies to enable robots to think and move. 来源
- 2026-07-29 partnership Together AI and Moonshot AI announced a strategic partnership where Together AI will serve as the launch platform for Moonshot's open-weight models, starting with Kimi K3. 来源
- 2026-07-23 product_launch Together AI launched an updated inference platform for open-weight AI models. 来源
- 2026-07-20 partnership Together AI and Y Combinator partnered to launch a dedicated GPU cluster for YC startups. 来源
- 2026-07-20 partnership Together AI and Y Combinator partnered to launch a dedicated GPU cluster for YC startups. 来源
24 天有情绪数据
Together AI significantly bolsters inference capacity with H100/H200 GPU expansion
The addition of one thousand NVIDIA H100 and H200 GPUs to Together AI's infrastructure represents a substantial investment in inference capabilities. This move directly supports the growing demand for high-throughput AI model serving and is likely intended to power both their internal services and external customer workloads.
Together AI to offer ATLAS as a distinct inference optimization service
Given the significant performance gains demonstrated by ATLAS, Together AI may soon offer this adaptive-learning inference system as a standalone service or an add-on feature for their existing GPU offerings. This would allow customers to leverage ATLAS's dynamic optimization without needing to manage the underlying infrastructure themselves.
Together AI's ATLAS system demonstrates superior inference speed on par with specialized hardware
Together AI's newly launched ATLAS system, an adaptive-learning inference engine, is showing remarkable performance, achieving up to 500 TPS on DeepSeek-V3.1. This performance rivals that of specialized hardware like Groq, suggesting Together AI is effectively optimizing LLM inference beyond standard GPU capabilities.
Together AI to integrate NVIDIA Blackwell features into all core services
The 90% training speed boost achieved with NVIDIA Blackwell and custom kernels indicates a deep integration. It's likely Together AI will leverage Blackwell's capabilities across their entire platform, including their new instant clusters and fine-tuning services, to offer a performance edge over competitors.
Together AI's ATLAS system shows strong performance against specialized hardware
The reported performance of Together AI's ATLAS system, achieving up to 500 TPS on DeepSeek-V3.1 and outperforming specialized hardware like Groq, is a significant technical achievement. This suggests their adaptive inference approach is highly effective and could set a new benchmark for LLM inference speed and efficiency.
-
Together AI 将在网络研讨会中详细介绍 Kimi K3 生产部署
Together AI 将于 8 月 26 日星期三举办一场网络研讨会,重点介绍其 Kimi K3 模型的生产部署。会议将邀请 Together AI 的推理团队成员 Kevin Cui 和 Austin Silveria,他们将讨论 Kimi K3 服务所涉及的基础架构、质量保证流程和优化。
-
Together AI 的 GLM-5.3 在效率基准测试中优于 Fable 5
Together AI 的 GLM-5.3 模型在效率方面明显优于 Fable 5,在相同的 100 美元预算下,其完成的工作量是 Fable 5 的五倍以上。在 DeepSWE 基准测试中,GLM-5.3 解决了约 17 个任务,而 Fable 5 只解决了 3 个,尽管在初步尝试时性能相似。
-
Together AI 将举办 Kimi k3 生产网络研讨会 · 跟踪 4 个来源
Together AI 将于 8 月 26 日举办一场网络研讨会,其推理团队的成员将讨论 Kimi k3 的生产部署。会议将涵盖服务 Kimi k3 模型所采用的架构、质量保证和优化策略。本次活动是“推理时段”系列活动的一部分,旨在深入了解部署大型语言模型的实际方面。
-
GLM-5.3 在编码任务上挑战 GPT-5.6 Sol 和 Claude Fable 5
Together AI 在 DeepSWE 软件工程任务上对其 GLM-5.3 模型与 OpenAI 的 GPT-5.6 Sol 和 Anthropic 的 Claude Fable 5 进行了基准测试。GLM-5.3 表现出竞争力,在单次准确率上略微落后于 GPT-5.6 Sol,但在多次尝试和显著更低的成本下超越了它。与 Claude Fable 5 相比,GLM-5.3 在准确率上相当,但成本却低了五倍多,两种模型都表现出相似的…
-
Together AI 推广关于设计更好 AI 代理用户界面的讨论
Together AI 正在推广 nutlope 关于如何为 AI 代理创建视觉吸引力的用户界面的演讲。本次讨论旨在通过关注设计原则来指导开发者构建更好的 AI 应用。
-
Together AI 和 OpenRouter 在纽约举办开发者见面会
Together AI 与 OpenRouter 合作,将于 8 月 20 日在纽约举办开发者见面会。活动将重点讨论开放模型、模型选择策略以及人工智能采用的未来轨迹。
-
Together AI推出1kpapers.com,提供DeepSeek-V4 Flash摘要
Together AI 推出了 1kpapers.com,该平台可对过去一年的 1,000 篇顶级研究论文进行摘要和可视化。该项目利用 DeepSeek-V4 Flash 进行摘要,整个过程仅花费 4 美元。此举旨在让前沿研究更加易于获取。
-
Together AI 与 Higgsfield.ai 达成合作,提供推理服务
Together AI 已将 Higgsfield.ai 确立为新客户,利用其专用容器推理服务。Higgsfield.ai 以其AI视频和图像创作平台 Cinema Studio 而闻名,将利用 Together AI 的基础设施来处理其长时间运行、多GPU推理任务。此次合作是在 Higgsfield.ai 最近完成 4 亿美元 B 轮融资之后达成的。
-
Together AI 在 Y Combinator 支持下简化 GPU 访问
Together AI 与 Y Combinator 合作,为其团队简化了 GPU 访问。该计划确保 GPU 可根据要求随时提供,无需长期预订或等待。该系统旨在提供所需的即时计算资源。
-
AI 代理基础设施快速增长,Together AI 领先
正如 Brex 的夏季基准所示,支持 AI 代理经济的基础设施正在经历快速增长。Together AI 已成为服务于构建 AI 产品团队的最快增长的软件供应商。初创公司越来越多地采用开源模型,其速度远超往年,Together AI 在产品和代理的代币处理方面看到了大幅增长。
-
IBM 和 Together AI 签署 2.4 亿美元协议,在 IBM Cloud 上构建 AI 推理集群
IBM 和 Together AI 已达成一项为期多年、价值 2.4 亿美元的协议,将在 IBM Cloud 上建立一个重要的大规模 AI 推理集群。该集群预计将于 2027 年第一季度投入使用,将利用 NVIDIA 先进的 HGX 300 系统和 Spectrum-X 网络。其主要目的是为企业客户提供开源 AI 模型推理的便利。
-
Together AI 强调 NVIDIA 在供应受限背景下的 AI 基础设施倡议
Together AI 正在重点介绍 NVIDIA 的一项倡议,该倡议正在“循环融资”的背景下进行讨论。对话指出,对 AI 基础设施的需求正迅速超过供应,导致价格上涨和需求破坏。
-
Together AI 与 Relace AI 合作推出经济高效的代码生成模型
Together AI 正与 Relace AI 合作,利用专用的 Y Combinator GPU 集群训练专门的、经济高效的小型代码生成模型。这种方法旨在通过专注于定制模型而非更大、更通用的模型来提供显著的成本节约。
-
Together AI发布GLM 5.3,用于编码和网络安全
Together AI发布了GLM 5.3,这是一款专为编码和网络安全任务设计的模型。在演示中,GLM 5.3在创建3D自行车网站方面表现出色,在质量和成本效益方面均优于Fable 5。新模型基于743B基础构建,并在处理复杂编码挑战和加强网络安全防御能力方面取得了显著进步。
-
OpenRouter替代品:TokenPAPA在中文大模型方面领先,Groq在速度方面领先
多个平台提供了访问大语言模型的OpenRouter替代方案,各有优势。TokenPAPA因其以有竞争力的价格访问DeepSeek V4 Flash和Mimo V2.5等中文大模型,以及简化的注册流程而受到关注。DeepInfra被认为是开源模型经济实惠的选择,而Together AI则提供更全面但价格更高的全栈解决方案,包括微调。Groq凭借其LPU硬件专注于速度,适用于延迟敏感的应用,而Fireworks AI则提供了快速服务和微调能力的结合。
-
Together AI 支持在生产环境中对 LLM 端点进行 A/B 测试
Together AI 推出了一项新功能,允许用户直接在 LLM 的端点上进行 A/B 测试。此功能使开发人员能够将实时流量分配给一个对照模型和最多 20 个变体,并根据用户互动而非基准来衡量实际性能。该平台在端点级别管理流量路由和用户群划分,简化了流程,并防止实验逻辑与应用程序代码纠缠在一起。
-
Together AI 的 Navigator 模型提供更快、更便宜的浏览器自动化
Together AI 发布了其 Navigator 模型,该模型专为基于浏览器的代理而设计。该模型以快速的截图-操作-重复循环来完成任务。Together AI 声称 Navigator 在性能上优于前沿模型,推理速度提高一倍,成本显著降低,便宜四到五倍。
-
俄媒:俄导弹据称使用英伟达AI芯片瞄准乌克兰
据报道,一枚俄罗斯导弹配备了来自英伟达(NVIDIA)的AI芯片,以辅助瞄准乌克兰。这一事态发展促使基辅呼吁加强对外国硅芯片出口到莫斯科军方的管控。导弹制导系统中使用的具体芯片及其作用尚未完全详述,但此次事件凸显了对先进技术在持续冲突中应用的担忧。
-
Anthropic 将为 AI 输出添加水印;Mojo 1.0 发布;Together AI 达成 2.4 亿美元 IBM 交易
Anthropic 已承诺为其 AI 生成的内容实施水印,以帮助区分人类创作的材料,主要动机是欧盟的法规。此举旨在追溯 AI 输出的来源并遵守新兴规则。与此同时,Modular 发布了其 Mojo 编程语言的 1.0 版本,开发者们期待开源编译器发布,以澄清高通收购该公司后的不确定性。此外,Together AI 已与 IBM Cloud 达成 2.4 亿美元的交易,为其计划于 2027 年初部署的大规模 Nvidia HGX B30…
-
阿里巴巴 Qwen3.8-Max 模型发布,拥有 2.4T 参数、1M 上下文 · 跟踪 2 个来源
阿里巴巴的 Qwen 团队发布了 Qwen3.8-Max,这是一个拥有 2.4 万亿参数和 950 亿激活参数的大型混合专家模型。该新模型专为自动代理、重度编码和长上下文处理等要求苛刻的任务而设计,拥有高达 100 万个 token 的上下文窗口。Qwen3.8-Max 已在 Fireworks 和 Together AI 平台上线,这两个平台被列为首发合作伙伴。