Together AI
PulseAugur coverage of Together AI — every cluster mentioning Together AI across labs, papers, and developer communities, ranked by signal.
- employed by Dan Fu 95%
- instance of Provisioned Throughput 95%
- founded Together 90%
- uses llama 90%
- uses Minimax M3 90%
- used by cubic metre 90%
- developed Artificial Analysis 90%
- used by Nvidia Blackwell B200 90%
- developed GLM-5.2 90%
- developed DeepSeek V4-Pro 90%
- partners with Dan Fu 90%
- partners with Kimi K2 90%
- 2026-08-11 partnership Together AI entered into a $240 million deal with IBM Cloud to fund infrastructure deployments. source
- 2026-08-11 partnership Together AI and IBM Cloud have entered into a $240 million infrastructure deal to deploy AI workloads. source
- 2026-08-11 partnership Together AI, IBM, and NVIDIA have entered into a multi-year agreement to scale AI inference on IBM Cloud. source
- 2026-08-06 partnership Together AI partnered with Roomote to integrate Together AI as an inference provider within Roomote. source
- 2026-08-05 research_milestone Together AI achieved a 10,000x increase in token serving capacity, scaling from 30 billion to 400 trillion tokens per month. source
- 2026-07-29 product_launch Together AI launched a system connecting LLMs with robot policies to enable robots to think and move. source
- 2026-07-29 partnership Together AI and Moonshot AI announced a strategic partnership where Together AI will serve as the launch platform for Moonshot's open-weight models, starting with Kimi K3. source
- 2026-07-23 product_launch Together AI launched an updated inference platform for open-weight AI models. source
- 2026-07-20 partnership Together AI and Y Combinator partnered to launch a dedicated GPU cluster for YC startups. source
- 2026-07-20 partnership Together AI and Y Combinator partnered to launch a dedicated GPU cluster for YC startups. source
- 2026-07-18 funding Together AI closed a $800 million Series C funding round at an $8.3 billion valuation, led by Aramco Ventures. source
- 2026-07-17 product_launch Together AI launched Provisioned Throughput for its MiniMax M3 model, offering guaranteed capacity and lower costs. source
- 2026-07-11 product_launch Together AI launched a new product feature called "Call" that enables voice interaction with AI models. source
- 2026-07-08 product_launch Together AI launched Provisioned Throughput, a new serverless product offering guaranteed inference capacity. source
- 2026-07-08 product_launch Together AI introduced Provisioned Throughput, a service for reserved inference capacity for open models. source
26 day(s) with sentiment data
Together AI significantly bolsters inference capacity with H100/H200 GPU expansion
The addition of one thousand NVIDIA H100 and H200 GPUs to Together AI's infrastructure represents a substantial investment in inference capabilities. This move directly supports the growing demand for high-throughput AI model serving and is likely intended to power both their internal services and external customer workloads.
Together AI to offer ATLAS as a distinct inference optimization service
Given the significant performance gains demonstrated by ATLAS, Together AI may soon offer this adaptive-learning inference system as a standalone service or an add-on feature for their existing GPU offerings. This would allow customers to leverage ATLAS's dynamic optimization without needing to manage the underlying infrastructure themselves.
Together AI's ATLAS system demonstrates superior inference speed on par with specialized hardware
Together AI's newly launched ATLAS system, an adaptive-learning inference engine, is showing remarkable performance, achieving up to 500 TPS on DeepSeek-V3.1. This performance rivals that of specialized hardware like Groq, suggesting Together AI is effectively optimizing LLM inference beyond standard GPU capabilities.
Together AI to integrate NVIDIA Blackwell features into all core services
The 90% training speed boost achieved with NVIDIA Blackwell and custom kernels indicates a deep integration. It's likely Together AI will leverage Blackwell's capabilities across their entire platform, including their new instant clusters and fine-tuning services, to offer a performance edge over competitors.
Together AI's ATLAS system shows strong performance against specialized hardware
The reported performance of Together AI's ATLAS system, achieving up to 500 TPS on DeepSeek-V3.1 and outperforming specialized hardware like Groq, is a significant technical achievement. This suggests their adaptive inference approach is highly effective and could set a new benchmark for LLM inference speed and efficiency.
-
Nvidia backs open-source AI, fueling hardware spending; IBM partners for inference
Nvidia is emphasizing its support for open-source AI initiatives, which is expected to boost hardware spending. The company's Nemotron 3.5 Lightning model, an open-weight system with 3.6 billion parameters, demonstrates…
-
Together AI inks $240M IBM Cloud deal for Nvidia hardware
Together AI has secured a significant $240 million deal with IBM Cloud to deploy AI workloads. This partnership will involve a large-scale deployment of NVIDIA's HGX B300 systems, expected to launch in early 2027. The a…
-
Together AI partners with IBM and NVIDIA for enterprise AI inference
Together AI has announced a multi-year partnership with IBM and NVIDIA to enhance AI inference capabilities on IBM Cloud. This collaboration will leverage a dedicated cluster of NVIDIA B300 GPUs and Spectrum-X networkin…
-
IBM and Together AI partner on $240M Nvidia-powered AI cluster for open-source models
IBM and Together AI have announced a partnership to build a $240 million AI inference cluster. This cluster will be powered by NVIDIA hardware and is intended to support open-source models. The initiative aims to help b…
-
NVIDIA launches Nemotron 3.5 Lightning for efficient agentic AI · 10 sources tracked
NVIDIA has launched Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model designed for efficient, high-volume agentic AI tasks. This new model offers up to 4x faster throughput and 30% faster task comp…
-
Together AI launches Meta's Muse Glimmer for agentic tasks · 3 sources tracked
Together AI has launched Muse Glimmer, an open-weight model developed by Meta Superintelligence Labs. This model is designed for complex, long-running agentic tasks, capable of reasoning, using tools, and recovering fro…
-
Meta releases open-weight Muse Glimmer; Anthropic, OpenAI advance frontier capabilities · 1 source tracked
Meta has re-entered the open-weight model release arena with Muse Glimmer, a 30B multimodal model optimized for local agents and consumer hardware deployment. This release, announced by Mark Zuckerberg, emphasizes long-…
-
Cursor partners with Together AI for real-time coding inference
Cursor has partnered with Together AI to enhance its AI-powered coding capabilities. Together AI has developed the necessary infrastructure to provide real-time inference, enabling Cursor's in-editor agents to generate …
-
Together AI leads Kimi K3 inference benchmarks
Together AI has been recognized as a top-tier inference provider, outperforming competitors in benchmarks for the Kimi K3 model. The company achieved the highest or tied-for-highest ranking on three out of four key perf…
-
Together AI showcases Seedance 2.5's trailer generation capabilities
Together AI has showcased the capabilities of Seedance 2.5 by generating a 30-second "lost-cinema" trailer using a single prompt. This demonstration highlights the model's ability to produce complex video content from c…
-
Autoscaling LLM inference workloads requires specialized approaches
Autoscaling inference workloads for large language models (LLMs) presents unique challenges compared to traditional web services. The nature of peaky LLM inference demands specialized approaches to efficiently manage fl…
-
Together AI launches FLUX 3 multimodal video generation model · 3 sources tracked
Together AI has launched FLUX 3, a new multimodal model developed by bfl_ai. FLUX 3 is capable of generating video with synchronized audio, supporting up to 20-second clips, multiple scenes, and control via text, images…
-
Kimi K3 outperforms Claude Fable-5 on legal tasks · 2 sources tracked
Together AI's Kimi K3 model has demonstrated superior performance compared to Anthropic's Claude Fable-5 on complex legal tasks. In evaluations conducted by Harvey LAB-AA, Kimi K3 achieved nearly double the score of Cla…
-
Together AI integrates with Roomote for enhanced agent workflows
Together AI has partnered with Roomote to enable builders to use Together AI as an inference provider within Roomote. This integration allows for the assignment of different open models to specific tasks such as coding,…
-
DeepSeek-V4 Flash challenges GPT-5.6 Luna on coding benchmark with cost-efficiency
Together AI has released a comparative analysis of DeepSeek-V4 Flash and GPT-5.6 Luna on the DeepSWE coding benchmark. While GPT-5.6 Luna demonstrates superior performance across all quality metrics, DeepSeek-V4 Flash p…
-
Together AI teases upcoming announcement
Together AI is announcing a new development next week. The company shared a cryptic message on X, formerly Twitter, indicating that "Next week on Together AI" will bring significant news. Further details about the natur…
-
Together AI offers Kimi K3 access via Together Chat
Together AI is offering access to its Kimi K3 model through its Together Chat platform. Users can directly prompt the Kimi K3 model without needing to set up an API. The service is hosted on secure North American infras…
-
AI services offer 'no-login' access, but privacy varies greatly
Several services offer access to AI models without requiring user registration, but true privacy depends on data handling rather than just login requirements. Duck.ai stands out by detailing its privacy mechanisms, incl…
-
Together AI sees 10,000x growth in token serving to 400T/month
Together AI has experienced exponential growth, scaling from serving 30 billion tokens per month to 400 trillion tokens. This massive increase, over 10,000x, is attributed to AI-native companies and enterprises migratin…
-
Together AI targets 99.9% uptime with multi-data-center inference strategy
Together AI is implementing a multi-data-center deployment strategy to ensure 99.9% uptime for its inference services. This approach involves running live traffic across multiple facilities and maintaining sufficient ca…