Together AI
PulseAugur coverage of Together AI — every cluster mentioning Together AI across labs, papers, and developer communities, ranked by signal.
- 2026-09-16 funding Together AI announced a $800 million funding round at an $8.3 billion valuation. source
- 2026-09-11 product_launch Together AI launched an expanded fine-tuning service with new models, live metrics, and advanced controls. source
- 2026-09-10 product_launch Together AI launched a public preview of preemptible compute for its GPU Clusters. source
- 2026-09-09 partnership Together AI and MiniMax AI are collaborating to host an event discussing open-source AI deployment. source
- 2026-09-04 product_launch Together AI released a guide for deploying a chat API using their platform and Render. source
- 2026-09-01 partnership Together AI formed a multi-billion dollar alliance integrating Middle Eastern infrastructure with global open-source software. source
- 2026-08-31 partnership Together AI announced Greptile as a new customer and its primary inference provider. source
- 2026-08-31 partnership Together AI is partnering with Humain to build a 250-megawatt data center in Saudi Arabia. source
- 2026-08-31 partnership Together AI has partnered with Saudi Arabia and humain to build a new data center. source
- 2026-08-19 product_launch Together AI launched 1kpapers.com, a platform that summarizes and visualizes the top 1,000 research papers from the past year. source
- 2026-08-19 partnership Together AI has partnered with Higgsfield.ai to provide Dedicated Container Inference services. source
- 2026-08-17 partnership Together AI announced a partnership with Relace AI to train specialized small models for code generation. source
- 2026-08-13 product_launch Together AI and OpenRouter are co-hosting a meetup in New York City on August 20th. source
- 2026-08-12 product_launch Together AI has made the Qwen3.8-2.4T-A95B model available via its Serverless Inference service. source
- 2026-08-11 partnership Together AI entered into a $240 million deal with IBM Cloud to fund infrastructure deployments. source
19 day(s) with sentiment data
Together AI significantly bolsters inference capacity with H100/H200 GPU expansion
The addition of one thousand NVIDIA H100 and H200 GPUs to Together AI's infrastructure represents a substantial investment in inference capabilities. This move directly supports the growing demand for high-throughput AI model serving and is likely intended to power both their internal services and external customer workloads.
Together AI to offer ATLAS as a distinct inference optimization service
Given the significant performance gains demonstrated by ATLAS, Together AI may soon offer this adaptive-learning inference system as a standalone service or an add-on feature for their existing GPU offerings. This would allow customers to leverage ATLAS's dynamic optimization without needing to manage the underlying infrastructure themselves.
Together AI's ATLAS system demonstrates superior inference speed on par with specialized hardware
Together AI's newly launched ATLAS system, an adaptive-learning inference engine, is showing remarkable performance, achieving up to 500 TPS on DeepSeek-V3.1. This performance rivals that of specialized hardware like Groq, suggesting Together AI is effectively optimizing LLM inference beyond standard GPU capabilities.
Together AI to integrate NVIDIA Blackwell features into all core services
The 90% training speed boost achieved with NVIDIA Blackwell and custom kernels indicates a deep integration. It's likely Together AI will leverage Blackwell's capabilities across their entire platform, including their new instant clusters and fine-tuning services, to offer a performance edge over competitors.
Together AI's ATLAS system shows strong performance against specialized hardware
The reported performance of Together AI's ATLAS system, achieving up to 500 TPS on DeepSeek-V3.1 and outperforming specialized hardware like Groq, is a significant technical achievement. This suggests their adaptive inference approach is highly effective and could set a new benchmark for LLM inference speed and efficiency.
-
Together AI raises $800M at $8.3B valuation
Together AI has secured $800 million in funding at an $8.3 billion valuation. The company's product team reportedly shared details about their operational repositories, including discussions on output backfiring and the…
-
Together AI outlines strategy for migrating to open-source models
Together AI's blog post outlines a strategy for migrating from closed-source to open-source AI models, emphasizing that such migrations can be faster and less complex than traditional ones, especially when utilizing man…
-
Together AI adds MiniMax H3 omni-modal video generation model
Together AI has announced the availability of MiniMax H3, a 33 billion parameter omni-modal video model. This model is capable of generating video clips ranging from 4 to 15 seconds at resolutions up to 2K, with native …
-
Deel fine-tunes models on Together AI for global compliance
Deel has developed a fine-tuned model, built on Together AI's infrastructure, to handle compliance, payroll, and HR inquiries across 150 countries. This approach allows Deel to provide faster answers by leveraging its i…
-
Together AI shares inference engine talk slides · 1 source tracked
Together AI has shared the full talk slides from zainhas regarding the inner workings of inference engines. The repost on X highlights the technical details of how these systems function.
-
Together AI shows top-tier performance for agentic workloads on OpenRouter
Together AI is showcasing strong performance for agentic workloads, serving GLM 5.3 and GLM 5.3 Flash models. According to data from OpenRouter, Together AI's models are performing at the top decile for metrics like tra…
-
Together AI releases Kimi K3 developer guide
Together AI has released a comprehensive developer guide for its Kimi K3 model. The guide aims to provide developers with the necessary information to effectively utilize the Kimi K3 model for various applications.
-
Together AI expands fine-tuning with new models and live tracking
Together AI has enhanced its fine-tuning service by incorporating a wider array of open-weight models, including advanced options like GLM 5.3 and Kimi K2.7, alongside cost-effective choices such as Qwen 3.8-27B and Gem…
-
Together AI optimizes ThunderKittens for NVIDIA Vera Rubin Blackwell GPUs
Together AI has gained access to NVIDIA's Vera Rubin NVL72 platform, which is based on the Blackwell architecture. Their team has updated their ThunderKittens software to leverage new features of the Vera Rubin chip, sp…
-
Together AI offers preemptible GPU compute at 50% discount
Together AI has launched a public preview of preemptible compute for its GPU Clusters, available on Kubernetes. This new offering allows users to access the same GPU infrastructure at half the on-demand price for interr…
-
MiniMax AI highlights community optimization of open-source H3 model
MiniMax AI is highlighting community contributions to its open-source H3 model, noting that after its initial release at 28 steps, the community has developed optimized versions. These accelerated versions have performe…
-
Together AI and MiniMax AI to discuss open-source AI economics in London
Together AI and MiniMax AI are collaborating to host an event in London focused on the economics of deploying open-source AI models. The discussion will cover cost-effectiveness, performance optimization, model selectio…
-
DeepSeek releases V4.1 Flash with efficient MoE architecture
DeepSeek has officially released its V4.1 Flash model, a 552 billion parameter Mixture-of-Experts (MoE) model featuring a Causal-Encoder-Decoder (CED) architecture and native multimodal capabilities. This new model is d…
-
MiniMax API pricing clarified, resellers largely match official rates
MiniMax's API pricing for its M3 and M2.7 models has been clarified, revealing that the standard rate for M3 is $0.30 per million input tokens and $1.20 per million output tokens, with a stated permanent 50% discount fr…
-
Together AI's GLM-5.3 Flash outperforms Claude Fable 5.1 on automation tasks
Together AI's GLM-5.3 Flash model has demonstrated superior performance in agentic automation tasks compared to Anthropic's Claude Fable 5.1. The new model achieves this feat at a significantly lower cost, reportedly 99…
-
Together AI hosts Nvidia, DTCP Capital during SF Tech Week
Together AI is hosting an event during SF Tech Week, coinciding with Fleet Week, featuring Nvidia and DTCP Capital. The event will take place at a private waterfront chalet with a view of the Blue Angels air show. The t…
-
Kimi K2 model pricing varies widely across platforms, impacting total cost
The pricing for Moonshot's Kimi K2 model varies significantly across different platforms, with output costs ranging from $3.20 to $4.50 per million tokens for the same model variant. This price discrepancy arises becaus…
-
Together AI's GLM-5.3 Flash matches GPT-5.6 Terra performance at lower cost
Together AI has released GLM-5.3 Flash, which matches the performance of GPT-5.6 Terra on Artificial Analysis's intelligence index. Notably, GLM-5.3 Flash achieves this comparable performance at an 82% lower cost per ta…
-
Together AI highlights token burn in AI inference on X
Together AI has reposted a message on X (formerly Twitter) from user Zainab that discusses work being measured in tokens burned. This implies a focus on the computational cost and resource consumption associated with AI…
-
Adaption Labs launches API to generate AI training data from task descriptions
Adaption Labs has launched 'Invent a Dataset,' a new feature that generates training data directly from a task description, eliminating the need for a seed corpus, schema, or manual labeling. This tool aims to improve m…