togethercompute
PulseAugur coverage of togethercompute — every cluster mentioning togethercompute across labs, papers, and developer communities, ranked by signal.
- 2026-07-02 research_milestone Together AI announced significant improvements in inference speed and cost reduction, alongside a commitment to weekly model releases. source
-
Together AI signals upcoming announcement, urges user engagement
Together AI has announced a new development, urging users to keep their notifications active for upcoming updates. The specific nature of this announcement remains undisclosed, but it suggests a significant upcoming eve…
-
MiniMax AI emphasizes open models at Paris AI event
MiniMax AI participated in RAISE week in Paris, highlighting the growing importance of open models in frontier AI development. The company's president discussed open models and multimodality, while MiniMax also hosted p…
-
MiniMax AI to host executive gathering at RAISE Week
MiniMax AI is participating in RAISE Week, an executive gathering focused on frontier multimodal AI. The company will host a private event on July 9th featuring executives from TogetherCompute, Cast AI, and Magnific AI …
-
Together AI achieves 6x cost reduction and sub-400ms latency
Together AI has announced significant improvements to its inference capabilities, achieving a sixfold reduction in cost per turn and a p95 latency under 400 milliseconds. The company is also committed to shipping new mo…
-
Together and MiniMax AI discuss AI optimizations
Together and MiniMax AI are co-hosting a talk discussing sparse attention and kernel optimizations. The event, promoted on X, aims to share insights on these advanced AI techniques.
-
Together AI releases free Brrrrr inference model
Together AI has released Brrrrr, a new inference model that is available for free use. Early benchmarks show the model achieving 131 tokens per second.
-
Together AI offers fast GLM-5.2 inference with optimized serving
Together AI is now offering GLM-5.2, a model that is reportedly fast and capable of handling long-context coding and agent workloads. The company emphasizes its optimized serving infrastructure, which allows for high th…