PulseAugur
EN
LIVE 19:21:56

Together launches next-gen inference platform for open models

Together has launched the next generation of its inference platform, designed to help users run open models in production with enhanced control and safety. The platform allows for testing changes on live traffic without impacting users, swapping models behind a single endpoint, and autoscaling based on real-time signals like TTFT and latency. This update is informed by Together's experience serving over 400 trillion tokens monthly. AI

IMPACT Enhances production deployment and testing capabilities for open-source AI models.

RANK_REASON Product launch for an inference platform, not a frontier model release.

Read on X — Together (inference / OSS) →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

Together launches next-gen inference platform for open models

COVERAGE [4]

  1. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    Try it yourself: https://t.co/2ilKtqodxX

    Try it yourself: https://t.co/2ilKtqodxX

  2. X — Together (inference / OSS) TIER_1 Dansk(DA) · togethercompute ·

    Blog: https://t.co/M2swebEtKW

    Blog: https://t.co/M2swebEtKW Webinar (Aug 6): https://t.co/ote2Bgw5xU

  3. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    What's new:

    What's new: -> Test on real traffic, zero user impact -> Swap models behind one stable endpoint -> Autoscale on real signals: TTFT, latency, throughput -> Pre-optimized deployment profiles

  4. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    The next generation of our inference platform is here, shaped by what we've learned serving over 400 trillion tokens a month.

    The next generation of our inference platform is here, shaped by what we've learned serving over 400 trillion tokens a month. Run open models in production with full control, and test every change on live traffic before users see it. https://t.co/5bIT1y7BOP