PulseAugur
EN
LIVE 19:31:40

Poolside releases Laguna S 2.1 with 256K context, achieving high token speeds

Poolside has released Laguna S 2.1, a new model that fits within 256K context at F16 quantization. Early performance tests show impressive speeds, with prefill reaching 400-600 tokens/sec and generation at 16-20 tokens/sec. These results were achieved using three V620 Aurigae GPUs, providing a total of 96 GB of VRAM, in a Dell PowerEdge R740 server. AI

IMPACT This release offers a new option for users needing large context windows and high token generation speeds.

RANK_REASON New model release from a frontier lab. [lever_c_demoted from frontier_release: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Poolside releases Laguna S 2.1 with 256K context, achieving high token speeds

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/_TheWolfOfWalmart_ ·

    Today was the perfect day for Poolside to drop Laguna S 2.1 because I just got these in! Finally have a half decent amount of VRAM. 3x V620 = 96 GB.

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1v353sh/today_was_the_perfect_day_for_poolside_to_drop/"> <img alt="Today was the perfect day for Poolside to drop Laguna S 2.1 because I just got these in! Finally have a half decent amount of VRAM. 3x V620 =…