Poolside has released Laguna S 2.1, a new model that fits within 256K context at F16 quantization. Early performance tests show impressive speeds, with prefill reaching 400-600 tokens/sec and generation at 16-20 tokens/sec. These results were achieved using three V620 Aurigae GPUs, providing a total of 96 GB of VRAM, in a Dell PowerEdge R740 server. AI
IMPACT This release offers a new option for users needing large context windows and high token generation speeds.
RANK_REASON New model release from a frontier lab. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →