PulseAugur
EN
LIVE 23:14:05

Poolside's Laguna S 2.1 benchmarked as browser-agent planner

Poolside's new Laguna S 2.1 (118B-A8B) model has been benchmarked as a browser-agent planner. It achieved 71% alignment with Sonnet, performing comparably to Tencent's Hy3 and Minimax M3, but falling short of Gemma 4 31B Q. AI

IMPACT Provides performance data for agent planning tasks, useful for developers choosing models.

RANK_REASON The item benchmarks a specific model's performance on a task, comparing it to other models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Poolside's Laguna S 2.1 benchmarked as browser-agent planner

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Benchmarked Poolside's new Laguna S 2.1 (118B-A8B) as a browser-agent planner. 71% Sonnet alignment. Near Hy3/MiniMax M3 but below them, and below Gemma 4 31B Q

    Benchmarked Poolside's new Laguna S 2.1 (118B-A8B) as a browser-agent planner. 71% Sonnet alignment. Near Hy3/MiniMax M3 but below them, and below Gemma 4 31B QAT / Qwen 3.6 27B (~77%). Best US open-weight in its class, half the size of M3 👇 https://www. webbrain.one/blog/poolsid…