A comparison of Russian AI models Alice (Yandex) and GigaChat (Sber) was conducted across five benchmarks: τ³, PRAKT-120, MultiChallenge, WildBench, and Arena-Hard. The testing process was extensive, spanning approximately two weeks due to setup, waiting times, and model instabilities. The author notes that while Marusya AI was intended for inclusion, it is currently undergoing a significant refactor into a sovereign coding environment. Consequently, the results presented focus on Alice, GigaChat Ultra, and GigaChat Max, with Alice accessed via its web connector. AI
影响 Provides insights into the performance of Russian AI models, potentially guiding development and adoption.
排序理由 Comparison of AI models on benchmarks. [lever_c_demoted from research: ic=1 ai=1.0]
在 Mastodon — mastodon.social 阅读 →
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →