A comparison of Russian AI models Alice (Yandex) and GigaChat (Sber) was conducted across five benchmarks: τ³, PRAKT-120, MultiChallenge, WildBench, and Arena-Hard. The testing process was extensive, spanning approximately two weeks due to setup, waiting times, and model instabilities. The author notes that while Marusya AI was intended for inclusion, it is currently undergoing a significant refactor into a sovereign coding environment. Consequently, the results presented focus on Alice, GigaChat Ultra, and GigaChat Max, with Alice accessed via its web connector. AI
IMPACT Provides insights into the performance of Russian AI models, potentially guiding development and adoption.
RANK_REASON Comparison of AI models on benchmarks. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →