AgiBot WITA-Omni, a new full-modal AI model, has achieved the top position on the DailyOmni Global Leaderboard for embodied cross-modal understanding. The model outperformed competitors such as Google Gemini, ByteDance Doubao, and Alibaba Qwen. AgiBot WITA-Omni utilizes a Thinker-Talker-Actor architecture, which synchronizes speech, action, and expression on a unified timeline, and secured first place in six out of eight performance indicators. AI
IMPACT Sets a new benchmark for embodied cross-modal understanding, potentially influencing future multimodal AI development.
RANK_REASON The cluster reports on a new AI model achieving a top score on a specific benchmark, surpassing competitors. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →