A new community-driven benchmark for open-weight language models has been launched, featuring live results for coding and agentic tasks. The benchmark, accessible at beta.locallm.top, aims to provide a comprehensive evaluation across various domains and model types. While DeepSeek-V4-Vision-Exp currently leads in some areas, the project acknowledges limitations such as incomplete coverage for newer models and a developing user interface. The creator is seeking community assistance for compute resources, evaluations, and development to expand the benchmark's capabilities. AI
IMPACT Provides a new platform for evaluating and comparing open-weight LLMs, potentially accelerating development and adoption.
RANK_REASON Community-driven benchmark release for open-weight models. [lever_c_demoted from research: ic=1 ai=1.0]
- DeepSeek-V4-Vision-Exp
- Falcon
- Gemma
- Llama 3
- methylpropyltryptamine
- Mistral AI
- Openchat
- Phi 3
- Qwen
- Qwen3.8-27B
- StableLM
- Zephyr
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →