Jebadiah v2.1, a new set of open-weight models, has been released with versions at 27B and 9B parameters. These models are designed to score every allowed label from logits, treating decisions as a closed-set scoring problem rather than text generation. The v2.1 updates show improved performance on the Decision Index benchmark, particularly in knowledge, language, and retrieval tasks, though a regression was noted in tool-use capabilities. AI
IMPACT Provides new open-weight models for decision-scoring tasks, with benchmark results and code available for further research.
RANK_REASON Release of open-weight models with benchmark results and code. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →