PulseAugur
EN
LIVE 08:28:02

New low-bit and ternary AI models released with performance updates · 1 source tracked

Several new low-bit and ternary models have been released and are being tracked, including Bonsai's 1-bit and 1.58-bit (ternary) versions, with a 27B parameter model now running on mainline backends. Updates to llama.cpp have improved CUDA and Vulkan performance for these models. Other notable releases include BitCPM-CANN, Tencent's Hy-MT1.5, DeepGrove's Maple-Preview, SyzygyResearch's Mach-1-Additive-35B, FermionResearch's Neutrino-8B, and Doses-AI's Pestle-27B-Ternary, with some models achieving high inference speeds on consumer hardware. AI

IMPACT These low-bit and ternary models offer potential for more efficient AI deployment on consumer hardware.

RANK_REASON The item discusses the release and tracking of various low-bit and ternary AI models, including performance updates and new versions, which falls under research and model releases. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New low-bit and ternary AI models released with performance updates · 1 source tracked

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/pmttyji ·

    1-bit / 2-bit / Ternary / Bitnet Models - Updates & Tracking

    <!-- SC_OFF --><div class="md"><h1>Bonsai / Ternary Bonsai</h1> <p>During April Bonsai came with bunch of models .... <a href="https://huggingface.co/collections/prism-ml/bonsai">1-bit</a> &amp; <a href="https://huggingface.co/collections/prism-ml/ternary-bonsai">1.58-bit(Ternary…