PulseAugur
EN
LIVE 18:34:38

CohereLabs releases 2.4B open-weight vision-language model

CohereLabs has released North-Micro-Vision-Instruct, a 2.4 billion parameter open-weight vision-language model. This model supports native-resolution image processing and is licensed under Apache 2.0, making it suitable for prototyping and specialized multimodal applications. It offers broad image understanding capabilities, including VQA, captioning, and OCR, with support for multiple languages and images. AI

IMPACT Provides a compact, customizable vision-language model for researchers and developers.

RANK_REASON Release of an open-weight model from a research lab. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

CohereLabs releases 2.4B open-weight vision-language model

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/pmttyji ·

    CohereLabs/North-Micro-Vision-Instruct · Hugging Face

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vmjmna/coherelabsnorthmicrovisioninstruct_hugging_face/"> <img alt="CohereLabs/North-Micro-Vision-Instruct · Hugging Face" src="https://external-preview.redd.it/R9TifauXCeIBOuonFF3__kl1e0_4fNHlfv78XrlQe0M.png…