H Company has introduced NeoMME, a new family of multimodal encoders designed for efficient document retrieval. These models, available in 260M and 800M parameter sizes, uniquely integrate text and image processing within a single transformer tower, eliminating the need for separate vision towers and causal decoders. The NeoMME-Retriever variant demonstrates competitive performance on benchmarks like ViDoRe v3, achieving strong results with significantly fewer parameters than other leading models. The checkpoints are released under Apache 2.0 and are compatible with Hugging Face Transformers, enabling practical deployment. AI
IMPACT Offers a more parameter-efficient approach to multimodal retrieval, potentially lowering deployment costs and enabling wider adoption.
RANK_REASON Release of a new family of multimodal encoders with a novel single-tower architecture and competitive benchmark results. [lever_c_demoted from research: ic=1 ai=1.0]
- ALBERT
- BEIR-15
- ColPali
- ColQwen2.5-v0.2
- FLORES-200+
- H Company
- Hugging Face Transformers
- ModernBERT
- NeoMME
- NeoMME-Retriever
- NVIDIA H100
- NVIDIA L40S
- SigLIP2
- ViDoRe v3
- Vultron Retriever Flash
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →