Florence 2
PulseAugur coverage of Florence 2 — every cluster mentioning Florence 2 across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
CapMap-MS-TTA system ranks 3rd in LSVOS Challenge multimodal track
A research paper details CapMap-MS-TTA, a system that achieved third place in the MUMU track of the 8th LSVOS Challenge at ECCV 2026. This track required a single multimodal model to perform image tagging, open-vocabula…
-
New FZ-VLM framework aids radiologists in pulmonary nodule characterization
Researchers have developed FZ-VLM, a novel two-stage vision-language model framework designed to assist radiologists in characterizing pulmonary nodules from CT scans. The first stage, utilizing a fine-tuned Florence-2 …
-
AI assistant Click2Poly streamlines manual vector mapping in QGIS
Researchers have developed Click2Poly, a new AI assistant that integrates with QGIS to streamline the manual process of vector mapping buildings and walls. This tool, built upon the Florence-2 Vision Language Model (VLM…
-
User study finds improved robot interaction system perceptible to users
A new study published on arXiv explores the perceptual differences users experience when interacting with a multimodal human-robot system. The research compared a baseline system using Whisper, Florence-2, and Llama 3.1…
-
Krea2 Trainer simplifies local AI model training for users
A new, user-friendly interface called Krea2 Trainer has been developed to simplify the local training of Krea2 models, a process that was previously complex. This tool, created by a non-developer, aims to make Krea2 tra…
-
Crucible launches as open-source local dataset manager for diffusion models
Crucible is a new, open-source, local application designed for managing datasets used in diffusion models. It runs entirely on user hardware, avoiding cloud dependencies and subscriptions. The tool offers features like …
-
Fashion Florence model extracts structured clothing attributes
Researchers have developed Fashion Florence, a vision-language model based on Florence-2, specifically fine-tuned for extracting structured fashion attributes from images. This model can generate a JSON object detailing…
-
Alibaba launches Qwen3.7-Plus multimodal agent model
Alibaba's Qwen team has released Qwen3.7-Plus, a new multimodal agent model designed to integrate vision and language capabilities for versatile agentic tasks. This release is part of a broader trend highlighted by Hugg…