English(EN)Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography
新的AI方法提高了3D医学影像分析的效率和准确性 · 跟踪7个来源
作者PulseAugur 编辑部·[7 个来源]·
研究人员正在开发新方法,以提高3D医学影像视觉语言模型(VLM)的效率和准确性。MedPruner引入了一个无需训练的框架,用于修剪3D医学影像中的冗余token,在保持性能的同时显著降低计算负荷。另一种方法,在“疾病中心视觉语言预训练”论文中详细介绍,利用混合CNN-ViT编码器和疾病级别对比学习,以更好地将CT扫描中的视觉特征与特定疾病对齐。GreenRFM专注于放射学基础模型的资源高效预训练,仅需最少的GPU资源,并展示了强大的可迁移性。Jolia采用概念级别对齐策略,以增强3D CT数据的对比学习,提高了分类、生成和跨中心迁移任务的性能。
AI
arXiv:2603.11625v2 Announce Type: replace-cross Abstract: While specialized Medical Vision-Language Models (VLMs) have achieved remarkable success in interpreting 2D and 3D medical modalities, their deployment for 3D volumetric data remains constrained by significant computationa…
Vision-language pre-training (VLP) holds great promise for general-purpose medical AI by leveraging radiology reports as rich textual supervision, yet existing methods struggle with 3D CT imaging due to inefficient visual backbones and coarse semantic alignment. To address these …
arXiv:2603.06467v3 Announce Type: replace Abstract: Radiology foundation models (RFMs) have largely inherited the scale-first recipe of natural-image vision--language pre-training. This recipe is difficult to deploy in 3D radiology, where training corpora are smaller, reports var…
arXiv:2606.25546v1 Announce Type: new Abstract: Vision-language pre-training (VLP) holds great promise for general-purpose medical AI by leveraging radiology reports as rich textual supervision, yet existing methods struggle with 3D CT imaging due to inefficient visual backbones …
Vision-language pre-training (VLP) holds great promise for general-purpose medical AI by leveraging radiology reports as rich textual supervision, yet existing methods struggle with 3D CT imaging due to inefficient visual backbones and coarse semantic alignment. To address these …
arXiv cs.CV
TIER_1English(EN)·Julien Khlaut, Charles Corbi\`ere, Baptiste Callard, Amaury Prat, Leo Butsanets, Antoine Saporta, Th\'eo Danielou, Leo Machado, Korentin Le Floch, Tom Boeken, Pierre Manceron, Corentin Dancette·
arXiv:2606.24570v1 Announce Type: new Abstract: Vision-language contrastive pretraining has become the dominant recipe for 3D medical foundation models, leveraging the large volumes of paired scans and reports produced in clinical practice. However, medical images usually span do…
Vision-language contrastive pretraining has become the dominant recipe for 3D medical foundation models, leveraging the large volumes of paired scans and reports produced in clinical practice. However, medical images usually span dozens of organs, and radiological reports are muc…