PulseAugur
实时 17:37:32
Deutsch(DE) Google veroeffentlicht TIPSv2 ViT-g/14: 40-Layer-Bildencoder mit zwei CLS-Tokens plus 12-Layer-Text-Encoder fuer kontrastives Zero-Shot-Image-Classification. Na

Google发布TIPS v2 ViT-g/14用于零样本图像分类

Google发布了TIPS v2 ViT-g/14,这是一款用于零样本图像分类的新模型。该模型包含一个40层图像编码器(带有两个CLS token)和一个12层文本编码器。它在224px的原始分辨率下运行,并根据Apache 2.0许可证发布,其预处理过程显著省略了ImageNet归一化。 AI

影响 该模型的架构和零样本能力有望推动图像分类任务和多模态AI研究的进步。

排序理由 前沿实验室模型发布,附带系统卡。[lever_c_降级自frontier_release: ic=1 ai=1.0]

在 Mastodon — mastodon.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Google发布TIPS v2 ViT-g/14用于零样本图像分类

报道来源 [1]

  1. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    Google发布TIPS v2 ViT-g/14:具有两个CLS token的40层图像编码器和用于对比式零样本图像分类的12层文本编码器。Na

    Google veroeffentlicht TIPSv2 ViT-g/14: 40-Layer-Bildencoder mit zwei CLS-Tokens plus 12-Layer-Text-Encoder fuer kontrastives Zero-Shot-Image-Classification. Native Aufloesung 224px, Apache-2.0-Lizenz, Preprocessing ohne ImageNet-Normalisierung. https:// huggingface.co/google/tip…