PulseAugur
EN
LIVE 17:36:52
Deutsch(DE) Google veroeffentlicht TIPSv2 ViT-g/14: 40-Layer-Bildencoder mit zwei CLS-Tokens plus 12-Layer-Text-Encoder fuer kontrastives Zero-Shot-Image-Classification. Na

Google releases TIPS v2 ViT-g/14 for zero-shot image classification

Google has released TIPS v2 ViT-g/14, a new model for zero-shot image classification. This model features a 40-layer image encoder with two CLS tokens and a 12-layer text encoder. It operates at a native resolution of 224px and is released under the Apache 2.0 license, with preprocessing that notably omits ImageNet normalization. AI

IMPACT This model's architecture and zero-shot capabilities could advance image classification tasks and multimodal AI research.

RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=1 ai=1.0]

Read on Mastodon — mastodon.social →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Google releases TIPS v2 ViT-g/14 for zero-shot image classification

COVERAGE [1]

  1. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    Google releases TIPS v2 ViT-g/14: 40-layer image encoder with two CLS tokens plus 12-layer text encoder for contrastive zero-shot image classification. Na

    Google veroeffentlicht TIPSv2 ViT-g/14: 40-Layer-Bildencoder mit zwei CLS-Tokens plus 12-Layer-Text-Encoder fuer kontrastives Zero-Shot-Image-Classification. Native Aufloesung 224px, Apache-2.0-Lizenz, Preprocessing ohne ImageNet-Normalisierung. https:// huggingface.co/google/tip…