Google has released TIPS v2 ViT-g/14, a new model for zero-shot image classification. This model features a 40-layer image encoder with two CLS tokens and a 12-layer text encoder. It operates at a native resolution of 224px and is released under the Apache 2.0 license, with preprocessing that notably omits ImageNet normalization. AI
IMPACT This model's architecture and zero-shot capabilities could advance image classification tasks and multimodal AI research.
RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →