Researchers from Kaiming He's team at MIT have developed NAT-ARC, a novel approach to solving the ARC (Abstraction and Reasoning Corpus) challenge using a purely visual method. Unlike previous methods that relied on large language models, NAT-ARC leverages a vision encoder pre-trained on ImageNet, specifically using the Masked Autoencoder (MAE) technique. This pre-training allows the model to learn general visual understanding from natural images like cats and dogs, which then transfers effectively to abstract reasoning tasks on colored grids. NAT-ARC's best single model achieved a 63.4% pass@2 score on ARC-1, with an ensemble reaching 70.2%, demonstrating that visual-only solutions can now rival LLM-based approaches in this difficult benchmark. AI
IMPACT Demonstrates the potential of pre-trained visual models for abstract reasoning, challenging the dominance of LLMs in benchmarks like ARC.
RANK_REASON Paper introducing a new method for a challenging AI benchmark. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →