Researchers have conducted a large-scale study on scaling pixel-wise Earth-observation foundation models, involving 395 training runs on 1,024 NVIDIA GH200 superchips. The study found that pretraining loss is a poor predictor of downstream performance, suggesting that focusing compute on larger encoders and more data, rather than larger projectors, is more effective. The team trained a family of models, including the 21-million-parameter TESSERA v2-1B-M, which outperforms larger open and proprietary models and offers efficient serving through Matryoshka representations. AI
IMPACT Provides a data-driven recipe for efficiently scaling Earth-observation foundation models, potentially accelerating deployment and improving performance.
RANK_REASON The cluster describes a research paper detailing a large-scale empirical study on scaling foundation models for Earth observation.
Read on Hugging Face Daily Papers →
- Barlow Twins
- Earth-observation
- GH200 superchips
- TESSERA v2
- Earth Foundation Models
- GeoTESSERA
- Matryoshka
- NVIDIA GH200
- Sentinel-1
- Sentinel-2
- TESSERA v2-1B-M
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →