A new system called Tessera has been developed to improve the performance and cost-efficiency of running large AI models on heterogeneous GPU clusters. Unlike previous methods that operated at a coarse granularity, Tessera disaggregates workloads at the kernel level, recognizing that different kernels have varying resource demands. This approach allows for more precise alignment of computation with hardware capabilities, leading to significant improvements in serving throughput and cost efficiency. Tessera also demonstrates generalization to new model architectures and can even outperform homogeneous GPU setups at a lower cost. AI
IMPACT Optimizes AI model serving on diverse hardware, potentially lowering inference costs and increasing throughput.
RANK_REASON The cluster contains a research paper detailing a new system for optimizing AI workloads on heterogeneous GPUs.
- arXiv
- graphics processing unit
- Parallel Thread Execution
- tessera
- Tiancheng Hu
- CoLab
- CUDA
- Cutileiro
- Flash Attention
- Gelu
- NVIDIA
- PyTorch
- TileGym
- Triton
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →