A new framework called STAGE has been developed to synthesize high-fidelity execution graphs for large language models (LLMs) and Mixture-of-Experts (MoEs). This framework aims to optimize distributed AI workloads by modeling various parallelization strategies, enabling exploration of different model architectures and system configurations without requiring direct access to large-scale infrastructure. STAGE has demonstrated its scalability by generating traces for over 128,000 GPUs, preserving tensor-level accuracy in compute, memory, and communication. AI
IMPACT Enables scalable exploration of distributed LLM training and inference configurations without requiring direct access to large-scale AI infrastructure.
RANK_REASON The cluster describes a new framework for synthesizing LLM execution graphs, detailed in an arXiv paper.
- A100 80GB
- AMD Strix Halo
- H100 SXM5
- H200 GPUs
- Hugging Face Jobs
- Llama 3.1 70B
- RDMA
- RoCE v2
- Tensor Parallelism
- vLLM
- AMD
- arXiv
- dev.to
- Hugging Face
- LLMs
- MoEs
- OpenAI
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →