cuDNN: Efficient Primitives for Deep Learning
PulseAugur coverage of cuDNN: Efficient Primitives for Deep Learning — every cluster mentioning cuDNN: Efficient Primitives for Deep Learning across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New im2win convolution method boosts GPU performance and memory efficiency
Researchers have developed an enhanced version of the im2win convolution method, designed for greater memory efficiency and performance on GPUs. This updated method supports full precision on CUDA cores and half precisi…
-
New frameworks optimize GPU kernels for deep learning and HPC · 3 sources tracked
Three new research papers introduce advanced frameworks for optimizing GPU kernels, crucial for deep learning and high-performance computing. HIERA focuses on workload-aware planning across different implementation spac…
-
Chinese Parsers DeepDoc, MinerU Crossover in Japanese RAG Performance
A comparative analysis of two Chinese open-source document parsers, DeepDoc and MinerU, for Japanese RAG systems reveals a crossover performance based on the retrieval method used. DeepDoc demonstrated superior results …
-
New method uses world models to speed up tensor program optimization
Researchers have developed a novel approach to optimize tensor programs for machine learning systems by modeling schedule evaluation as latent dynamics. This method, inspired by world models, uses a lightweight transiti…
-
Together AI releases FlashAttention-3 and -4 for faster LLM processing
Together AI has released FlashAttention-3 and FlashAttention-4, significant upgrades to their GPU-accelerated attention mechanism for large language models. FlashAttention-3, designed for Hopper GPUs, achieves up to 75%…
-
NVIDIA open-sources cuDNN kernels after 12 years, including MoE and sparse attention
NVIDIA has open-sourced parts of its cuDNN library, a significant move after 12 years of it being closed-source. This release includes over 20 Mixture-of-Experts (MoE) kernels and NSA sparse attention kernels. The codeb…