cuBLASLt
PulseAugur coverage of cuBLASLt — every cluster mentioning cuBLASLt across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New LLM inference techniques boost GPU utilization and efficiency
Researchers have developed a new method to dissect GPU utilization for LLM inference, moving beyond a single percentage to provide eight detailed views derived from Nsight Compute reports. This approach maps utilization…
-
AI system CUDA-L2 surpasses NVIDIA's cuBLAS for matrix multiplication
Researchers have developed CUDA-L2, a system that leverages large language models and reinforcement learning to automatically optimize matrix multiplication CUDA kernels. This system significantly outperforms existing b…
-
CUDA/C++ inference engine built for NVIDIA's DVLT 3D model
A new inference engine called dvlt.cu has been developed from scratch using CUDA/C++ for NVIDIA's DVLT 3D transformer model. This standalone 5MB binary has minimal dependencies, relying only on cuBLASLt and the header-o…