Researchers have developed Kernel Forge, an open-source agentic harness that uses large language models to automatically generate and optimize CUDA kernels for PyTorch models. This tool aims to reduce the need for expert engineers to manually write low-level GPU code. Kernel Forge supports various workloads, including vision, diffusion, and LLM models, and employs Monte Carlo Tree Search for optimization. It has demonstrated significant speedups, outperforming PyTorch's eager mode on several kernels, with notable improvements on models like ResNet-50 and Gemma 4-E2B. AI
IMPACT Accelerates GPU kernel optimization for AI models, potentially reducing inference latency and costs across various workloads.
RANK_REASON The cluster describes a new open-source tool and research paper detailing an agentic harness for optimizing CUDA kernels using LLMs.
Read on Mastodon — sigmoid.social →
- CUDA
- DGX Spark
- Gemma 4-E2B
- Kernel Forge
- Monte Carlo tree search
- NVIDIA
- NVIDIA GB10 Grace Blackwell Superchip
- PyTorch
- Qwen 3.5-35B-A3B
- ResNet-50
- Stable Diffusion 3.5 Medium
- University of Michigan
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →