Researchers are developing advanced methods for optimizing GPU kernel generation using large language models (LLMs). One approach, presented at MLSys 2026, uses a harness-centered system to constrain, validate, and profile LLM-generated code, achieving significant speedups over baseline implementations. Another study introduces Atrex-Bench, a benchmark derived from production inference traces, to evaluate LLM-generated GPU kernels, revealing that current models only reach about 10% of hardware potential and often rely on fallbacks. To address this, an optimization agent was developed that successfully converts these fallbacks into production-ready kernels. Additionally, a system called ATSInfer has been created for hybrid CPU-GPU LLM inference on consumer devices, improving throughput by up to 3.29x through tensor-level scheduling and asynchronous coordination. AI
IMPACT These advancements could significantly improve the efficiency and performance of AI model inference on various hardware, potentially lowering costs and increasing accessibility.
RANK_REASON The cluster contains multiple research papers detailing novel methods for LLM-driven GPU kernel generation and optimization, including benchmarks and new agent systems.
Read on Hugging Face Daily Papers →
- arXiv
- Atrex-Bench
- Atrex-Kernel-Agent
- DagsHub
- FlyDSL
- graphics processing unit
- Hugging Face
- PyTorch
- Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices
- central processing unit
- ATSinfer
- Claude Code
- codex
- FlashInfer
- LLM
- MLSys 2026
- Nvidia Blackwell B200 GPUs
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →