A new benchmark called D2K-Bench has been developed to evaluate how effectively Large Language Model (LLM) agents can translate expert design guidance into efficient GPU kernels. The benchmark, comprising 26 tasks and 85 workloads, assesses guidance across high-level algorithms, dataflow design, and low-level optimizations. Results on NVIDIA B200 GPUs showed that expert guidance significantly improved correctness and performance, with frontier models like GPT-6-ASTRA, Claude Opus 4-8, and GPT 5.6 "Sol" achieving substantial speedups. AI
IMPACT This benchmark could accelerate the development of more efficient AI agents capable of complex code generation for specialized hardware.
RANK_REASON The cluster contains an academic paper introducing a new benchmark for evaluating LLM agents. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →