Researchers have developed VEHBench, a new diagnostic benchmark designed to evaluate Large Language Models (LLMs) in the context of designing vibration energy harvesters (VEHs). This benchmark addresses the limitations of existing engineering benchmarks by focusing on LLM performance across different stages of the design process, rather than just the final artifact. VEHBench comprises 763 tasks grounded in literature and scored by a physical oracle, assessing LLMs in roles such as specification triage and corrupted-state recovery. Initial experiments indicate that LLM capabilities are highly dependent on the specific design stage, with no single model excelling across all tasks. AI
IMPACT Provides a stage-aware foundation for evaluating and improving LLMs in engineering design workflows.
RANK_REASON The item describes a new research benchmark for evaluating LLMs in a specific engineering domain. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →