Researchers have developed Code-MUE, a novel framework designed to measure the uncertainty of code Large Language Models (LLMs). This purely black-box system utilizes execution-based Semantic Interaction Graphs to assess uncertainty by analyzing runtime behavior and calculating the Von Neumann entropy of the solution space. Empirical studies involving eight state-of-the-art LLMs show that Code-MUE effectively correlates with functional correctness, outperforming traditional lexical and embedding-based methods for risk detection in software engineering workflows. AI
IMPACT This framework could improve the reliability and safety of code generation tools by better quantifying model uncertainty.
RANK_REASON The cluster describes a new research paper introducing a novel framework for evaluating code LLMs.
- alphaXiv
- arXiv
- CatalyzeX
- Code-MUE
- DagsHub
- Gotit.pub
- Hugging Face
- Large Language Models
- ScienceCast
- Semantic Interaction Graphs
- Spearman's rank correlation coefficient
- von Neumann entropy
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →