Researchers have introduced STEP-KTODER, a novel framework designed to enhance code generation models through function-level process supervision. This method defines 'steps' as module-level functions within decomposed programs and utilizes automatically generated unit tests to assign binary correctness labels. By combining function-level supervision with outcome-level feedback, STEP-KTODER aims to improve upon existing methods like Direct Preference Optimization (DPO) and outcome-only KTO. Evaluations on benchmarks such as HumanEval(+) and MBPP(+) demonstrate STEP-KTODER's effectiveness, highlighting the critical role of execution-based labels over LLM-as-a-judge annotations, which were found to degrade performance. AI
IMPACT This research could lead to more robust and accurate AI code generation models by improving how they learn from intermediate execution feedback.
RANK_REASON Academic paper detailing a new method for code generation optimization. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- BigCodeBench
- CatalyzeX Code Finder for Papers
- CORE Recommender
- DagsHub
- Direct Preference Optimization
- Gotit.pub
- Hugging Face
- Influence Flower
- LiveCodeBench
- ScienceCast
- STEP-KTODER
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →