Researchers have introduced Invariant Pretraining (InvPT), a novel method to enhance the robustness of encoder-based code representation models. These models, commonly used for tasks like clone detection and code classification, often degrade in performance when faced with semantically equivalent code written in different syntactic forms. InvPT addresses this by employing a code-only continued pretraining strategy that combines masked language modeling with multi-positive supervised contrastive learning. This approach treats all syntactic variations of the same function as positive examples, improving robustness by up to 19 percentage points on specific tasks while maintaining or enhancing standard accuracy. AI
IMPACT Enhances the reliability of code analysis tools, potentially improving developer productivity and code security.
RANK_REASON The cluster contains a research paper detailing a new method for improving code representations. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- IArxiv Recommender
- Influence Flower
- Invariant Pretraining for Robust Code Representations
- InvPT
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →