Researchers have developed a new benchmark and an efficient bi-encoder for natural language code retrieval specifically for the 1C:Enterprise ecosystem. This system addresses the scarcity of open datasets and specialized models for this domain, which combines Russian syntax with specific terminology. The approach utilizes synthetic data generated by google/gemma-4-26B-A4B-it and Matryoshka Representation Learning (MRL) to achieve strong performance, outperforming baseline architectures. AI
IMPACT This research could lead to more efficient code retrieval tools for specialized software ecosystems.
RANK_REASON Academic paper detailing a new benchmark and model for a specific domain. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →