Researchers have developed a new method using Bayesian optimization to efficiently identify strong single experts within large language models, a process known as gradient-free post-training. This approach, which applies Bayesian optimization within a random linear embedding of weight space and uses a Gaussian process surrogate, requires no backpropagation. Experiments on reasoning benchmarks with Qwen2.5-Instruct models demonstrated that this method achieves comparable or superior results to RandOpt with significantly fewer candidate evaluations, reducing the cost of post-training while yielding stronger deployable models. AI
IMPACT Reduces the computational cost of fine-tuning LLMs, potentially accelerating the deployment of specialized models.
RANK_REASON The cluster contains an academic paper detailing a new method for optimizing large language models. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Bayesian optimization
- Gaussian process
- Hugging Face
- Neural Thickets
- Nigel Bastian Cendra
- Qwen2.5-Instruct
- RANDOPT: Optimisation par algorithmes stochastiques
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →