This paper introduces a novel framework for regret minimization in online learning scenarios involving piecewise linear reward functions, applicable to areas like contract design and auctions. The proposed algorithm achieves a tight regret bound of $\widetilde{O}(\sqrt{nT})$ under a monotonicity assumption, addressing open problems in learning optimal linear contracts and setting prices in posted-price auctions. Additionally, a separate paper explores optimizing regret by developing derivative theory for the covariance regret functional, suggesting applications in portfolio tilting and LLM-based allocation strategies. Another study focuses on preference optimization for LLMs, proposing a regularization technique to prevent over-optimization and improve generation quality, demonstrated with gains on Llama-3.1-8B-Instruct. AI
IMPACT These papers advance theoretical understanding in optimization and LLM alignment, potentially leading to more efficient and controllable AI systems.
RANK_REASON Multiple academic papers published on arXiv discussing theoretical advancements in optimization and LLM alignment.
- AlpacaEval2
- arXiv
- CBCP19
- Cesa-Bianchi et al.
- Direct Preference Optimization
- Francesco Bacchiocchi
- Llama 3.1 8B-Instruct
- Zhu+23
- Hugging Face
- LLM-based allocation strategies
- piecewise linear rewards
- portfolio tilting
- Thompson sampling
AI-generated summary · Google Gemini · from 5 sources. How we write summaries →