Thompson sampling
PulseAugur coverage of Thompson sampling — every cluster mentioning Thompson sampling across labs, papers, and developer communities, ranked by signal.
- used by Multi-armed bandits for adjudicating documents in pooling-based evaluation of information retrieval systems 90%
- instance of alphaXiv 70%
- instance of CatalyzeX 70%
- authored by alphaXiv 70%
- used by CORE Recommender 70%
- instance of Gotit.pub 70%
- authored by Gotit.pub 60%
- authored by ScienceCast 50%
9 day(s) with sentiment data
-
New framework boosts LLM heuristic design with Bayesian MCTS
Researchers have developed Clade-AHD, a novel framework designed to enhance the efficiency of Monte Carlo Tree Search (MCTS) in the context of Automatic Heuristic Design (AHD) for large language models. This new approac…
-
New research explores memory-augmented evolution for code optimization
Two new research papers propose novel approaches to enhance evolutionary algorithms for code optimization and automated algorithm design. EvoMem introduces a persistent memory architecture to capture and reuse successfu…
-
DocMemo framework enhances long-document understanding with dynamic memory
Researchers have introduced DocMemo, a novel memory-guided framework designed to enhance multi-modal document understanding, particularly for long documents. This system addresses limitations in static retrieval and fra…
-
LLMs enhance cold-start recommendation with Bayesian priors · 2 sources tracked
Researchers have developed a method to improve cold-start performance in comment recommendation systems by leveraging large language models (LLMs). The approach uses LLMs to extract semantic signals from comment text, c…
-
Conformal Bandits framework integrates statistical validity with reward efficiency
Researchers have introduced Conformal Bandits, a new framework that integrates Conformal Prediction into bandit problems for sequential decision-making. This approach aims to provide statistical validity and improve rew…
-
New Bayesian Optimization Method Enhances Spectroscopic Data Analysis
Researchers have developed a new method for selecting optimal wavelengths in near-infrared spectroscopy, crucial for improving the accuracy and interpretability of spectral data in tasks like sugar content estimation. T…
-
AI research uses multi-armed bandits to prune neural networks
Researchers have developed a novel method for pruning feature maps in convolutional neural networks (CNNs) to reduce computational costs and storage requirements. This approach utilizes multi-armed bandit algorithms, sp…
-
New framework classifies Thompson Sampling under model misspecification
This paper introduces a novel stochastic stability framework to analyze Thompson Sampling (TS) algorithms in dynamic decision-making scenarios where the underlying model might be misspecified. The research provides a de…
-
New algorithm PBTS tackles periodically non-stationary bandit problems
Researchers have introduced Periodic Bootstrap Thompson Sampling (PBTS), a novel algorithm designed to address bandit problems with periodic non-stationarity. Unlike traditional Thompson Sampling, which can become biase…
-
New research explores regret minimization and LLM preference optimization
This paper introduces a novel framework for regret minimization in online learning scenarios involving piecewise linear reward functions, applicable to areas like contract design and auctions. The proposed algorithm ach…
-
New Stochastic Reset Pathfinding framework introduced for graph-based learning
Researchers have introduced Stochastic Reset Pathfinding (SRP), a new episodic learning problem designed for scenarios involving unknown edge success probabilities on directed graphs. This framework is applicable to div…
-
New causal bandit methods leverage structural relationships for better decision-making
Researchers have developed new methods for causal bandits, which leverage structural relationships between variables to improve decision-making. The proposed techniques, Information-Directed Sampling (IDS) and causal va…
-
Thompson Sampling Proven 2-Competitive for Mistakes in Bayesian Bandit Models
A new paper published on arXiv details a theoretical advancement in Bayesian bandit models, proving that Thompson sampling is 2-competitive in terms of mistakes. This means Thompson sampling makes at most twice the expe…
-
New framework tackles Low Autocorrelation Binary Sequences Problem
Researchers have developed a novel hybrid search framework to tackle the complex Low Autocorrelation Binary Sequences Problem (LABS). This new method integrates Thompson sampling with parallel self-avoiding walks, allow…
-
New Joint-Thompson Sampling algorithm improves communication link adaptation
Researchers have introduced a new algorithm called Joint-Thompson Sampling (Joint-TS) for link adaptation in communication systems. This algorithm models the problem as a multi-armed bandit, where each modulation and co…
-
Bayesian Optimization needs optimal initial points, study finds
A new paper on arXiv explores the optimal number of initial points required for Bayesian Optimization (BO). The research indicates that the total cost of finding a global optimum exhibits a U-shaped relationship with th…
-
New bandit algorithms research tackles heavy tails and non-stationarity · 4 sources tracked
Three new research papers explore advancements in bandit algorithms. One paper analyzes the regret of Thompson sampling in linear-Gaussian bandits, showing a decoupling of prior-dependent and minimax regret terms. Anoth…
-
AI robot wins garment folding challenge with novel RL policy
A novel reinforcement learning approach has won first place in the online and second place in the offline rounds of the LeHome Challenge 2026, a competition focused on bimanual garment folding. The system utilizes a vis…
-
New Thompson Sampling methods tackle non-stationary and private contextual bandits
Two new research papers introduce novel approaches to Thompson sampling for contextual bandits. One paper, "Flow-Corrected Thompson Sampling for Non-Stationary Contextual Bandits," proposes a Bayesian method that reuses…
-
LLM framework generates verifiable PCB schematics without unit tests
Researchers have developed PCBSchemaGen, a novel framework designed to enable large language models (LLMs) to generate verifiable code for printed circuit board (PCB) schematic designs. Unlike typical code synthesis ben…