Get Started
Research questionHow can variance-inflated Thompson Sampling be analyzed for stochastic generalized linear bandits without tractable posterior approximations?Thompson Sampling analyses for these bandits often require posterior variance inflation, while prior guarantees may rely on tractable posterior approximations. Establishing regret bounds without that approximation assumption is therefore difficult.
Machine Learning
Research Paper
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.Posterior Tempering Explains Variance Inflation in Linear and Generalized Linear Thompson SamplingThe setting is α-TS, which uses a fractional posterior for stochastic generalized linear bandits. The supplied evidence gives theoretical regularity conditions on priors and reward distributions, upper and lower regret results, and bounds for exponential and sub-Gaussian reward families, including the α proportional to d⁻¹ regime; it does not provide deployment or empirical evidence.research paper · Sep 2, 2026
Related questions
How can sequential treatment allocation learn unknown outcome variances while preserving efficient ATE inference?How can Adam’s error be bounded for strongly convex stochastic optimization without assuming bounded iterates?How many linear samples are needed to approximate Lipschitz operators under Gaussian measures, and how fast can errors decay?How can simultaneous bounds preserve index-wise variation in Banach-valued processes with mixed-tail increments?
Home
Topics
Search
Library