\begin{abstract} \label{abstract}
Classic no-regret multi-armed bandit algorithms, including the Upper Confidence Bound (\textsc{UCB}), \textsc{Hedge}, and \textsc{EXP3}, are inherently unfair by design. Their unfairness stems from their very objective of playing the most rewarding arm as frequently as possible while ignoring the rest. In this paper, we consider a fair prediction problem in the stochastic setting with a guaranteed minimum rate of accrual of rewards for each arm. We study the problem in both full-information and bandit feedback settings. Combining queueing-theoretic techniques with adversarial bandits, we propose a new online policy called \BQ that achieves the target reward rates while conceding a regret and target rate violation penalty of at most $O(T^{\nicefrac{3}{4}}).$ The regret bound in the full-information setting can be further improved to $O(\sqrt{T})$ under either a monotonicity assumption or when considering time-averaged regret. The proposed policy is efficient and admits a black-box reduction from the fair prediction problem to the standard adversarial MAB problem. The analysis of the \BQ policy involves a new self-bounding inequality, which might be of independent interest. 
%
%\textcolor{red}{In our regret definition, we consider a somewhat weaker benchmark where the reward and constraint vectors are averaged over a window of constant size $w \geq 1$ instead over the entire horizon of length $T.$ }
	\end{abstract}
