Active Assessment of Prediction Services as Accuracy Surface Over Attribute Combinations

Vihari Piratla; Soumen Chakrabarti; Sunita Sarawagi

Active Assessment of Prediction Services as Accuracy Surface Over Attribute Combinations

Vihari Piratla, Soumen Chakrabarti, Sunita Sarawagi

Published: 09 Nov 2021, Last Modified: 26 May 2025NeurIPS 2021 PosterReaders: Everyone

Keywords: model evaluation, probabilistic models, active estimation, gaussian processes

Abstract: Our goal is to evaluate the accuracy of a black-box classification model, not as a single aggregate on a given test data distribution, but as a surface over a large number of combinations of attributes characterizing multiple test data distributions. Such attributed accuracy measures become important as machine learning models get deployed as a service, where the training data distribution is hidden from clients, and different clients may be interested in diverse regions of the data distribution. We present Attributed Accuracy Assay (AAA) --- a Gaussian Process (GP)-based probabilistic estimator for such an accuracy surface. Each attribute combination, called an 'arm' is associated with a Beta density from which the service's accuracy is sampled. We expect the GP to smooth the parameters of the Beta density over related arms to mitigate sparsity. We show that obvious application of GPs cannot address the challenge of heteroscedastic uncertainty over a huge attribute space that is sparsely and unevenly populated. In response, we present two enhancements: pooling sparse observations, and regularizing the scale parameter of the Beta densities. After introducing these innovations, we establish the effectiveness of AAA both in terms of its estimation accuracy and exploration efficiency, through extensive experiments and analysis.

Code Of Conduct: I certify that all co-authors of this work have read and commit to adhering to the NeurIPS Statement on Ethics, Fairness, Inclusivity, and Code of Conduct.

Supplementary Material: pdf

TL;DR: Prediction Services should go beyond scalar accuracy and instead report performance on attribute combinations characterizing a client’s utility. Read our paper for how we can estimate accuracy for combinatorially large attribute combinations.

Code: https://github.com/vihari/AAA

Community Implementations: [![CatalyzeX](/images/catalyzex_icon.svg) 4 code implementations](https://www.catalyzex.com/paper/active-assessment-of-prediction-services-as/code)

8 Replies

Loading