Cross-model consensus of explanations and beyond for image classification models: an empirical study

DOI10.1007/S10994-023-06312-1arXiv2109.00707OpenAlexW3198826142MaRDI QIDQ6103575FDOQ6103575

Dejing Dou, Xuhong Li, Siyu Huang, Haoyi Xiong, Shilei Ji

Publication date: 27 June 2023

Published in: Machine Learning (Search for Journal in Brave)

Abstract: Existing interpretation algorithms have found that, even deep models make the same and right predictions on the same image, they might rely on different sets of input features for classification. However, among these sets of features, some common features might be used by the majority of models. In this paper, we are wondering what are the common features used by various models for classification and whether the models with better performance may favor those common features. For this purpose, our works uses an interpretation algorithm to attribute the importance of features (e.g., pixels or superpixels) as explanations, and proposes the cross-model consensus of explanations to capture the common features. Specifically, we first prepare a set of deep models as a committee, then deduce the explanation for every model, and obtain the consensus of explanations across the entire committee through voting. With the cross-model consensus of explanations, we conduct extensive experiments using 80+ models on 5 datasets/tasks. We find three interesting phenomena as follows: (1) the consensus obtained from image classification models is aligned with the ground truth of semantic segmentation; (2) we measure the similarity of the explanation result of each model in the committee to the consensus (namely consensus score), and find positive correlations between the consensus score and model performance; and (3) the consensus score coincidentally correlates to the interpretability.

Full work available at URL: https://arxiv.org/abs/2109.00707

zbMATH Keywords

interpretability semantic segmentation and visualization explanations of deep neural networks

Mathematics Subject Classification ID

Learning and adaptive systems in artificial intelligence (68T05)

Cites Work

Cited In (1)

Distilling ensemble of explanations for weakly-supervised pre-training of image segmentation models

This page was built for publication: Cross-model consensus of explanations and beyond for image classification models: an empirical study

Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6103575)