arXiv cs.AIOctober 7, 2026
When Are Concept Bottleneck Model Explanations Faithful and Compact?
Excerpt
arXiv:2610.06285v1 Announce Type: cross Abstract: Concept bottleneck models (CBMs) are neural classifiers that allow to explain their decisions via high-level concepts, potentially enabling understanding, steering and debugging. However, their explanations are often derived heuristically. Building on formal explainability, we argue they should also be faithful, i.e., not misreport which concepts actually matter. We show that, for widespread CBM architectures, including recent VLM-based variants,