← Back to all articles
arXiv cs.AIOctober 7, 2026

When Are Concept Bottleneck Model Explanations Faithful and Compact?

Excerpt

arXiv:2610.06285v1 Announce Type: cross Abstract: Concept bottleneck models (CBMs) are neural classifiers that allow to explain their decisions via high-level concepts, potentially enabling understanding, steering and debugging. However, their explanations are often derived heuristically. Building on formal explainability, we argue they should also be faithful, i.e., not misreport which concepts actually matter. We show that, for widespread CBM architectures, including recent VLM-based variants,