arXiv cs.AIAugust 18, 2026
Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
Excerpt
arXiv:2608.16747v1 Announce Type: cross Abstract: Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But what constitutes a "good" explanation? In this work, we evaluate explanations through the lens of counterfactual simulatability-whether the explanation is useful for predicting model behaviors on related counterfactual inputs. To this end, we introduce CHIVE (Counterfactual Hypothesis Investigation Via Edits),