arXiv cs.LGOctober 2, 2026
Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models
Excerpt
arXiv:2610.01821v1 Announce Type: new Abstract: Understanding information processing in large language models (LLMs) requires dissecting the geometric organization of their internal token representations. While existing mechanistic interpretability (MI) methods seek to extract concepts, they are constrained by a strong linearity assumption challenged by evidence of non-linear feature manifolds. We move beyond linear concepts by adapting Non-Linear Multi-Dimensional Concept Discovery (NLMCD) from