arXiv cs.LGOctober 1, 2026
Concept Unlearning via Cross-Attention Activation Projection for Diffusion Models
Excerpt
arXiv:2605.25765v2 Announce Type: replace-cross Abstract: Existing closed-form methods for concept unlearning in text-to-image diffusion models typically derive editing directions from fixed text embeddings, which may not fully capture how concepts are expressed across latent states, timesteps, and layers. To capture this variation, we investigate cross-attention activations collected during denoising. In controlled probing experiments using the same anchor prompts, activation-derived bases achi