arXiv cs.AIOctober 7, 2026
Reflections and Fragments: Securing LLMs Against Sequential Mosaic Attacks
Excerpt
arXiv:2610.05346v1 Announce Type: cross Abstract: Self-play red-teaming improves language-model safety by pitting attacker and defender roles against each other in a zero-sum game. However, real adversaries increasingly use mosaic attacks: multi-turn sequences whose individual fragments are innocuous in isolation yet assemble into a harmful payload. We develop a theory of mosaic defense that characterizes what is required to prevent such attacks without sacrificing helpfulness. We first show tha