← Back to all articles
arXiv cs.AIOctober 2, 2026

TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety

Excerpt

arXiv:2610.01323v1 Announce Type: new Abstract: Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which suppression at supervised single-turn contexts yields a bound on multi-turn trajectory risk. The bound accounts for coverage, trans