arXiv cs.AIOctober 7, 2026
Behavior Uncloning: Distilling Mode Redirection into Policy Weights without Inference-Time Steering
Excerpt
arXiv:2606.29201v2 Announce Type: replace-cross Abstract: Behavior-cloned policies often learn multiple behavior modes from demonstration datasets, including modes that are unsafe or otherwise undesired at deployment. For example, a policy trained on diverse handover demonstrations may learn to pass a knife blade-first. Standard remedies such as data curation and inference-time steering either require access to the original demonstrations for full retraining or add substantial inference-time ove