← Back to all articles
arXiv cs.AIOctober 7, 2026

Bidirectional Preference Synthesis: Learning Prompt-Conditioned Preferences from Boundary Failures

Excerpt

arXiv:2610.04328v1 Announce Type: new Abstract: Correction-based offline preference pipelines commonly treat model failures only as rejected responses under the original prompt. This supervision is incomplete for boundary failures: responses that violate the given instruction yet coherently satisfy a nearby intent or constraint setting. We introduce Bidirectional Preference Synthesis (BPS), a data-construction method for standard Direct Preference Optimization (DPO) that makes this missing promp