← Back to all articles
arXiv cs.LGOctober 7, 2026

Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators

Excerpt

arXiv:2509.03647v3 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other models. This bias undermines fairness and reliability in evaluation pipelines, particularly for tasks like preference tuning and model routing. We investigate whether lightweight steering vectors can mitigate this problem at inference time without retraining. We int