← Back to all articles
arXiv cs.CLSeptember 11, 2026

Alignment Reduces Expressed but Not Encoded Gender Bias: A Unified Framework and Study

Excerpt

arXiv:2603.24125v3 Announce Type: replace Abstract: During training, Large Language Models (LLMs) learn social regularities that can lead to gender bias in downstream applications. Most mitigation efforts focus on reducing bias in generated outputs, typically evaluated on structured benchmarks, which raises two concerns: output-level evaluation does not reveal whether alignment modifies the model's underlying representations, and structured benchmarks may not reflect realistic usage scenarios. W