← Back to all articles
arXiv cs.AIOctober 7, 2026

When Verifiable Counts Depend on Wording: Auditing Wording Robustness in Instruction Following

Excerpt

arXiv:2610.05278v1 Announce Type: cross Abstract: Verifiable instruction-following benchmarks often express each constraint through one fixed template. We test whether scores remain stable when the operational requirement is unchanged but its wording varies. We introduce WISE, a matched evaluation suite and reporting protocol instantiated on exact word count, keyword inclusion exactly once, and an inclusive 8--12 word range. Across 100 matched tasks, up to thirteen models from seven providers, a