arXiv cs.AIOctober 7, 2026
When Verifiable Counts Depend on Wording: Auditing Wording Robustness in Instruction Following
Excerpt
arXiv:2610.05278v1 Announce Type: cross Abstract: Verifiable instruction-following benchmarks often express each constraint through one fixed template. We test whether scores remain stable when the operational requirement is unchanged but its wording varies. We introduce WISE, a matched evaluation suite and reporting protocol instantiated on exact word count, keyword inclusion exactly once, and an inclusive 8--12 word range. Across 100 matched tasks, up to thirteen models from seven providers, a