← Back to all articles
arXiv cs.CLAugust 19, 2026

What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations

Excerpt

arXiv:2608.17719v1 Announce Type: cross Abstract: Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically rely on aggregate benchmark scores, which compress heterogeneous item-level behaviour into a single net figure. Objective: We measure what that compression conceals. Method: On three pairwise upgrades in the GPT-5.4 to GPT-5.6 Sol product sequence, we query 900 public b