arXiv cs.LGOctober 2, 2026
Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard
Excerpt
arXiv:2610.00651v1 Announce Type: cross Abstract: Agent evaluations are increasingly used to compare LLMs and inform deployment decisions, yet ranks can reflect not only the model but also the effects of the evaluation conditions such as the scaffolds or tasks. This makes reliability claim-dependent: an evaluation that reliably ranks deployed systems may not reliably rank underlying models. We ask which conclusions current agent evaluations reliably support and what additional evaluation would i