arXiv cs.CLSeptember 22, 2026
BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents
Excerpt
arXiv:2609.23490v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly execute multi-step workflows through tool use and interaction with users and environments. However, current agent evaluations are largely English-centric, limiting our understanding of agent capabilities in multilingual settings. We introduce BabelFlow, a benchmark-general agentic workflow that adapts existing agent benchmarks to new languages by analyzing runtime dependencies, coordinating structure-p