← Back to all articles
arXiv cs.CLSeptember 22, 2026

BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents

Excerpt

arXiv:2609.23490v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly execute multi-step workflows through tool use and interaction with users and environments. However, current agent evaluations are largely English-centric, limiting our understanding of agent capabilities in multilingual settings. We introduce BabelFlow, a benchmark-general agentic workflow that adapts existing agent benchmarks to new languages by analyzing runtime dependencies, coordinating structure-p