arXiv cs.AIOctober 7, 2026
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
Excerpt
arXiv:2601.11854v3 Announce Type: replace-cross Abstract: Agentic task-oriented dialogue (TOD) requires systems to track concurrent goals, dependencies, and long-horizon state. We examine goal-lifecycle recovery from fixed dialogue trajectories. ATOD contains 1,000 synthetic dialogues annotated for six advanced-TOD properties, and ATOD-Eval defines metrics for dependency-sensitive completion, memory recall, and proactivity. We implement a symbolic-vector memory evaluator for ATOD-Eval. Among fiv