← Back to all articles
arXiv cs.CLAugust 19, 2026

TSQueryBench: LLM-as-a-Judge for Time Series Explanations

Excerpt

arXiv:2604.02118v2 Announce Type: replace-cross Abstract: Natural language explanations of time series data are increasingly produced by foundation models in high stakes domains, making factual correctness critical. Evaluating such explanations differs fundamentally from standard natural language generation: correctness requires verifying numerical claims against structured data rather than similarity to reference text. While LLM as a Judge has emerged as a scalable paradigm for text evaluation,