← Back to all articles
arXiv cs.AIOctober 2, 2026

Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory

Excerpt

arXiv:2601.22984v3 Announce Type: replace Abstract: Diagnosing failure patterns in Deep Research Agents (DRAs) remains a critical challenge. Existing benchmarks predominantly rely on end-to-end evaluation, obscuring intermediate hallucinations that accumulate throughout the research trajectory. To bridge this gap, we propose a shift from outcome-based to process-aware evaluation by auditing hallucinations in the full plan-search-summarize trajectory. We introduce the PING Taxonomy, which categor