← Back to all articles
arXiv cs.CLSeptember 11, 2026

DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports

Excerpt

arXiv:2601.08536v3 Announce Type: replace Abstract: Deep Research Agents (DRA) aim to help users search the web, synthesize information, and deliver comprehensive investigative reports. Prior benchmarks often either under-evaluate a system's ability to produce meaningful insights and high-quality writing, or adopt coarse or LLM-defined criteria that are hard to verify and can diverge from human expert judgment. To address these issues, we introduce Deep Research Bench II, a new benchmark for eva