← Back to all articles
arXiv cs.AIAugust 18, 2026

BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics

Excerpt

arXiv:2608.16211v1 Announce Type: new Abstract: Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research artifacts. Medical imaging workflows are multi-stage and data-sensitive, while expert trajectories remain scarce and difficult to share. Structured benchmarks can localize failures through stage-level rubrics, but standard post-training discards these diagnostics before the next training round. We present Benchmark-as-Teacher (BaT), a recursive