arXiv cs.LGOctober 7, 2026
Dynamic Budget Allocation for LLM Evaluation under Hard Resource Constraints
Excerpt
arXiv:2610.07362v1 Announce Type: new Abstract: We evaluate large language models (LLMs) in multi-turn interactions through their time-to-event: the number of interaction steps required to produce an event of interest, such as a successful jailbreak or agentic task completion. Under limited compute, interactions may be terminated before the event occurs, so that event times are only partially observed (censored). Existing allocation methods for calibrating time-to-event bounds satisfy the budget