arXiv cs.LGOctober 7, 2026
Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads
Excerpt
arXiv:2610.07094v1 Announce Type: cross Abstract: LLM deployment is shifting from single-turn completion to agentic trajectories in which a model plans, calls tools, reads results and reasons at test time before acting. This inverts the economics of inference hardware: chat serving amortises weight reads across large batches, whereas agent trajectories are sequentially dependent, run at effective batch one, and make per-token decode latency (TPOT) the dominant term in task completion time. Using