← Back to all articles
arXiv cs.LGOctober 7, 2026

SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding

Excerpt

arXiv:2610.07086v1 Announce Type: new Abstract: LLM agents interact with external systems by generating structured tool calls. Given a user request, conversational context, and a catalog of tool schemas, a tool-calling model must select tools and generate their arguments, potentially producing multiple calls in a single response. Standard autoregressive decoding generates these calls token by token, incurring substantial latency for requests involving multiple calls or many argument fields. The