arXiv cs.CLSeptember 21, 2026
Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction
Excerpt
arXiv:2609.21392v1 Announce Type: new Abstract: Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental question underexplored: can a model correctly infer the user's underlying demand from complex multimodal interaction? Real-world user