← Back to all articles
arXiv cs.CLSeptember 24, 2026

VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark

Excerpt

arXiv:2508.13680v5 Announce Type: replace Abstract: We introduce VMMU, a Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark designed to evaluate how vision-language models (VLMs) interpret and reason over visual and textual information beyond English. VMMU consists of 2.5k multimodal questions across 7 tasks, covering a diverse range of problem contexts, including STEM problem solving, data interpretation, rule-governed visual reasoning, and abstract visual reasoning. All ques