← Back to all articles
arXiv cs.LGAugust 18, 2026

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning

Excerpt

arXiv:2603.12478v3 Announce Type: replace-cross Abstract: Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is highly uneven. We present Goal-Driven Data Optimization (GDO), a framework that computes six sample descriptors for each candidate and constructs optimized 1$\times$ training subsets for different goals. Under a fixed one-epoch Qwen3-VL-8B-Instruct training and evaluation recipe on 8 H20 GPUs