← Back to all articles
arXiv cs.LGOctober 7, 2026

DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models

Excerpt

arXiv:2610.08341v1 Announce Type: cross Abstract: Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic degradation due to their task-agnostic design or unreliable attention estimates. Based on our empirical analysis, we have found that this issue arises because salient tokens in shallow layers persistently suppress emerging