arXiv cs.AIOctober 2, 2026
MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs
Excerpt
arXiv:2610.01434v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units, leaving finer-grained redundancy underexplored. We find that redundancy varies both across modality-interaction paths within the