arXiv cs.CLSeptember 24, 2026
Small Cues, Big Consequences: Learning Pivotal Cues for Multimodal Meme Classification
Excerpt
arXiv:2609.26907v1 Announce Type: cross Abstract: Memes often derive their harmful, hateful, or sarcastic meaning from small but decisive visual, textual, or cross-modal cues. Existing multimodal classifiers can miss such evidence when relying mainly on global image-text representations. We introduce MemeCF, a cue-focused benchmark of 9,895 memes across harm, hate, and sarcasm, with annotations identifying the modality and rationale of the pivotal evidence. We also propose MemePIVOT, a local-glo