← Back to all articles
arXiv cs.AIAugust 18, 2026

Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans

Excerpt

arXiv:2608.16514v1 Announce Type: cross Abstract: Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath. Whether multimodal large language models (MLLMs), given the same foveated input, search as humans do bears on their use as models of human vision and on attention-alignment scores. We compare three general-purpose MLLMs with human eye-movement scanpaths on goal-directed search (COCO-Search18), driving each model fixation by fixatio