arXiv cs.LGOctober 1, 2026
STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction
Excerpt
arXiv:2609.40245v1 Announce Type: cross Abstract: Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation. However, it remains unclear whether VLMs can accurately understand complex social n