arXiv cs.AIOctober 7, 2026
The TIME Machine: On The Power of Motion for Efficient Perception
Excerpt
arXiv:2605.23045v4 Announce Type: replace-cross Abstract: Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and the success of self-supervised models trained on next-frame prediction. While these factors have pushed the boundaries of what video models can do, they also introduce their own set of limitations. First, scaling video models can reach prohibitive costs, with recent models needing hundreds o