← Back to all articles
arXiv cs.AIOctober 7, 2026

The TIME Machine: On The Power of Motion for Efficient Perception

Excerpt

arXiv:2605.23045v4 Announce Type: replace-cross Abstract: Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and the success of self-supervised models trained on next-frame prediction. While these factors have pushed the boundaries of what video models can do, they also introduce their own set of limitations. First, scaling video models can reach prohibitive costs, with recent models needing hundreds o