arXiv cs.CLAugust 19, 2026
AVA-Encoder: Towards Agent-Native Video Representation Learning
Excerpt
arXiv:2608.12313v2 Announce Type: replace-cross Abstract: Video creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both faithful to film content and directly usable for agentic reasoning and manipulation. To address the challenge, we propose the Agentic Video Auto-Encoder (AVA-Encoder), a novel auto-encoding framework driven by age