← Back to all articles
arXiv cs.AIAugust 18, 2026

MLLM-Guided Semantic Correction for Text-to-Video Generation

Excerpt

arXiv:2608.16513v1 Announce Type: cross Abstract: Recent advances in diffusion models and Transformer architectures have led to significant progress in text-to-video generation. However, these models often suffer from semantic errors such as missing objects, incorrect attributes, or mismatched actions. Although some semantic correction methods perform optimization before sampling or refinement after sampling, how to detect and correct semantic deviations during the video generation process remai