论文概要
研究领域: CV 作者: Ye-Chan Kim, Seunghee Choi, SeungJu Cha 发布时间: 2026-09-06 arXiv: 2509.04290
中文摘要
弱监督密集视频字幕生成旨在仅给定每个视频的事件级字幕有序集合的情况下,对未修剪视频中的多个事件进行定位和描述。近期工作通过LLM合成辅助过渡字幕以提供额外的视觉-语言对齐,但这些字幕缺乏视觉基础,并被刚性分配到固定位置和时长的每个事件间隔中。为解决这些问题,我们提出了”先观察再合成”(SBS)框架,仅在需要时自适应地提供有视觉基础的语言指导。利用VLM,我们为事件间隔生成帧级叙述,并通过语义变化检测过渡。对于识别出的过渡,我们通过融合时间中点和语义变化点来细化事件间隔的时间掩码,并选择使视觉-语言对齐最大化的宽度。在ActivityNet Captions和YouCook2上的实验表明,该框架在字幕生成和定位方面均达到了最先进性能。
原文摘要
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only an ordered set of event-level captions per video. Recent work synthesizes auxiliary transition captions via LLM to provide additional vision-language alignment, but these captions lack visual grounding and are rigidly assigned to every inter-event gap at a fixed location and duration. To address these, we propose Seeing Before Synthesizing (SBS), a framework that adaptively provides visually grounded linguistic guidance only where warranted. Leveraging a VLM, we generate frame-level narratives for the inter-event gaps and detect transitions from the semantic variation across them. For identified transitions, we then refine inter-event temporal masks by blending the temporal…
— 自动采集于 2026-09-07
#论文 #arXiv #CV #小凯
