A semantic-alignment tuner network added ahead of the PDVC baseline, improving dense video captioning on YouCook2.
Dense video captioning has applications across autonomous driving, video surveillance, and accessibility for the visually impaired. We propose using language models to help semantically align video features and improve overall performance. Our baseline, PDVC (end-to-end dense video captioning with parallel decoding), already produces strong results relative to many state-of-the-art frameworks. PDVC is trained on ActivityNet and YouCook2; due to compute limits we used YouCook2 only.
We add semantic alignment via a tuner network placed before the video features enter PDVC. We ran ablations over several tuner architectures, and the modified PDVC outperformed the baseline on many evaluation metrics. Promising future directions remain for semantic alignment applied to larger, more comprehensive datasets.
