arXiv Analytics

Sign in

arXiv:2006.11693 [cs.CV]AbstractReferencesReviewsResources

Dense-Captioning Events in Videos: SYSU Submission to ActivityNet Challenge 2020

Teng Wang, Huicheng Zheng, Mingjing Yu

Published 2020-06-21Version 1

This technical report presents a brief description of our submission to the dense video captioning task of ActivityNet Challenge 2020. Our approach follows a two-stage pipeline: first, we extract a set of temporal event proposals; then we propose a multi-event captioning model to capture the event-level temporal relationships and effectively fuse the multi-modal information. Our approach achieves a 9.28 METEOR score on the test set.

Related articles: Most relevant | Search more
arXiv:1907.12223 [cs.CV] (Published 2019-07-29)
Multi-Granularity Fusion Network for Proposal and Activity Localization: Submission to ActivityNet Challenge 2019 Task 1 and Task 2
arXiv:1705.00754 [cs.CV] (Published 2017-05-02)
Dense-Captioning Events in Videos
arXiv:2206.10861 [cs.CV] (Published 2022-06-22)
UniCon+: ICTCAS-UCAS Submission to the AVA-ActiveSpeaker Task at ActivityNet Challenge 2022