Unsupervised Video Summarization Based on Deep Reinforcement Learning with Interpolation

Ui Nyoung Yoon; Myung Duk Hong; Geun-Sik Jo

doi:10.3390/s23073384

Unsupervised Video Summarization Based on Deep Reinforcement Learning with Interpolation

Sensors (Basel). 2023 Mar 23;23(7):3384. doi: 10.3390/s23073384.

Authors

Ui Nyoung Yoon¹, Myung Duk Hong¹, Geun-Sik Jo¹

Affiliation

¹ Artificial Intelligence Laboratory, Department of Electrical and Computer Engineering, Inha University, Incheon 22212, Republic of Korea.

Abstract

Individuals spend time on online video-sharing platforms searching for videos. Video summarization helps search through many videos efficiently and quickly. In this paper, we propose an unsupervised video summarization method based on deep reinforcement learning with an interpolation method. To train the video summarization network efficiently, we used the graph-level features and designed a reinforcement learning-based video summarization framework with a temporal consistency reward function and other reward functions. Our temporal consistency reward function helped to select keyframes uniformly. We present a lightweight video summarization network with transformer and CNN networks to capture the global and local contexts to efficiently predict the keyframe-level importance score of the video in a short length. The output importance score of the network was interpolated to fit the video length. Using the predicted importance score, we calculated the reward based on the reward functions, which helped select interesting keyframes efficiently and uniformly. We evaluated the proposed method on two datasets, SumMe and TVSum. The experimental results illustrate that the proposed method showed a state-of-the-art performance compared to the latest unsupervised video summarization methods, which we demonstrate and analyze experimentally.

Keywords: piecewise linear interpolation; reinforcement learning; unsupervised learning; video summarization.

Grants and funding

National Research Foundation of Korea (NRF) and INHA UNIVERSITY Research Grant