Joint learning of images and videos with a single Vision Transformer

Shuki Shimizu, Toru Tamaki

Published: 2023, Last Modified: 29 Sept 2023MVA 2023Readers: Everyone

Abstract: In this study, we propose a method for jointly learning of images and videos using a single model. In general, images and videos are often trained by separate models. We propose in this paper a method that takes a batch of images as input to Vision Transformer (IV-ViT), and also a set of video frames with temporal aggregation by late fusion. Experimental results on two image datasets and two action recognition datasets are presented.

0 Replies