Learning the abstract motion semantics of verbs from captioned videos

Stefan Mathe, Afsaneh Fazly, Sven J. Dickinson, Suzanne Stevenson

2008 (modified: 10 Nov 2022)CVPR Workshops 2008Readers: Everyone

Abstract: We propose an algorithm for learning the semantics of a (motion) verb from videos depicting the action expressed by the verb, paired with sentences describing the action participants and their roles. Acknowledging that commonalities among example videos may not exist at the level of the input features, our approximation algorithm efficiently searches the space of more abstract features for a common solution. We test our algorithm by using it to learn the semantics of a sample set of verbs; results demonstrate the usefulness of the proposed framework, while identifying directions for further improvement.

0 Replies