Sound-Indicated Visual Object Detection for Robotic ExplorationDownload PDFOpen Website

2019 (modified: 28 Jan 2024)ICRA 2019Readers: Everyone
Abstract: Robots are usually equipped with microphones and cameras to perceive and understand the physical world. Though visual object detection technology has achieved great success, the detection in other modalities remains unsolved. In this paper, we establish a novel robotic sound-indicated visual object detection framework, and develop a two-stream weakly-supervised deep learning architecture to connect the visual and audio modalities for localizing the sounding object. A dataset is constructed from the AudioSet to validate the proposed method and some promising applications are demonstrated on robotic platforms.
0 Replies

Loading