BatVision: Learning to See 3D Spatial Layout with Two Ears

Christensen, Jesper Haahr; Hornauer, Sascha; Yu, Stella

Computer Science > Computer Vision and Pattern Recognition

arXiv:1912.07011v1 (cs)

[Submitted on 15 Dec 2019 (this version), latest version 19 Mar 2020 (v3)]

Title:BatVision: Learning to See 3D Spatial Layout with Two Ears

Authors:Jesper Haahr Christensen, Sascha Hornauer, Stella Yu

View PDF

Abstract:Virtual camera images showing the correct layout of a space ahead can be generated by purely listening to the reflections of chirping sounds. Many species evolved sophisticated non-visual perception while artificial systems fall behind. Radar and ultrasound are used where cameras fail, but provide very limited information or require large, complex and expensive sensors. Yet sound is used effortlessly by dolphins, bats, wales and humans as a sensor modality with many advantages over vision. However, it is challenging to harness useful and detailed information for machine perception. We train a network to generate representations of the world in 2D and 3D only from sounds, sent by one speaker and captured by two microphones. Inspired by examples from nature, we emit short frequency modulated sound chirps and record returning echoes through an artificial human pinnae pair. We then learn to generate disparity-like depth maps and grayscale images from the echoes in an end-to-end fashion. With only low-cost equipment, our models show good reconstruction performance while being robust to errors and even overcoming limitations of our vision-based ground truth. Finally, we introduce a large dataset consisting of binaural sound signals synchronised in time with both RGB images and depth maps.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:1912.07011 [cs.CV]
	(or arXiv:1912.07011v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1912.07011

Submission history

From: Jesper Christensen [view email]
[v1] Sun, 15 Dec 2019 09:33:04 UTC (6,176 KB)
[v2] Fri, 13 Mar 2020 12:14:36 UTC (5,957 KB)
[v3] Thu, 19 Mar 2020 07:57:28 UTC (5,957 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:BatVision: Learning to See 3D Spatial Layout with Two Ears

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:BatVision: Learning to See 3D Spatial Layout with Two Ears

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators