Visually-Guided Spatial Audio Generation for 360° In-the-Wild Speech Scenes

Accepted at Interspeech 2026

Abstract

Spatial audio is a key component of immersive 360◦ media, yet high-quality spatial capture remains limited in real-world speechdominant scenes. We study visually guided First-Order Ambisonics (FOA) speech spatialization in the wild: given aligned 360◦ video and an omnidirectional audio track, we recover the missing directional FOA components. To support this task, we introduce YT-SPEECH, a speech-oriented 360◦ video-FOA dataset curated from YouTube. We propose a two-stage Localizer-Renderer framework, where an audio-visual segmentation backbone provides frame-wise spatial heatmaps and a conditional complexdomain U-Net reconstructs directional FOA signals from the omnidirectional channel. A confidence-based gating strategy stabilizes conditioning under ambiguous acoustic conditions. Experiments show improved reconstruction fidelity, spatial accuracy, and perceptual speech quality relative to ablated variants and prior approaches.

Overview

Framework overview

Overview of the proposed framework.

Guidelines for Listening Demo Videos

  • Please use your headphones 🎧 for the best audio experience.
  • Please feel free to drag the screen to experience the 360-degree audio 🎥
  • Please use Play to start/pause the spatial audio; Rewind to restart.

YT-SPEECH Examples

NOTE: Input shows the model input, a 360° video with mono audio (W); Ground-Truth and Prediction are full FOA.