Unifying General Audio Understanding and Spatial Awareness
Overview
MiDashengLM-Spatial, to our knowledge, is the first open-source end-to-end unified
audio-language model that supports both general audio understanding and
spatial awareness within a single architecture. It extends
MiDashengLM with
Spatial-Dasheng through a hierarchical semantic-to-spatial conditioning
(HSSC) module to separately capture and effectively integrate semantic and spatial audio
information. This design enables binaural perception and spatial audio understanding capabilities
while preserving general audio understanding capabilities of the base model.
Highlights
Unified general audio understanding and spatial awareness — the unified audio-language model that supports both general audio understanding and spatial awareness within a single architecture
Spatial-Dasheng — a spatial audio encoder that performs frame-wise detection and localization of overlapping sound events
Hierarchical Semantic-to-Spatial Conditioning (HSSC) — a conditioning mechanism that hierarchically conveys intermediate semantic representations from the semantic branch to corresponding layers of the spatial branch, allowing spatial modeling to leverage semantic context while preserving the functional separation of the two branches
Scalable spatial-scene synthesis — a data pipeline that constructs 1M spatial acoustic scenes (~13,000 hours) involving environmental sounds, speech, and music, together with rich scene-level spatial descriptions and 6M fact-verified QA pairs
Robust spatial awareness without compromising generality — extensive experiments demonstrate that Spatial-Dasheng delivers robust spatial perception and sim-to-real generalization, while MiDashengLM-Spatial acquires spatial awareness without sacrificing general audio understanding
Architecture
MiDashengLM-Spatial: Spatial-Dasheng encoder integrated with MiDashengLM via hierarchical semantic-to-spatial conditioning (HSSC).
Binaural Audio Demonstrations
🎧 Best experienced with headphones.
Citation
@article{hu2026midashenglmspatial,
title={MiDashengLM-Spatial: Unifying General Audio Understanding and Spatial Awareness},
author={Jinbo Hu and Hang Su and Lichun Fan and Heinrich Dinkel and Gang Li and Zhanchen Dai and Yiru Zhang and Chang Liu and Peng Wang and Junnan Wu and Jian Luan and Cong Zou and Heng Qu},
year={2026},
journal={arXiv preprint arxiv:2610.11156},
}