Computational Modeling of Soundscape and Visual Imagery in Beijing Old City Communities Based on Multi-modal Data
DOI:
https://doi.org/10.67541/jdf2603Keywords:
Soundscape; Visual imagery; Multimodal learning; Cross-modal correlation; Beijing old city communityAbstract
As an important carrier of the spirit of place, soundscape lacks systematic quantitative evaluation tools. The fundamental bottleneck lies in the lack of an accurate cross-modal alignment calculation framework between soundscape temporal sequence signals and visual image spatial semantics. In this paper, we propose a joint embedding model driven by cross-modal contrastive learning to realize the dynamic coupling of acoustic physical-perceptual features and visual global-local semantic features. We introduce a cross-modal attention forgetting gate to suppress modality-specific noise, design a hierarchical contrastive loss function to embed spatial structure with both instance-level and category-level constraints, and construct a spatio-temporal adaptive weight regulator to realize spatially differentiated modulation of the modality fusion coefficient. Based on 8,432 groups of measured multimodal data (including day-night differences) from four typical community types in Beijing’s old city, our experiments show that the proposed model achieves a cross-modal retrieval mAP of 76.82%, a perception score prediction RMSE as low as 0.318, and an intraclass correlation coefficient of 0.824 with manual scores, all of which are significantly better than those of the comparison methods (p < 0.001). Ablation experiments verify the independent contributions of the three innovations. Feature response analysis reveals that the partial correlation coefficient between soundscape roughness and visual building density exceeds 0.78, indicating a deep sensory coordination mechanism between auditory texture density and visual spatial envelopment perception. The proposed model provides an effective computational tool for quantitative diagnosis of multi-sensory quality in old city communities.
References
[1] Wang, S., Zhang, J., Wang, F., & Dong, Y. (2023). How to achieve a balance between functional improvement and heritage conservation? A case study on the renewal of old Beijing city. Sustainable Cities and Society, 98, 104790. https://doi.org/10.1016/j.scs.2023.104790
[2] Zhang, R., Martí Casanovas, M., Bosch González, M., & Sun, S. (2024). Revitalizing heritage: The role of urban morphology in creating public value in China’s historic districts. Land, 13(11), 1919. https://doi.org/10.3390/land13111919
[3] Amir, S., Sadoway, D., & Dommaraju, P. (2023). Taming the noise: Soundscape and livability in a technocratic city-state. East Asian Science, Technology and Society: An International Journal, 17(1), 88-104. https://doi.org/10.1080/18752160.2021.1936749
[4] Gil-Sayas, S., Di Pierro, G., Tansini, A., Serra, S., Currò, D., Broatch, A., & Fontaras, G. (2024). Energy consumption of mobile air-conditioning systems in electrified vehicles under different ambient temperatures. International Journal of Engine Research, 25(2), 293-304. https://dx.doi.org/10.1177/14680874231171303
[5] Bahgat, G., Al-Makhlasawy, R. M., Khairy, M., Nour, M., & Abdelfattah, A. (2026). Energy saving for air conditioning devices based on the amalgamation of intelligent models and occupancy detection. Journal of Electrical Systems and Information Technology, 13(1), 78. https://doi.org/10.1186/s43067-026-00379-1
[6] Zhang, D., Ni, J., & Shi, X. (2024). Study of low-temperature energy consumption optimization of battery electric vehicle air conditioning systems considering blower efficiency. Processes, 12(7), 1495. https://doi.org/10.3390/pr12071495
[7] Muhammad, K., Hussain, T., Ullah, H., Del Ser, J., Rezaei, M., Kumar, N., ... & De Albuquerque, V. H. C. (2022). Vision-based semantic segmentation in scene understanding for autonomous driving: Recent achievements, challenges, and outlooks. IEEE Transactions on Intelligent Transportation Systems, 23(12), 22694-22715. https://doi.org/10.1109/TITS.2022.3207665
[8] Emek Soylu, B., Guzel, M. S., Bostanci, G. E., Ekinci, F., Asuroglu, T., & Acici, K. (2023). Deep-learning-based approaches for semantic segmentation of natural scene images: A review. Electronics, 12(12), 2730. https://doi.org/10.3390/electronics12122730
[9] Liang, Y., Mitchell, A., Kang, J., & Aletta, F. (2026). A Review of Soundscape Datasets: Challenges and Prospects for Multimodal Research. IEEE Transactions on Affective Computing. https://doi.org/10.1109/TAFFC.2026.3659084
[10] Chen, M., Lin, Z., Song, X., Luo, Y., Duan, X., Li, S., & Xie, C. (2026). Theoretical Framework, Technical Evolution, and Future Prospects of Cross-Modal Mapping and Controllable Image Generation Under Multi-Source Heterogeneous Collaboration. Sensors (Basel, Switzerland), 26(10), 2972. https://doi.org/10.3390/s26102972
[11] Kothinti, S. R., & Elhilali, M. (2023). Are acoustics enough? Semantic effects on auditory salience in natural scenes. Frontiers in Psychology, 14, 1276237. https://doi.org/10.3389/fpsyg.2023.1276237
[12] Sun, J., Deng, L., Afouras, T., Owens, A., & Davis, A. (2023). Eventfulness for interactive video alignment. ACM Transactions on Graphics (TOG), 42(4), 1-10. https://doi.org/10.1145/3592118
[13] Zhuang, Y., Kang, Y., Fei, T., Bian, M., & Du, Y. (2024). From hearing to seeing: Linking auditory and visual place perceptions with soundscape-to-image generative artificial intelligence. Computers, Environment and Urban Systems, 110, 102122. https://doi.org/10.1016/j.compenvurbsys.2024.102122
[14] Chen, P., Huang, X., Fei, T., & Wang, S. (2026). Cross‐Modal Urban Sensing: Evaluating Sound–Vision Alignment Across Street‐Level and Aerial Imagery. Transactions in GIS, 30(2), e70246. https://doi.org/10.1111/tgis.70246
[15] Zhang, Y., Wu, M., & Cai, X. (2025). A dynamic cross-modal learning framework for joint text-to-audio grounding and acoustic scene classification in smart city environments. Digital Signal Processing, 167, 105444. https://doi.org/10.1016/j.dsp.2025.105444
[16] Wang, T., Li, F., Zhu, L., Li, J., Zhang, Z., & Shen, H. T. (2025). Cross-modal retrieval: a systematic review of methods and future directions. Proceedings of the IEEE, 112(11), 1716-1754. https://doi.org/10.48550/arXiv.2308.14263
[17] Yang, Y., Wang, D., Hu, Y., & Liang, L. (2025, June). Research on Time Synchronization Technology of Non-safety DCS System Based on IEEE1588v2 Timing Protocol. In International Conference on Nuclear Engineering (pp. 163-176). Singapore: Springer Nature Singapore. https://doi.org/10.1007/978-981-95-2921-6_13
[18] Patoli, A. A., & Fortino, G. (2025). FPGA-based system implementation of IEEE 1588 precision time protocol: A review. IEEE Sensors Journal, 25(11), 18624-18642. https://doi.org/10.1109/JSEN.2025.3557277
[19] Iturbe-Martin, Z., Martín-Garín, A., & Casado-Rezola, A. (2026). Soundscape-Informed Urban Planning and Architecture in Historic Centers: A Multi-Layer Method for Soundscape Characterization Applied to Bilbao Old Town. Applied Sciences, 16(8), 3630. https://doi.org/10.3390/app16083630
[20] Versümer, S., Steffens, J., & Weinzierl, S. (2023). Day-to-day loudness assessments of indoor soundscapes: Exploring the impact of loudness indicators, person, and situation. The Journal of the Acoustical Society of America, 153(5), 2956-2956. https://doi.org/10.1121/10.0019413
[21] Chen, Y., Lin, Y., Xu, R., & Vela, P. A. (2023, October). Wdiscood: Out-of-distribution detection via whitened linear discriminant analysis. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 5275-5284). IEEE. https://doi.org/10.48550/arXiv.2303.07543
[22] Gao, S., Yang, K., Shi, H., Wang, K., & Bai, J. (2022). Review on panoramic imaging and its applications in scene understanding. IEEE Transactions on Instrumentation and Measurement, 71, 1-34. https://doi.org/10.1109/TIM.2022.3216675
[23] Taha, K. (2026). Generative AI for multimodal content: a survey with empirical and experimental evaluations. Artificial Intelligence Review, 59(6), 140. https://doi.org/10.1007/s10462-026-11525-6
[24] Yan, T., Zhao, S., Hu, M., Wang, M., Zhang, X., Luo, Z., & Wang, M. (2024). HCL: A hierarchical contrastive learning framework for zero-shot relation extraction. IEEE Transactions on Neural Networks and Learning Systems, 36(3), 5694-5705. https://doi.org/10.1109/TNNLS.2024.3379527
[25] Zhong, B., Wang, P., & Wang, X. (2024, August). Ts-HCL: hierarchical layer-wise contrastive learning for unsupervised domain adaptation on time-series. In Asia-Pacific Web (APWeb) and Web-Age Information Management (WAIM) Joint International Conference on Web and Big Data (pp. 31-45). Singapore: Springer Nature Singapore. https://doi.org/10.1007/978-981-97-7238-4_3
[26] Xie, H., He, Y., Wu, X., & Lu, Y. (2022). Interplay between auditory and visual environments in historic districts: A big data approach based on social media. Environment and Planning B: Urban Analytics and City Science, 49(4), 1245-1265. https://doi.org/10.1177/23998083211059838
[27] Yao, X. W., Zheng, J. M., Li, Q., Yang, K. H., Yao, Z. H., & Shang, Y. H. (2025, September). MSTDFN: Multi-modal Spatio-Temporal Dynamic Fusion Network for Traffic Flow Prediction. In China Conference on Wireless Sensor Networks (pp. 189-202). Singapore: Springer Nature Singapore. https://doi.org/10.1007/978-981-95-9615-7_12
[28] Zhao, T., Chen, G., Suraphee, S., Phoophiwfa, T., & Busababodhin, P. (2025). A hybrid TCN-XGBoost model for agricultural product market price forecasting. PLoS One, 20(5), e0322496. https://doi.org/10.1371/journal.pone.0322496
[29] Zhao, T., Chen, G., Pang, C., & Busababodhin, P. (2025). Application and performance optimization of SLHS-TCN-XGBoost model in power demand forecasting. Comput. Model. Eng. Sci, 143(3), 2883-2917. https://doi.org/10.32604/cmes.2025.066442
[30] Zhao, T., Chen, G., Pang, C., Seenoi, P., Papukdee, N., & Busababodhin, P. (2025). Time-lapse earthquake difference prediction based on physics-informed long short-term memory coupled with interpretability boosting. Journal of Seismic Exploration, 34(3), 25. https://doi.org/10.36922/JSE025310049
[31] Zhao, T., Chen, G., Pang, C., Li, L., & Busababodhin, P. (2026). Forecasting global agricultural trade imbalances using a hybrid deep learning and gradient boosting framework. Discover Computing, 29(1), 443. https://doi.org/10.1007/s10791-026-10367-8
Published
Data Availability Statement
The data that support the findings of this study are available upon request from the corresponding authors, L.M.
Issue
Section
License
Copyright (c) 2026 Ziyan Shen, Xuanyi Qi, Mengyu Liu (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.