Computational Modeling of Soundscape and Visual Imagery in Beijing Old City Communities Based on Multi-modal Data

Ziyan Shen1 , Xuanyi Qi2 , Mengyu Liu3*
1 College of Mechanical and Electrical Engineering, Beijing University of Chemical Technology, Chaoyang District, Beijing, China
2 Department of Art and Design, Beijing City University, Shunyi District, Beijing, China
3 Digital Media Art Specialty, Department of Art and Design, Beijing City University, Shunyi District, Beijing, China
* Corresponding author: Mengyu Liu. Email: daimou971231@naver.com
Journal of Digital Frontier 2026, Vol. 1, No. 1, pp. 61-82
DOI: 10.67541/jdf2603
Received: 31 May 2026; Revised: 15 July 2026; Accepted: 23 July 2026; Published: 6 August 2026
Abstract

As an important carrier of the spirit of place, soundscape lacks systematic quantitative evaluation tools. The fundamental bottleneck lies in the lack of an accurate cross-modal alignment calculation framework between soundscape temporal sequence signals and visual image spatial semantics. In this paper, we propose a joint embedding model driven by cross-modal contrastive learning to realize the dynamic coupling of acoustic physical-perceptual features and visual global-local semantic features. We introduce a cross-modal attention forgetting gate to suppress modality-specific noise, design a hierarchical contrastive loss function to embed spatial structure with both instance-level and category-level constraints, and construct a spatio-temporal adaptive weight regulator to realize spatially differentiated modulation of the modality fusion coefficient. Based on 8,432 groups of measured multimodal data (including day-night differences) from four typical community types in Beijing's old city, our experiments show that the proposed model achieves a cross-modal retrieval mAP of 76.82%, a perception score prediction RMSE as low as 0.318, and an intraclass correlation coefficient of 0.824 with manual scores, all of which are significantly better than those of the comparison methods (p < 0.001). Ablation experiments verify the independent contributions of the three innovations. Feature response analysis reveals that the partial correlation coefficient between soundscape roughness and visual building density exceeds 0.78, indicating a deep sensory coordination mechanism between auditory texture density and visual spatial envelopment perception. The proposed model provides an effective computational tool for quantitative diagnosis of multi-sensory quality in old city communities.

Keywords
Soundscape Visual imagery Multimodal learning Cross-modal correlation Beijing old city community
References
  1. Wang, S., Zhang, J., Wang, F., & Dong, Y. (2023). How to achieve a balance between functional improvement and heritage conservation? A case study on the renewal of old Beijing city. Sustainable Cities and Society, 98, 104790. DOI: 10.1016/j.scs.2023.104790
  2. Zhang, R., Martí Casanovas, M., Bosch González, M., & Sun, S. (2024). Revitalizing heritage: The role of urban morphology in creating public value in China's historic districts. Land, 13(11), 1919. DOI: 10.3390/land13111919
  3. Amir, S., Sadoway, D., & Dommaraju, P. (2023). Taming the noise: Soundscape and livability in a technocratic city-state. East Asian Science, Technology and Society: An International Journal, 17(1), 88-104. DOI: 10.1080/18752160.2021.1936749
  4. Gil-Sayas, S., Di Pierro, G., Tansini, A., Serra, S., Currò, D., Broatch, A., & Fontaras, G. (2024). Energy consumption of mobile air-conditioning systems in electrified vehicles under different ambient temperatures. International Journal of Engine Research, 25(2), 293-304. DOI: 10.1177/14680874231171303
  5. Bahgat, G., Al-Makhlasawy, R. M., Khairy, M., Nour, M., & Abdelfattah, A. (2026). Energy saving for air conditioning devices based on the amalgamation of intelligent models and occupancy detection. Journal of Electrical Systems and Information Technology, 13(1), 78. DOI: 10.1186/s43067-026-00379-1
  6. Zhang, D., Ni, J., & Shi, X. (2024). Study of low-temperature energy consumption optimization of battery electric vehicle air conditioning systems considering blower efficiency. Processes, 12(7), 1495. DOI: 10.3390/pr12071495
  7. Muhammad, K., Hussain, T., Ullah, H., Del Ser, J., Rezaei, M., Kumar, N., ... & De Albuquerque, V. H. C. (2022). Vision-based semantic segmentation in scene understanding for autonomous driving: Recent achievements, challenges, and outlooks. IEEE Transactions on Intelligent Transportation Systems, 23(12), 22694-22715. DOI: 10.1109/TITS.2022.3207665
  8. Emek Soylu, B., Guzel, M. S., Bostanci, G. E., Ekinci, F., Asuroglu, T., & Acici, K. (2023). Deep-learning-based approaches for semantic segmentation of natural scene images: A review. Electronics, 12(12), 2730. DOI: 10.3390/electronics12122730
  9. Liang, Y., Mitchell, A., Kang, J., & Aletta, F. (2026). A Review of Soundscape Datasets: Challenges and Prospects for Multimodal Research. IEEE Transactions on Affective Computing. DOI: 10.1109/TAFFC.2026.3659084
  10. Chen, M., Lin, Z., Song, X., Luo, Y., Duan, X., Li, S., & Xie, C. (2026). Theoretical Framework, Technical Evolution, and Future Prospects of Cross-Modal Mapping and Controllable Image Generation Under Multi-Source Heterogeneous Collaboration. Sensors, 26(10), 2972. DOI: 10.3390/s26102972
  11. Kothinti, S. R., & Elhilali, M. (2023). Are acoustics enough? Semantic effects on auditory salience in natural scenes. Frontiers in Psychology, 14, 1276237. DOI: 10.3389/fpsyg.2023.1276237
  12. Sun, J., Deng, L., Afouras, T., Owens, A., & Davis, A. (2023). Eventfulness for interactive video alignment. ACM Transactions on Graphics (TOG), 42(4), 1-10. DOI: 10.1145/3592118
  13. Zhuang, Y., Kang, Y., Fei, T., Bian, M., & Du, Y. (2024). From hearing to seeing: Linking auditory and visual place perceptions with soundscape-to-image generative artificial intelligence. Computers, Environment and Urban Systems, 110, 102122. DOI: 10.1016/j.compenvurbsys.2024.102122
  14. Chen, P., Huang, X., Fei, T., & Wang, S. (2026). Cross‐Modal Urban Sensing: Evaluating Sound–Vision Alignment Across Street‐Level and Aerial Imagery. Transactions in GIS, 30(2), e70246. DOI: 10.1111/tgis.70246
  15. Zhang, Y., Wu, M., & Cai, X. (2025). A dynamic cross-modal learning framework for joint text-to-audio grounding and acoustic scene classification in smart city environments. Digital Signal Processing, 167, 105444. DOI: 10.1016/j.dsp.2025.105444
  16. Wang, T., Li, F., Zhu, L., Li, J., Zhang, Z., & Shen, H. T. (2025). Cross-modal retrieval: a systematic review of methods and future directions. Proceedings of the IEEE, 112(11), 1716-1754. DOI: 10.48550/arXiv.2308.14263
  17. Yang, Y., Wang, D., Hu, Y., & Liang, L. (2025). Research on Time Synchronization Technology of Non-safety DCS System Based on IEEE1588v2 Timing Protocol. In International Conference on Nuclear Engineering (pp. 163-176). Singapore: Springer Nature Singapore. DOI: 10.1007/978-981-95-2921-6_13
  18. Patoli, A. A., & Fortino, G. (2025). FPGA-based system implementation of IEEE 1588 precision time protocol: A review. IEEE Sensors Journal, 25(11), 18624-18642. DOI: 10.1109/JSEN.2025.3557277
  19. Iturbe-Martin, Z., Martín-Garín, A., & Casado-Rezola, A. (2026). Soundscape-Informed Urban Planning and Architecture in Historic Centers: A Multi-Layer Method for Soundscape Characterization Applied to Bilbao Old Town. Applied Sciences, 16(8), 3630. DOI: 10.3390/app16083630
  20. Versümer, S., Steffens, J., & Weinzierl, S. (2023). Day-to-day loudness assessments of indoor soundscapes: Exploring the impact of loudness indicators, person, and situation. The Journal of the Acoustical Society of America, 153(5), 2956. DOI: 10.1121/10.0019413
  21. Chen, Y., Lin, Y., Xu, R., & Vela, P. A. (2023). Wdiscood: Out-of-distribution detection via whitened linear discriminant analysis. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 5275-5284). IEEE. DOI: 10.48550/arXiv.2303.07543
  22. Gao, S., Yang, K., Shi, H., Wang, K., & Bai, J. (2022). Review on panoramic imaging and its applications in scene understanding. IEEE Transactions on Instrumentation and Measurement, 71, 1-34. DOI: 10.1109/TIM.2022.3216675
  23. Taha, K. (2026). Generative AI for multimodal content: a survey with empirical and experimental evaluations. Artificial Intelligence Review, 59(6), 140. DOI: 10.1007/s10462-026-11525-6
  24. Yan, T., Zhao, S., Hu, M., Wang, M., Zhang, X., Luo, Z., & Wang, M. (2024). HCL: A hierarchical contrastive learning framework for zero-shot relation extraction. IEEE Transactions on Neural Networks and Learning Systems, 36(3), 5694-5705. DOI: 10.1109/TNNLS.2024.3379527
  25. Zhong, B., Wang, P., & Wang, X. (2024). Ts-HCL: hierarchical layer-wise contrastive learning for unsupervised domain adaptation on time-series. In Asia-Pacific Web (APWeb) and Web-Age Information Management (WAIM) Joint International Conference on Web and Big Data (pp. 31-45). Singapore: Springer Nature Singapore. DOI: 10.1007/978-981-97-7238-4_3
  26. Xie, H., He, Y., Wu, X., & Lu, Y. (2022). Interplay between auditory and visual environments in historic districts: A big data approach based on social media. Environment and Planning B: Urban Analytics and City Science, 49(4), 1245-1265. DOI: 10.1177/23998083211059838
  27. Yao, X. W., Zheng, J. M., Li, Q., Yang, K. H., Yao, Z. H., & Shang, Y. H. (2025). MSTDFN: Multi-modal Spatio-Temporal Dynamic Fusion Network for Traffic Flow Prediction. In China Conference on Wireless Sensor Networks (pp. 189-202). Singapore: Springer Nature Singapore. DOI: 10.1007/978-981-95-9615-7_12
  28. Zhao, T., Chen, G., Suraphee, S., Phoophiwfa, T., & Busababodhin, P. (2025). A hybrid TCN-XGBoost model for agricultural product market price forecasting. PLoS One, 20(5), e0322496. DOI: 10.1371/journal.pone.0322496
  29. Zhao, T., Chen, G., Pang, C., & Busababodhin, P. (2025). Application and performance optimization of SLHS-TCN-XGBoost model in power demand forecasting. Computer Modeling in Engineering & Sciences, 143(3), 2883-2917. DOI: 10.32604/cmes.2025.066442
  30. Zhao, T., Chen, G., Pang, C., Seenoi, P., Papukdee, N., & Busababodhin, P. (2025). Time-lapse earthquake difference prediction based on physics-informed long short-term memory coupled with interpretability boosting. Journal of Seismic Exploration, 34(3), 25. DOI: 10.36922/JSE025310049
  31. Zhao, T., Chen, G., Pang, C., Li, L., & Busababodhin, P. (2026). Forecasting global agricultural trade imbalances using a hybrid deep learning and gradient boosting framework. Discover Computing, 29(1), 443. DOI: 10.1007/s10791-026-10367-8