Computational Modeling of Soundscape and Visual Imagery in Beijing Old City Communities Based on Multi-modal Data
Ziyan Shen1, Xuanyi Qi2, Mengyu Liu3*
1 College of Mechanical and Electrical Engineering, Beijing University of Chemical Technology, Chaoyang District, Beijing, China 2 Department of Art and Design, Beijing City University, Shunyi District, Beijing, China 3 Digital Media Art Specialty, Department of Art and Design, Beijing City University, Shunyi District, Beijing, China
Journal of Digital Frontier2026, Vol. 1, No. 1, pp. 61-82
DOI: 10.67541/jdf2603
Received: 31 May 2026; Revised: 15 July 2026; Accepted: 23 July 2026; Published: 6 August 2026
Abstract
As an important carrier of the spirit of place, soundscape lacks systematic quantitative evaluation tools. The fundamental bottleneck lies in the lack of an accurate cross-modal alignment calculation framework between soundscape temporal sequence signals and visual image spatial semantics. In this paper, we propose a joint embedding model driven by cross-modal contrastive learning to realize the dynamic coupling of acoustic physical-perceptual features and visual global-local semantic features. We introduce a cross-modal attention forgetting gate to suppress modality-specific noise, design a hierarchical contrastive loss function to embed spatial structure with both instance-level and category-level constraints, and construct a spatio-temporal adaptive weight regulator to realize spatially differentiated modulation of the modality fusion coefficient. Based on 8,432 groups of measured multimodal data (including day-night differences) from four typical community types in Beijing's old city, our experiments show that the proposed model achieves a cross-modal retrieval mAP of 76.82%, a perception score prediction RMSE as low as 0.318, and an intraclass correlation coefficient of 0.824 with manual scores, all of which are significantly better than those of the comparison methods (p < 0.001). Ablation experiments verify the independent contributions of the three innovations. Feature response analysis reveals that the partial correlation coefficient between soundscape roughness and visual building density exceeds 0.78, indicating a deep sensory coordination mechanism between auditory texture density and visual spatial envelopment perception. The proposed model provides an effective computational tool for quantitative diagnosis of multi-sensory quality in old city communities.
Keywords
SoundscapeVisual imageryMultimodal learningCross-modal correlationBeijing old city community
References
Wang, S., Zhang, J., Wang, F., & Dong, Y. (2023). How to achieve a balance between functional improvement and heritage conservation? A case study on the renewal of old Beijing city. Sustainable Cities and Society, 98, 104790. DOI: 10.1016/j.scs.2023.104790
Zhang, R., Martí Casanovas, M., Bosch González, M., & Sun, S. (2024). Revitalizing heritage: The role of urban morphology in creating public value in China's historic districts. Land, 13(11), 1919. DOI: 10.3390/land13111919
Amir, S., Sadoway, D., & Dommaraju, P. (2023). Taming the noise: Soundscape and livability in a technocratic city-state. East Asian Science, Technology and Society: An International Journal, 17(1), 88-104. DOI: 10.1080/18752160.2021.1936749
Gil-Sayas, S., Di Pierro, G., Tansini, A., Serra, S., Currò, D., Broatch, A., & Fontaras, G. (2024). Energy consumption of mobile air-conditioning systems in electrified vehicles under different ambient temperatures. International Journal of Engine Research, 25(2), 293-304. DOI: 10.1177/14680874231171303
Bahgat, G., Al-Makhlasawy, R. M., Khairy, M., Nour, M., & Abdelfattah, A. (2026). Energy saving for air conditioning devices based on the amalgamation of intelligent models and occupancy detection. Journal of Electrical Systems and Information Technology, 13(1), 78. DOI: 10.1186/s43067-026-00379-1
Zhang, D., Ni, J., & Shi, X. (2024). Study of low-temperature energy consumption optimization of battery electric vehicle air conditioning systems considering blower efficiency. Processes, 12(7), 1495. DOI: 10.3390/pr12071495
Muhammad, K., Hussain, T., Ullah, H., Del Ser, J., Rezaei, M., Kumar, N., ... & De Albuquerque, V. H. C. (2022). Vision-based semantic segmentation in scene understanding for autonomous driving: Recent achievements, challenges, and outlooks. IEEE Transactions on Intelligent Transportation Systems, 23(12), 22694-22715. DOI: 10.1109/TITS.2022.3207665
Emek Soylu, B., Guzel, M. S., Bostanci, G. E., Ekinci, F., Asuroglu, T., & Acici, K. (2023). Deep-learning-based approaches for semantic segmentation of natural scene images: A review. Electronics, 12(12), 2730. DOI: 10.3390/electronics12122730
Liang, Y., Mitchell, A., Kang, J., & Aletta, F. (2026). A Review of Soundscape Datasets: Challenges and Prospects for Multimodal Research. IEEE Transactions on Affective Computing. DOI: 10.1109/TAFFC.2026.3659084
Chen, M., Lin, Z., Song, X., Luo, Y., Duan, X., Li, S., & Xie, C. (2026). Theoretical Framework, Technical Evolution, and Future Prospects of Cross-Modal Mapping and Controllable Image Generation Under Multi-Source Heterogeneous Collaboration. Sensors, 26(10), 2972. DOI: 10.3390/s26102972
Kothinti, S. R., & Elhilali, M. (2023). Are acoustics enough? Semantic effects on auditory salience in natural scenes. Frontiers in Psychology, 14, 1276237. DOI: 10.3389/fpsyg.2023.1276237
Sun, J., Deng, L., Afouras, T., Owens, A., & Davis, A. (2023). Eventfulness for interactive video alignment. ACM Transactions on Graphics (TOG), 42(4), 1-10. DOI: 10.1145/3592118
Zhuang, Y., Kang, Y., Fei, T., Bian, M., & Du, Y. (2024). From hearing to seeing: Linking auditory and visual place perceptions with soundscape-to-image generative artificial intelligence. Computers, Environment and Urban Systems, 110, 102122. DOI: 10.1016/j.compenvurbsys.2024.102122
Chen, P., Huang, X., Fei, T., & Wang, S. (2026). Cross‐Modal Urban Sensing: Evaluating Sound–Vision Alignment Across Street‐Level and Aerial Imagery. Transactions in GIS, 30(2), e70246. DOI: 10.1111/tgis.70246
Zhang, Y., Wu, M., & Cai, X. (2025). A dynamic cross-modal learning framework for joint text-to-audio grounding and acoustic scene classification in smart city environments. Digital Signal Processing, 167, 105444. DOI: 10.1016/j.dsp.2025.105444
Wang, T., Li, F., Zhu, L., Li, J., Zhang, Z., & Shen, H. T. (2025). Cross-modal retrieval: a systematic review of methods and future directions. Proceedings of the IEEE, 112(11), 1716-1754. DOI: 10.48550/arXiv.2308.14263
Yang, Y., Wang, D., Hu, Y., & Liang, L. (2025). Research on Time Synchronization Technology of Non-safety DCS System Based on IEEE1588v2 Timing Protocol. In International Conference on Nuclear Engineering (pp. 163-176). Singapore: Springer Nature Singapore. DOI: 10.1007/978-981-95-2921-6_13
Patoli, A. A., & Fortino, G. (2025). FPGA-based system implementation of IEEE 1588 precision time protocol: A review. IEEE Sensors Journal, 25(11), 18624-18642. DOI: 10.1109/JSEN.2025.3557277
Iturbe-Martin, Z., Martín-Garín, A., & Casado-Rezola, A. (2026). Soundscape-Informed Urban Planning and Architecture in Historic Centers: A Multi-Layer Method for Soundscape Characterization Applied to Bilbao Old Town. Applied Sciences, 16(8), 3630. DOI: 10.3390/app16083630
Versümer, S., Steffens, J., & Weinzierl, S. (2023). Day-to-day loudness assessments of indoor soundscapes: Exploring the impact of loudness indicators, person, and situation. The Journal of the Acoustical Society of America, 153(5), 2956. DOI: 10.1121/10.0019413
Chen, Y., Lin, Y., Xu, R., & Vela, P. A. (2023). Wdiscood: Out-of-distribution detection via whitened linear discriminant analysis. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 5275-5284). IEEE. DOI: 10.48550/arXiv.2303.07543
Gao, S., Yang, K., Shi, H., Wang, K., & Bai, J. (2022). Review on panoramic imaging and its applications in scene understanding. IEEE Transactions on Instrumentation and Measurement, 71, 1-34. DOI: 10.1109/TIM.2022.3216675
Taha, K. (2026). Generative AI for multimodal content: a survey with empirical and experimental evaluations. Artificial Intelligence Review, 59(6), 140. DOI: 10.1007/s10462-026-11525-6
Yan, T., Zhao, S., Hu, M., Wang, M., Zhang, X., Luo, Z., & Wang, M. (2024). HCL: A hierarchical contrastive learning framework for zero-shot relation extraction. IEEE Transactions on Neural Networks and Learning Systems, 36(3), 5694-5705. DOI: 10.1109/TNNLS.2024.3379527
Zhong, B., Wang, P., & Wang, X. (2024). Ts-HCL: hierarchical layer-wise contrastive learning for unsupervised domain adaptation on time-series. In Asia-Pacific Web (APWeb) and Web-Age Information Management (WAIM) Joint International Conference on Web and Big Data (pp. 31-45). Singapore: Springer Nature Singapore. DOI: 10.1007/978-981-97-7238-4_3
Xie, H., He, Y., Wu, X., & Lu, Y. (2022). Interplay between auditory and visual environments in historic districts: A big data approach based on social media. Environment and Planning B: Urban Analytics and City Science, 49(4), 1245-1265. DOI: 10.1177/23998083211059838
Yao, X. W., Zheng, J. M., Li, Q., Yang, K. H., Yao, Z. H., & Shang, Y. H. (2025). MSTDFN: Multi-modal Spatio-Temporal Dynamic Fusion Network for Traffic Flow Prediction. In China Conference on Wireless Sensor Networks (pp. 189-202). Singapore: Springer Nature Singapore. DOI: 10.1007/978-981-95-9615-7_12
Zhao, T., Chen, G., Suraphee, S., Phoophiwfa, T., & Busababodhin, P. (2025). A hybrid TCN-XGBoost model for agricultural product market price forecasting. PLoS One, 20(5), e0322496. DOI: 10.1371/journal.pone.0322496
Zhao, T., Chen, G., Pang, C., & Busababodhin, P. (2025). Application and performance optimization of SLHS-TCN-XGBoost model in power demand forecasting. Computer Modeling in Engineering & Sciences, 143(3), 2883-2917. DOI: 10.32604/cmes.2025.066442
Zhao, T., Chen, G., Pang, C., Seenoi, P., Papukdee, N., & Busababodhin, P. (2025). Time-lapse earthquake difference prediction based on physics-informed long short-term memory coupled with interpretability boosting. Journal of Seismic Exploration, 34(3), 25. DOI: 10.36922/JSE025310049
Zhao, T., Chen, G., Pang, C., Li, L., & Busababodhin, P. (2026). Forecasting global agricultural trade imbalances using a hybrid deep learning and gradient boosting framework. Discover Computing, 29(1), 443. DOI: 10.1007/s10791-026-10367-8