Research on Rehabilitation Training Movement Recognition and Real-time Feedback Model Based on Computer Vision
1 College of Arts, Shandong Jianzhu University, Jinan, Shandong, China
* Corresponding author: Mingxiang Yang. Email: yangmingxiang10@outlook.com
Abstract
Aiming at the key problems such as the separation of action recognition and quality assessment tasks, coarse feedback granularity and high labeling cost in the automatic evaluation of rehabilitation training, this paper proposes PDDS Net. The framework takes human skeleton sequence as input and realizes action classification and location-level deviation location synchronously through the collaborative architecture of a global action recognition stream and a local part evaluation stream. In the local flow, the joint nodes are divided into three functional part groups: upper limb, trunk and lower limb. independent graph convolutional subnetworks are used for decoupled coding, and a contrastive learning strategy is used to drive the attention map to focus on abnormal joints without frame level annotations. This is mainly achieved through the adaptive fusion mechanism to dynamically integrate the confidence of the deep network and the matching score of dynamic time warping template to maintain decision stability under conditions of pose degradation. Experiments on the NTU RGB+D and PKU MMD public datasets show that the action recognition accuracy of PDDS Net reaches 93.7%, the average precision of position level feedback is 0.656, 82.3% of the original AUC is still maintained at a 40% joint dropout rate, and the single frame inference latency is 41.8 ms, which meets the requirements of real-time interaction. This study provides a replicable technical path for the validation of rehabilitation evaluation algorithms without clinical data collection.
Keywords
Computer vision
Rehabilitation training
Motion quality assessment
Graph convolutional network
Contrastive learning
References
- Kamenov, K., Mills, J. A., Chatterji, S., & Cieza, A. (2019). Needs and unmet needs for rehabilitation services: a scoping review. Disability and Rehabilitation, 41(10), 1227-1237. DOI: 10.1080/09638288.2017.1422036
- Al Imam, M. H., Jahan, I., Das, M. C., Muhit, M., Akbar, D., Badawi, N., & Khandaker, G. (2022). Situation analysis of rehabilitation services for persons with disabilities in Bangladesh: identifying service gaps and scopes for improvement. Disability and Rehabilitation, 44(19), 5571-5584. DOI: 10.1080/09638288.2021.1939799
- Dutta, D., Sen, S., Aruchamy, S., & Mandal, S. (2022). Prevalence of post-stroke upper extremity paresis in developing countries and significance of m-Health for rehabilitation after stroke-A review. Smart Health, 23, 100264. DOI: 10.1016/j.smhl.2022.100264
- Martinez-Martin, E., & Cazorla, M. (2019). Rehabilitation technology: assistance from hospital to home. Computational Intelligence and Neuroscience, 2019(1), 1431509. DOI: 10.1155/2019/1431509
- Kyriazakos, S., Schlieter, H., Gand, K., Caprino, M., Corbo, M., Tropea, P., ... & Lynggaard, V. (2020). A novel virtual coaching system based on personalized clinical pathways for rehabilitation of older adults—requirements and implementation plan of the vCare project. Frontiers in Digital Health, 2, 546562. DOI: 10.3389/fdgth.2020.546562
- Zhou, Y., Rashid, F. A. N., Mat Daud, M., Hasan, M. K., & Chen, W. (2025). Machine learning-based computer vision for depth camera-based physiotherapy movement assessment: A systematic review. Sensors, 25(5), 1586. DOI: 10.3390/s25051586
- Hulleck, A. A., Menoth Mohan, D., Abdallah, N., El Rich, M., & Khalaf, K. (2022). Present and future of gait assessment in clinical practice: Towards the application of novel trends and technologies. Frontiers in Medical Technology, 4, 901331. DOI: 10.3389/fmedt.2022.901331
- Jeyabharathi, J., SK, G. P., VenuMadhav, G., Dinesh, G., & Saikumarreddy, C. V. V. (2025). Real Time Gameplay using Pose Detection. In 2025 International Conference on Multi-Agent Systems for Collaborative Intelligence (ICMSCI) (pp. 801-805). IEEE. DOI: 10.1109/ICMSCI62561.2025.10894142
- Sykes, E. R. (2025). Next-generation fall detection: harnessing human pose estimation and transformer technology. Health Systems, 14(2), 85-103. DOI: 10.1080/20476965.2024.2395574
- Yarko, E. I. (2025). Comparative Analysis of Libraries for Human Pose Detection in Mobile Device Environments. Automatic Documentation and Mathematical Linguistics, 59(Suppl 3), S220-S228. DOI: 10.3103/S0005105525700955
- Zhang, H. B., Cai, J. J., Qiu, H. M., Lei, Q., Liu, J. H., & Zhang, M. H. (2026). A survey of deep learning-based action quality assessment: methods and applications. Journal of Big Data, 13(1), 70. DOI: 10.1186/s40537-026-01409-5
- Trigeorgis, G., Nicolaou, M. A., Schuller, B. W., & Zafeiriou, S. (2017). Deep canonical time warping for simultaneous alignment and representation learning of sequences. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(5), 1128-1138. DOI: 10.1109/TPAMI.2017.2710047
- Wang, Z., Zhang, Z., Su, T., Ding, Z., & Zhao, T. (2024). Research on supply chain network optimisation based on the CNNs-BiLSTM model. In 2024 International Conference on Information Technology, Comunication Ecosystem and Management (ITCEM) (pp. 197-202). IEEE. DOI: 10.1109/ITCEM65710.2024.00044
- Zhao, T., Chen, G., Pang, C., Li, L., & Busababodhin, P. (2026). Forecasting global agricultural trade imbalances using a hybrid deep learning and gradient boosting framework. Discover Computing, 29(1), 443. DOI: 10.1007/s10791-026-10367-8
- Chen, G., Zhao, T., Pang, C., & Busababodhin, P. (2026). Integrated CNN–LSTM–XGBoost hybrid model predicts shale oil seismic attributes and global oil price trends. Scientific Reports. DOI: 10.1038/s41598-026-55910-1
- Kamal, S., Alshehri, M., AlQahtani, Y., Alshahrani, A., Almujally, N. A., Jalal, A., & Liu, H. (2025). A novel multi-modal rehabilitation monitoring over human motion intention recognition. Frontiers in Bioengineering and Biotechnology, 13, 1568690. DOI: 10.3389/fbioe.2025.1568690
- Momin, M. S., Sufian, A., Barman, D., Dutta, P., Dong, M., & Leo, M. (2022). In-home older adults' activity pattern monitoring using depth sensors: A review. Sensors, 22(23), 9067. DOI: 10.3390/s22239067
- Zhang, J., Wu, C., & Wang, Y. (2020). Human fall detection based on body posture spatio-temporal evolution. Sensors, 20(3), 946. DOI: 10.3390/s20030946
- Wang, H., Ho, E. S., Shum, H. P., & Zhu, Z. (2019). Spatio-temporal manifold learning for human motions via long-horizon modeling. IEEE Transactions on Visualization and Computer Graphics, 27(1), 216-227. DOI: 10.1109/TVCG.2019.2936810
- Yu, Q., Yang, G., Wang, X., Shi, Y., Feng, Y., & Liu, A. (2025). A review of time series forecasting and spatio-temporal series forecasting in deep learning. The Journal of Supercomputing, 81(10), 1160. DOI: 10.1007/s11227-025-07632-w
- Saravana, M. K., Roopa, M. S., Arunalatha, J. S., & Venugopal, K. R. (2026). Transformers for multivariate time series forecasting: Comprehensive analysis, challenges, research opportunities and future prospects. IEEE Access. DOI: 10.1109/ACCESS.2026.3654408
- Tian, Z., Zhao, F., Liu, H., & Bian, X. (2025). HDMFNet: Learning complementary spatiotemporal representations via heterogeneous dual-branch multimodal fusion network for human action recognition. Applied Soft Computing, 114391. DOI: 10.1016/j.asoc.2025.114391
- Chen, J., Zhang, Z., Sun, X., Zhou, Y., Zhou, Y., Zhao, Y., & Shi, J. (2026). A Multi-Source Data Fusion-Based Method for Safety Monitoring of Construction Workers on Concrete Placement Surfaces. Buildings, 16(6), 1165. DOI: 10.3390/buildings16061165
- Phu, K. A., & Hoang, V. D. (2025). Predicting occluded skeletal joints via tracking-based feature extraction. Neurocomputing, 131004. DOI: 10.1016/j.neucom.2025.131004
- Shahroudy, A., Liu, J., Ng, T. T., & Wang, G. (2016). Ntu rgb+ d: A large scale dataset for 3d human activity analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 1010-1019). DOI: 10.1109/CVPR.2016.115
- Wu, H., Wang, C., Duan, Y., & Song, R. (2026). Motion Saliency Guided Fine-Grained Multimodal Interactive Learning for RGB-D Action Recognition. IEEE Transactions on Circuits and Systems for Video Technology. DOI: 10.1109/TCSVT.2026.3708969
- Zhao, T., Chen, G., & Busababodhin, P. (2026). An intelligent expert system for logistics disruption prediction and mitigation in global supply chains: Integrating graph-based risk inference with ensemble forecasting. Scientific Reports. DOI: 10.1038/s41598-026-61617-0