Research on Rehabilitation Training Movement Recognition and Real-time Feedback Model Based on Computer Vision

Mingxiang Yang1*
1 College of Arts, Shandong Jianzhu University, Jinan, Shandong, China
* Corresponding author: Mingxiang Yang. Email: yangmingxiang10@outlook.com
Journal of Discovery Core 2026, Vol. 1, No. 1, pp. 101-128
DOI: 10.67541/jdc2605
Received: 20 June 2026; Revised: 18 July 2026; Accepted: 6 August 2026; Published: 18 August 2026
Abstract

Aiming at the key problems such as the separation of action recognition and quality assessment tasks, coarse feedback granularity and high labeling cost in the automatic evaluation of rehabilitation training, this paper proposes PDDS Net. The framework takes human skeleton sequence as input and realizes action classification and location-level deviation location synchronously through the collaborative architecture of a global action recognition stream and a local part evaluation stream. In the local flow, the joint nodes are divided into three functional part groups: upper limb, trunk and lower limb. independent graph convolutional subnetworks are used for decoupled coding, and a contrastive learning strategy is used to drive the attention map to focus on abnormal joints without frame level annotations. This is mainly achieved through the adaptive fusion mechanism to dynamically integrate the confidence of the deep network and the matching score of dynamic time warping template to maintain decision stability under conditions of pose degradation. Experiments on the NTU RGB+D and PKU MMD public datasets show that the action recognition accuracy of PDDS Net reaches 93.7%, the average precision of position level feedback is 0.656, 82.3% of the original AUC is still maintained at a 40% joint dropout rate, and the single frame inference latency is 41.8 ms, which meets the requirements of real-time interaction. This study provides a replicable technical path for the validation of rehabilitation evaluation algorithms without clinical data collection.

Keywords
Computer vision Rehabilitation training Motion quality assessment Graph convolutional network Contrastive learning
References
  1. Kamenov, K., Mills, J. A., Chatterji, S., & Cieza, A. (2019). Needs and unmet needs for rehabilitation services: a scoping review. Disability and Rehabilitation, 41(10), 1227-1237. DOI: 10.1080/09638288.2017.1422036
  2. Al Imam, M. H., Jahan, I., Das, M. C., Muhit, M., Akbar, D., Badawi, N., & Khandaker, G. (2022). Situation analysis of rehabilitation services for persons with disabilities in Bangladesh: identifying service gaps and scopes for improvement. Disability and Rehabilitation, 44(19), 5571-5584. DOI: 10.1080/09638288.2021.1939799
  3. Dutta, D., Sen, S., Aruchamy, S., & Mandal, S. (2022). Prevalence of post-stroke upper extremity paresis in developing countries and significance of m-Health for rehabilitation after stroke-A review. Smart Health, 23, 100264. DOI: 10.1016/j.smhl.2022.100264
  4. Martinez-Martin, E., & Cazorla, M. (2019). Rehabilitation technology: assistance from hospital to home. Computational Intelligence and Neuroscience, 2019(1), 1431509. DOI: 10.1155/2019/1431509
  5. Kyriazakos, S., Schlieter, H., Gand, K., Caprino, M., Corbo, M., Tropea, P., ... & Lynggaard, V. (2020). A novel virtual coaching system based on personalized clinical pathways for rehabilitation of older adults—requirements and implementation plan of the vCare project. Frontiers in Digital Health, 2, 546562. DOI: 10.3389/fdgth.2020.546562
  6. Zhou, Y., Rashid, F. A. N., Mat Daud, M., Hasan, M. K., & Chen, W. (2025). Machine learning-based computer vision for depth camera-based physiotherapy movement assessment: A systematic review. Sensors, 25(5), 1586. DOI: 10.3390/s25051586
  7. Hulleck, A. A., Menoth Mohan, D., Abdallah, N., El Rich, M., & Khalaf, K. (2022). Present and future of gait assessment in clinical practice: Towards the application of novel trends and technologies. Frontiers in Medical Technology, 4, 901331. DOI: 10.3389/fmedt.2022.901331
  8. Jeyabharathi, J., SK, G. P., VenuMadhav, G., Dinesh, G., & Saikumarreddy, C. V. V. (2025). Real Time Gameplay using Pose Detection. In 2025 International Conference on Multi-Agent Systems for Collaborative Intelligence (ICMSCI) (pp. 801-805). IEEE. DOI: 10.1109/ICMSCI62561.2025.10894142
  9. Sykes, E. R. (2025). Next-generation fall detection: harnessing human pose estimation and transformer technology. Health Systems, 14(2), 85-103. DOI: 10.1080/20476965.2024.2395574
  10. Yarko, E. I. (2025). Comparative Analysis of Libraries for Human Pose Detection in Mobile Device Environments. Automatic Documentation and Mathematical Linguistics, 59(Suppl 3), S220-S228. DOI: 10.3103/S0005105525700955
  11. Zhang, H. B., Cai, J. J., Qiu, H. M., Lei, Q., Liu, J. H., & Zhang, M. H. (2026). A survey of deep learning-based action quality assessment: methods and applications. Journal of Big Data, 13(1), 70. DOI: 10.1186/s40537-026-01409-5
  12. Trigeorgis, G., Nicolaou, M. A., Schuller, B. W., & Zafeiriou, S. (2017). Deep canonical time warping for simultaneous alignment and representation learning of sequences. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(5), 1128-1138. DOI: 10.1109/TPAMI.2017.2710047
  13. Wang, Z., Zhang, Z., Su, T., Ding, Z., & Zhao, T. (2024). Research on supply chain network optimisation based on the CNNs-BiLSTM model. In 2024 International Conference on Information Technology, Comunication Ecosystem and Management (ITCEM) (pp. 197-202). IEEE. DOI: 10.1109/ITCEM65710.2024.00044
  14. Zhao, T., Chen, G., Pang, C., Li, L., & Busababodhin, P. (2026). Forecasting global agricultural trade imbalances using a hybrid deep learning and gradient boosting framework. Discover Computing, 29(1), 443. DOI: 10.1007/s10791-026-10367-8
  15. Chen, G., Zhao, T., Pang, C., & Busababodhin, P. (2026). Integrated CNN–LSTM–XGBoost hybrid model predicts shale oil seismic attributes and global oil price trends. Scientific Reports. DOI: 10.1038/s41598-026-55910-1
  16. Kamal, S., Alshehri, M., AlQahtani, Y., Alshahrani, A., Almujally, N. A., Jalal, A., & Liu, H. (2025). A novel multi-modal rehabilitation monitoring over human motion intention recognition. Frontiers in Bioengineering and Biotechnology, 13, 1568690. DOI: 10.3389/fbioe.2025.1568690
  17. Momin, M. S., Sufian, A., Barman, D., Dutta, P., Dong, M., & Leo, M. (2022). In-home older adults' activity pattern monitoring using depth sensors: A review. Sensors, 22(23), 9067. DOI: 10.3390/s22239067
  18. Zhang, J., Wu, C., & Wang, Y. (2020). Human fall detection based on body posture spatio-temporal evolution. Sensors, 20(3), 946. DOI: 10.3390/s20030946
  19. Wang, H., Ho, E. S., Shum, H. P., & Zhu, Z. (2019). Spatio-temporal manifold learning for human motions via long-horizon modeling. IEEE Transactions on Visualization and Computer Graphics, 27(1), 216-227. DOI: 10.1109/TVCG.2019.2936810
  20. Yu, Q., Yang, G., Wang, X., Shi, Y., Feng, Y., & Liu, A. (2025). A review of time series forecasting and spatio-temporal series forecasting in deep learning. The Journal of Supercomputing, 81(10), 1160. DOI: 10.1007/s11227-025-07632-w
  21. Saravana, M. K., Roopa, M. S., Arunalatha, J. S., & Venugopal, K. R. (2026). Transformers for multivariate time series forecasting: Comprehensive analysis, challenges, research opportunities and future prospects. IEEE Access. DOI: 10.1109/ACCESS.2026.3654408
  22. Tian, Z., Zhao, F., Liu, H., & Bian, X. (2025). HDMFNet: Learning complementary spatiotemporal representations via heterogeneous dual-branch multimodal fusion network for human action recognition. Applied Soft Computing, 114391. DOI: 10.1016/j.asoc.2025.114391
  23. Chen, J., Zhang, Z., Sun, X., Zhou, Y., Zhou, Y., Zhao, Y., & Shi, J. (2026). A Multi-Source Data Fusion-Based Method for Safety Monitoring of Construction Workers on Concrete Placement Surfaces. Buildings, 16(6), 1165. DOI: 10.3390/buildings16061165
  24. Phu, K. A., & Hoang, V. D. (2025). Predicting occluded skeletal joints via tracking-based feature extraction. Neurocomputing, 131004. DOI: 10.1016/j.neucom.2025.131004
  25. Shahroudy, A., Liu, J., Ng, T. T., & Wang, G. (2016). Ntu rgb+ d: A large scale dataset for 3d human activity analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 1010-1019). DOI: 10.1109/CVPR.2016.115
  26. Wu, H., Wang, C., Duan, Y., & Song, R. (2026). Motion Saliency Guided Fine-Grained Multimodal Interactive Learning for RGB-D Action Recognition. IEEE Transactions on Circuits and Systems for Video Technology. DOI: 10.1109/TCSVT.2026.3708969
  27. Zhao, T., Chen, G., & Busababodhin, P. (2026). An intelligent expert system for logistics disruption prediction and mitigation in global supply chains: Integrating graph-based risk inference with ensemble forecasting. Scientific Reports. DOI: 10.1038/s41598-026-61617-0