Front-View Chinese Sign Recognition in a Continuous-Like Setting: A Reproducible Baseline and Evaluation Protocol
DOI:
https://doi.org/10.54691/5a05sk76Keywords:
Chinese Sign Language; Language Resources; Reproducibility; Benchmark; Front-view RGB; Isolated-to-continuous Evaluation.Abstract
This paper introduces a reproducible front-view benchmark setting for Chinese sign recognition in a continuous-like single-sign scenario. Building on the NationalCSL-DP resource, only red, green and blue (RGB) frames from the front camera are used as input, and this paper investigates the performance of a frame-based isolated sign recognition pipeline when exposed to untrimmed sign sequences instead of action-centric samples. To support reuse, the original per-sequence folders are converted into ImageFolder-style training data, a fixed 30-frame sampling policy is used, and preprocessing metadata is saved together with the checkpoints used for inference. The baseline is a two-tier pipeline consisting of a binary activity model and a ResNet-18 gloss classifier. Among the 6,707 front-view sequences, the resulting continuous-like diagnostic evaluation achieved a Top-1 accuracy of 0.2646, a Top-5 accuracy of 0.5184, an average prediction confidence of 0.4777, and zero no-segment failures under the selected segmentation policy. Further ablation studies have been conducted on frame budget, fusion strategy and segmentation thresholds, and the gap between action-centric frame classification and sequence-level continuous-like evaluation has also been analysed. Based on the above, boundary noise, single-view ambiguity and the mismatch between isolated training and sequence-level inference are all still problems. This study will serve as a reusable baseline and evaluation protocol for future research in front-view Chinese sign language resources, but it will not be a continuous state-of-the-art sign language recognition system.
Downloads
References
[1] Jin, P., et al. (2025). A large-scale dataset of Chinese national sign language for dual-view isolated sign language recognition. Scientific Data, 12, Article 660. https://doi.org/10.1038/s41597-025-04986-x
[2] Gao, W., Fang, G., Zhao, D., & Chen, Y. (2004). A Chinese sign language recognition system based on SOFM/SRN/HMM. Pattern Recognition, 37(12), 2389–2402.
[3] Koller, O., Forster, J., & Ney, H. (2015). Continuous sign language recognition: Towards large vocabulary statistical recognition systems handling multiple signers. Computer Vision and Image Understanding, 141, 108–125.
[4] Camgoz, N. C., Hadfield, S., Koller, O., Ney, H., & Bowden, R. (2018). Neural sign language translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 7784–7793).
[5] Li, D., Rodriguez, C., Yu, X., & Li, H. (2020). Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) (pp. 1459–1469).
[6] Vaezi Joze, H. R., & Koller, O. (2019). MS-ASL: A large-scale data set and benchmark for understanding American Sign Language. In Proceedings of the British Machine Vision Conference (BMVC).
[7] Huang, J., Zhou, W., Zhang, Q., Li, H., & Li, W. (2018). Video-based sign language recognition without temporal segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 32, No. 1, pp. 2257–2264).
[8] He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 770–778).
[9] Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009). ImageNet: A large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 248–255).
[10] Carreira, J., & Zisserman, A. (2017). Quo Vadis, action recognition? A new model and the Kinetics dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 6299–6308).
[11] Rastgoo, R., Kiani, K., & Escalera, S. (2021). Sign language recognition: A deep survey. Expert Systems with Applications, 164, 113794.
[12] Koller, O. (2020). Quantitative Survey of the State of the Art in Sign Language Recognition. arXiv preprint arXiv:2008.09918.
[13] Kvanchiani, K., Kraynov, R., Petrova, E., Surovcev, P., Nagaev, A., & Kapitanov, A. (2024). Training strategies for isolated sign language recognition. arXiv preprint arXiv:2412.11553.
[14] Koller, O., Zargaran, S., & Ney, H. (2017). Re-sign: Re-aligned end-to-end sequence modelling with deep recurrent CNN-HMMs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 4297–4305).
[15] Graves, A., Fernandez, S., Gomez, F., & Schmidhuber, J. (2006). Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd International Conference on Machine Learning (ICML) (pp. 369–376).
[16] Cui, R., Liu, H., & Zhang, C. (2019). A deep neural framework for continuous sign language recognition by iterative training. IEEE Transactions on Multimedia, 21(7), 1880–1891.
[17] Cheng, K. L., Yang, Z., Chen, Q., & Tai, Y. W. (2020). Fully convolutional networks for continuous sign language recognition. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 697–714).
[18] Pu, J., Zhou, W., & Li, H. (2019). Iterative alignment network for continuous sign language recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 4165–4174).
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Scientific Journal of Intelligent Systems Research

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.




