Front-View Chinese Sign Recognition in a Continuous-Like Setting: A Reproducible Baseline and Evaluation Protocol

Authors

  • Mingze Han International Division, Beijing 101 High School, Beijing 100091, China

DOI:

https://doi.org/10.54691/5a05sk76

Keywords:

Chinese Sign Language; Language Resources; Reproducibility; Benchmark; Front-view RGB; Isolated-to-continuous Evaluation.

Abstract

This paper introduces a reproducible front-view benchmark setting for Chinese sign recognition in a continuous-like single-sign scenario. Building on the NationalCSL-DP resource, only red, green and blue (RGB) frames from the front camera are used as input, and this paper investigates the performance of a frame-based isolated sign recognition pipeline when exposed to untrimmed sign sequences instead of action-centric samples. To support reuse, the original per-sequence folders are converted into ImageFolder-style training data, a fixed 30-frame sampling policy is used, and preprocessing metadata is saved together with the checkpoints used for inference. The baseline is a two-tier pipeline consisting of a binary activity model and a ResNet-18 gloss classifier. Among the 6,707 front-view sequences, the resulting continuous-like diagnostic evaluation achieved a Top-1 accuracy of 0.2646, a Top-5 accuracy of 0.5184, an average prediction confidence of 0.4777, and zero no-segment failures under the selected segmentation policy. Further ablation studies have been conducted on frame budget, fusion strategy and segmentation thresholds, and the gap between action-centric frame classification and sequence-level continuous-like evaluation has also been analysed. Based on the above, boundary noise, single-view ambiguity and the mismatch between isolated training and sequence-level inference are all still problems. This study will serve as a reusable baseline and evaluation protocol for future research in front-view Chinese sign language resources, but it will not be a continuous state-of-the-art sign language recognition system.

Downloads

Download data is not yet available.

References

[1] Jin, P., et al. (2025). A large-scale dataset of Chinese national sign language for dual-view isolated sign language recognition. Scientific Data, 12, Article 660. https://doi.org/10.1038/s41597-025-04986-x

[2] Gao, W., Fang, G., Zhao, D., & Chen, Y. (2004). A Chinese sign language recognition system based on SOFM/SRN/HMM. Pattern Recognition, 37(12), 2389–2402.

[3] Koller, O., Forster, J., & Ney, H. (2015). Continuous sign language recognition: Towards large vocabulary statistical recognition systems handling multiple signers. Computer Vision and Image Understanding, 141, 108–125.

[4] Camgoz, N. C., Hadfield, S., Koller, O., Ney, H., & Bowden, R. (2018). Neural sign language translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 7784–7793).

[5] Li, D., Rodriguez, C., Yu, X., & Li, H. (2020). Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) (pp. 1459–1469).

[6] Vaezi Joze, H. R., & Koller, O. (2019). MS-ASL: A large-scale data set and benchmark for understanding American Sign Language. In Proceedings of the British Machine Vision Conference (BMVC).

[7] Huang, J., Zhou, W., Zhang, Q., Li, H., & Li, W. (2018). Video-based sign language recognition without temporal segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 32, No. 1, pp. 2257–2264).

[8] He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 770–778).

[9] Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009). ImageNet: A large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 248–255).

[10] Carreira, J., & Zisserman, A. (2017). Quo Vadis, action recognition? A new model and the Kinetics dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 6299–6308).

[11] Rastgoo, R., Kiani, K., & Escalera, S. (2021). Sign language recognition: A deep survey. Expert Systems with Applications, 164, 113794.

[12] Koller, O. (2020). Quantitative Survey of the State of the Art in Sign Language Recognition. arXiv preprint arXiv:2008.09918.

[13] Kvanchiani, K., Kraynov, R., Petrova, E., Surovcev, P., Nagaev, A., & Kapitanov, A. (2024). Training strategies for isolated sign language recognition. arXiv preprint arXiv:2412.11553.

[14] Koller, O., Zargaran, S., & Ney, H. (2017). Re-sign: Re-aligned end-to-end sequence modelling with deep recurrent CNN-HMMs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 4297–4305).

[15] Graves, A., Fernandez, S., Gomez, F., & Schmidhuber, J. (2006). Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd International Conference on Machine Learning (ICML) (pp. 369–376).

[16] Cui, R., Liu, H., & Zhang, C. (2019). A deep neural framework for continuous sign language recognition by iterative training. IEEE Transactions on Multimedia, 21(7), 1880–1891.

[17] Cheng, K. L., Yang, Z., Chen, Q., & Tai, Y. W. (2020). Fully convolutional networks for continuous sign language recognition. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 697–714).

[18] Pu, J., Zhou, W., & Li, H. (2019). Iterative alignment network for continuous sign language recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 4165–4174).

Downloads

Published

2026-06-23

Issue

Section

Articles

How to Cite

Han, Mingze. 2026. “Front-View Chinese Sign Recognition in a Continuous-Like Setting: A Reproducible Baseline and Evaluation Protocol”. Scientific Journal of Intelligent Systems Research 8 (6): 7-14. https://doi.org/10.54691/5a05sk76.