Skip to main navigation Skip to search Skip to main content

CNText2Sign and CNSign: Unified Chinese Sign Language Datasets for Bidirectional Accessibility

  • Yulong Li
  • , Yuxuan Zhang
  • , Feilong Tang
  • , Ming Hu
  • , Zhixiang Lu
  • , Haochen Xue
  • , Jianghao Wu
  • , Rui Chen
  • , Mian Zhou
  • , Kang Dang
  • , Chong Li
  • , Yifang Wang
  • , Imran Razzak*
  • , Jionglong Su*
  • *Corresponding author for this work
  • Xi'an Jiaotong-Liverpool University
  • Monash University
  • Mohamed Bin Zayed University of Artificial Intelligence

Research output: Chapter in Book or Report/Conference proceedingConference Proceedingpeer-review

Abstract

Sign language is the primary communication mode for 72 million hearing-impaired individuals worldwide, necessitating effective bidirectional Sign Language Production and Sign Language Translation systems. However, functional bidirectional systems require a unified linguistic environment, hindered by the lack of suitable unified datasets, particularly those providing the necessary pose information for accurate Sign Language Production (SLP) evaluation. Concurrently, current SLP evaluation methods like back-translation ignore pose accuracy, and high-quality coordinated generation remains challenging. To create this crucial environment and overcome these challenges, we introduce CNText2Sign and CNSign, which together constitute the first unified dataset aimed at supporting bidirectional accessibility systems for Chinese sign language; CNText2Sign provides 15,000 natural language-to-sign mappings and standardized skeletal keypoints for 8,643 vocabulary items supporting pose assessment. Building upon this foundation, we propose the AuraLLM model, which leverages a decoupled architecture with CNText2Sign's pose data for novel direct gesture accuracy assessment. The model employs retrieval augmentation and Cascading Vocabulary Resolution to handle semantic mapping and out-of-vocabulary words, and achieves all-scenario production with controllable coordination of gestures and facial expressions via pose-conditioned video synthesis. Concurrently, our Sign Language Translation model SignMST-C employs targeted self-supervised pretraining for dynamic feature capture, achieving new SOTA results on PHOENIX2014-T with BLEU-4 scores up to 32.08. AuraLLM establishes a strong performance baseline on CNText2Sign with a BLEU-4 score of 50.41 under direct evaluation.

Original languageEnglish
Title of host publicationKDD 2026 - Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1
PublisherAssociation for Computing Machinery
Pages2723-2734
Number of pages12
ISBN (Electronic)9798400722585
DOIs
Publication statusPublished - 20 Apr 2026
Event32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, KDD 2026 - Jeju Island, Korea, Republic of
Duration: 9 Aug 202613 Aug 2026

Publication series

NameProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
Volume1-A
ISSN (Print)2154-817X

Conference

Conference32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, KDD 2026
Country/TerritoryKorea, Republic of
CityJeju Island
Period9/08/2613/08/26

Keywords

  • all-scenario adaptability
  • bidirectional accessibility
  • out-of-vocabulary handling
  • sign language production and translation

Cite this