Skip to main navigation Skip to search Skip to main content

KD-MSLRT: Lightweight Sign Language Recognition Model Based on Mediapipe and 3D to 1D Knowledge Distillation

    • Xi'an Jiaotong-Liverpool University

    Research output: Chapter in Book or Report/Conference proceedingConference Proceedingpeer-review

    4 Citations (Scopus)

    Abstract

    Artificial intelligence has achieved notable results in sign language recognition and translation. However, relatively few efforts have been made to significantly improve the quality of life for the 72 million hearing-impaired people worldwide. Sign language translation models, relying on video inputs, involves with large parameter sizes, making it time-consuming and computationally intensive to be deployed. This directly contributes to the scarcity of human-centered technology in this field. Additionally, the lack of datasets in sign language translation hampers research progress in this area. To address these, we first propose a cross-modal multi-knowledge distillation technique from 3D to 1D and a novel end-to-end pre-training text correction framework. Compared to other pretrained models, our framework achieves significant advancements in correcting text output errors. Our model achieves a decrease in Word Error Rate (WER) of at least 1.4% on PHOENIX14 and PHOENIX14T datasets compared to the state-of-the-art CorrNet. Additionally, the TensorFlow Lite (TFLite) quantized model size is reduced to 12.93 MB, making it the smallest, fastest, and most accurate model to date. We have also collected and released extensive Chinese sign language datasets, and developed a specialized training vocabulary. To address the lack of research on data augmentation for landmark data, we have designed comparative experiments on various augmentation methods. Moreover, we performed a simulated deployment and prediction of our model on Intel platform CPUs and assessed the feasibility of deploying the model on other platforms.

    Original languageEnglish
    Title of host publicationSpecial Track on AI Alignment
    EditorsToby Walsh, Julie Shah, Zico Kolter
    PublisherAssociation for the Advancement of Artificial Intelligence
    Pages28177-28185
    Number of pages9
    Edition27
    ISBN (Electronic)157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978
    DOIs
    Publication statusPublished - 11 Apr 2025
    Event39th Annual AAAI Conference on Artificial Intelligence, AAAI 2025 - Philadelphia, United States
    Duration: 25 Feb 20254 Mar 2025

    Publication series

    NameProceedings of the AAAI Conference on Artificial Intelligence
    Number27
    Volume39
    ISSN (Print)2159-5399
    ISSN (Electronic)2374-3468

    Conference

    Conference39th Annual AAAI Conference on Artificial Intelligence, AAAI 2025
    Country/TerritoryUnited States
    CityPhiladelphia
    Period25/02/254/03/25

    Fingerprint

    Dive into the research topics of 'KD-MSLRT: Lightweight Sign Language Recognition Model Based on Mediapipe and 3D to 1D Knowledge Distillation'. Together they form a unique fingerprint.

    Cite this