|
Skeleton-based sign language recognition via bone-augmented graph neural networks and temporal modeling |
|---|---|
| รหัสดีโอไอ | |
| Title | Skeleton-based sign language recognition via bone-augmented graph neural networks and temporal modeling |
| Creator | Emma Anderson |
| Contributor | Cholwich Nattee, Advisor |
| Publisher | Thammasat University |
| Publication Year | 2568 |
| Keyword | Sign language recognition, Graph neural networks, Skeleton extraction, Temporal convolution, MediaPipe, WLASL dataset, Deep learning |
| Abstract | Sign language recognition is a challenging task that requires understanding both spatial relationships between body joints and temporal movement patterns. This paper proposes a hybrid skeleton-based architecture combining Graph Neural Networks (GNN) and 1D Convolution for American Sign Language (ASL) recognition, designed as a computationally efficient and interpretable alternative to appearance-based video models. Skeleton data is extracted using MediaPipe Holistic and enriched with bone features encoding spatial joint displacement to provide explicit structural context. Spatial relationships between joints are learned through residual DenseSAGEConv layers, while 1D convolution with max pooling captures temporal motion patterns across frames. Evaluated on the WLASL100 benchmark (100 classes, 2,038 samples), the model achieves 65.93% top-1, 87.25% top-5, and 91.67% top-10 accuracy. On top-5 and top-10 metrics, it surpasses the published I3D baseline, demonstrating that lightweight skeleton-based methods can be competitive with video-based approaches at higher recall levels. While the top-1 gap reflects the inherent difficulty of fine-grained disambiguation on a limited dataset, the consistent alignment between validation and test accuracy confirms stable generalization rather than overfitting. This work establishes that GNN architectures operating on skeleton graphs with bone features offer a practical, interpretable path toward deployable sign language recognition systems. |