Skl-MHI: Efficient Human Action Recognition through Skeleton Temporal Compression

Skeleton motion history based human action recognition using deep learning

2017-10-01
Cho Nilar Phyo, Thi Thi Zin, Pyke Tin
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a Human Action Recognition (HAR) system that combines Skeleton Motion History Images (Skl MHI) with a 2D Deep Convolutional Neural Network (2D-DCNN). By encoding temporal dynamics into a single 2D skeleton-based representation, the method achieves a high recognition accuracy of 97.82% across 10 distinct action categories using Kinect V2 sensor data.

TL;DR

The paper presents a streamlined approach to Human Action Recognition (HAR) by transforming 3D skeleton sequences into 2D Skeleton Motion History Images (Skl MHI). By training a lightweight 2D-DCNN on these compressed representations, the researchers achieved a 97.82% accuracy while significantly reducing the computational burden compared to traditional raw-video deep learning.

Problem & Motivation: The High Cost of Action Recognition

Recognizing human actions—like walking, bending, or waving—is crucial for elderly monitoring and human-robot interaction. However, the field has long struggled with two primary hurdles:

  1. Handcrafted Frailty: Older methods depend on manual feature extraction tailored to specific environments, making them brittle when background or lighting changes.
  2. Computational Complexity: Directly processing 3D video volumes (RGB-T) requires massive memory and GPU power.

The authors' insight is to leverage the Microsoft Kinect V2 to extract skeleton joint points first. This removes environmental noise (background clutter) and reduces the data dimensionality from millions of pixels to just 25 joint coordinates, focusing purely on human kinematics.

Methodology: From Skeleton Sequences to Skl MHI

The core innovation lies in how temporal information is packed into a static image.

1. The Skl MHI Generation

Instead of feeding a sequence of frames, the system takes 9 consecutive skeleton frames. It performs a binary OR-operation across these frames to create a "motion trail" or a history of the movement. This result is normalized into a 62×62 image, creating a signature of the action's trajectory.

Process flow of creating Skl MHI

2. Lightweight 2D-DCNN Architecture

Because the input is a compact 62x62 image, the neural network doesn't need to be deep. The proposed architecture uses:

  • 3 Convolutional Layers: Utilizing Gabor and Gaussian filters for feature extraction.
  • Max Pooling: For spatial invariance.
  • Soft-max Output: To classify the 10 specific actions (e.g., A1: Bending, A6: Waving).

Experiments & Results: Precision Meets Efficiency

The model was tested on a self-collected dataset featuring 10 actions performed by 6 different individuals.

SOTA Comparison and Accuracy

The system showed remarkable convergence speed. It hit 94.52% accuracy in just 10 epochs. When pushed to 1000 epochs, the accuracy stabilized at 97.82%.

Analyzing the overall accuracy depending on training epoch

As shown in the confusion matrix, actions like "Standing" (A4), "Pointing" (A9), and "Lying" (A10) achieved 100% accuracy, demonstrating that the Skl MHI captures unique geometric signatures for these postures perfectly. Some confusion was noted between "Bending" (A1) and "Taking Medicine" (A5) due to similar torso inclinations.

Sample Skl MHI of 10 actions

Critical Analysis & Conclusion

Takeaway

The beauty of this research is its simplicity. By converting a temporal problem into a spatial one through Skl MHI, the authors circumvent the need for expensive Recurrent Neural Networks (RNNs) or 3D-CNNs. This makes the system ideal for deployment on edge devices with limited processing power.

Limitations & Future Work

The current approach relies on a fixed viewpoint. If the person turns 90 degrees, the 2D "silhouette" of the skeleton changes significantly. The authors acknowledge this and plan to investigate view-invariant environments and more diverse action sets in future iterations. To truly move HAR forward, integrating this Skl MHI approach with Graph Convolutional Networks (GCNs) could further exploit the topological structure of the human skeleton.


Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Skeleton Motion History Images (Skl MHI) combined with Transformer-based architectures instead of traditional CNNs for action recognition.
  • Who first introduced the concept of Motion History Images (MHI) for video analysis, and how has the transition from pixel-level MHI to skeleton-level MHI improved robustness against lighting and background changes?
  • What are the state-of-the-art methods for cross-view human action recognition that use skeleton data to handle viewpoint variations more effectively than the Microsoft Kinect-based approach?
Contents
Skl-MHI: Efficient Human Action Recognition through Skeleton Temporal Compression
1. TL;DR
2. Problem & Motivation: The High Cost of Action Recognition
3. Methodology: From Skeleton Sequences to Skl MHI
3.1. 1. The Skl MHI Generation
3.2. 2. Lightweight 2D-DCNN Architecture
4. Experiments & Results: Precision Meets Efficiency
4.1. SOTA Comparison and Accuracy
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work