Multimodal Gait Recognition: Beyond the Visual Silhouette

Multimodal features fusion for gait, gender and shoes recognition

2016-05-06
F. M. Castro, M. Marín-Jiménez, N. Guil
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a multimodal fusion framework for gait-based person identification, gender recognition, and shoe-type classification. By integrating RGB, depth, and audio features using techniques like Fisher Vectors and Rank Minimization, the authors achieve SOTA results on the TUM GAID and CASIA-B datasets, including 100% identification accuracy in challenging temporal scenarios.

TL;DR

Gait recognition—identifying individuals by the way they walk—has long been haunted by the "clothing and carrying" problem. This paper breaks the reliance on 2D silhouettes by fusing RGB, Depth, and Audio information. By leveraging tracklet-based DCS features and sophisticated fusion strategies like Rank Minimization, the authors achieve a flawless 100% accuracy on the TUM GAID dataset, even when subjects change shoes or wear heavy winter coats.

Background & Motivation

Most gait recognition systems are built on Gait Energy Images (GEI)—essentially a temporal average of binary silhouettes. While elegant, GEIs fail when a person puts on a bulky jacket or carries a bag, as the silhouette's geometry changes fundamentally.

The authors' central insight is that walking isn't just a visual shape; it is a spatio-temporal motion pattern that produces distinct depth maps and unique acoustic signatures (the sound of footsteps). By capturing the "physics" of the walk through dense trajectories and the "acoustics" through spectral features, the system becomes significantly more resilient to appearance changes.

Methodology: The Multimodal Pipeline

The proposed system follows a structured four-stage pipeline:

  1. Low-level Feature Extraction:
    • Visual & Depth: Uses "tracklets" (short dense trajectories). It employs the DCS (Divergence-Curl-Shear) descriptor, which captures local expansion, rotation, and deformation of the walking person.
    • Audio: Captures MFCC and Mel-frequency spectra from footstep recordings.
  2. Mid-level Representation: These features are encoded into Fisher Vectors (FV). Unlike simple Bag-of-Words, FV captures the distribution of descriptors relative to a Gaussian Mixture Model (GMM), providing a much richer signature.
  3. Information Fusion: The heart of the paper. The authors compare Early Fusion (e.g., Multiple Kernel Learning - MKL) where features are mixed before classification, and Late Fusion (e.g., Rank Minimization) where individual scores are optimized post-classification.
  4. Classification: One-vs-all linear SVMs identify the subject.

Overall Pipeline Architecture Figure 1: The multimodal pipeline showing the path from raw Audio/RGB/Depth data to the final Identity assignment.

Key Results & Breakthroughs

1. The Power of "Visual + Depth"

The experiment on the TUM GAID dataset (which records people in different months wearing different clothes) was a "stress test." While standard visual descriptors struggle with "Elapsed Time" (TS) scenarios, the combination of DCS-Visual and DCS-Depth reached 100% accuracy.

2. Shoe and Gender Recognition

The study extended gait features to "Soft Biometrics":

  • Gender: Achieved >97% accuracy by combining RGB tracklets and Depth.
  • Shoes: Interestingly, Audio played its biggest role here. The acoustic signature of a high boot vs. a sneaker is more distinctive than its visual motion, leading to a significant performance boost when audio was added to the fusion mix.

Performance Comparison Table Table 1: Quantitative results showing our DCS-Visual and Depth methods outperforming traditional GEI and GEV approaches.

3. Rank Minimization: The "Secret Sauce"

One of the paper's most impressive technical contributions is the application of Rank Minimization (RM). RM treats the classification scores from various models as a matrix that should be "low-rank" (consistent). By removing noise and enforcing consistency across different views or modalities, RM boosted the accuracy of even single-modality systems by 3-28%.

Critical Insights

  • Why DCS? Unlike silhouettes, DCS descriptors are derived from motion trajectories. They are less focused on "what the person looks like" and more on "how the person moves," making them inherently more robust to clothing.
  • The Depth Advantage: Depth sensors (like Kinect) bypass color/lighting issues, providing a "clean" 3D silhouette that encodes the vigor and stride of the gait more accurately than 2D pixels.
  • The Audio Limitation: While audio helps identify shoe types, it is the weakest modality for identity. However, as a supplementary signal, it provides "orthogonal" information that RGB simply doesn't have.

Conclusion

This work signals a shift from "Gait as a Shape" to "Gait as a Multimodal Experience." The achieving of 100% accuracy on temporal datasets suggests that the future of biometric surveillance lies in sensors that can "see" in 3D and "hear" the environment. While the computational cost of dense trajectories is higher than simple silhouettes, the gains in robustness make it a viable path for high-security applications.

Future Work: Adapting this multi-modal approach to "in-the-wild" scenarios where audio might be noisy and depth sensors have limited range remains the next frontier.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Deep Learning based multimodal fusion (RGB-D-Audio) specifically for long-term gait verification in the wild.
  • Which paper first proposed the Rank Minimization method for score fusion in video event detection, and how has it been adapted for biometric identification?
  • Explore the application of Divergence-Curl-Shear (DCS) trajectories in modern action recognition tasks beyond human gait analysis.
Contents
Multimodal Gait Recognition: Beyond the Visual Silhouette
1. TL;DR
2. Background & Motivation
3. Methodology: The Multimodal Pipeline
4. Key Results & Breakthroughs
4.1. 1. The Power of "Visual + Depth"
4.2. 2. Shoe and Gender Recognition
4.3. 3. Rank Minimization: The "Secret Sauce"
5. Critical Insights
6. Conclusion