Eaves: Leveraging Acoustic Intelligence for Non-Intrusive Social Distancing

4757_Eaves An IoT-Based Acoustic Social Distancing Assistant for Pandemic-Like Situations.

Summary
Problem
Method
Results
Takeaways

The paper introduces "Eaves," an IoT-based acoustic social distancing assistant that uses ambient audio to estimate the distance of nearby individuals. By extracting human-voice-centric MFCC features and employing a Random Forest classifier, the system achieves nearly 100% prediction accuracy within the critical 2-meter social distancing range without requiring Line-of-Sight (LoS) or specialized apps on third-party devices.

TL;DR

The Eaves system is a breakthrough in pandemic-responsive technology that uses human voice acoustics rather than cameras or Bluetooth to enforce social distancing. By analyzing the "fingerprint" of sound attenuation through MFCC features, it achieves near-perfect accuracy within 2 meters, providing a privacy-friendly, hardware-agnostic solution for public safety.

Academic Positioning: This work sits at the intersection of IoT Signal Processing and Applied Machine Learning, offering a robust alternative to SOTA vision-based systems by eliminating the Line-of-Sight (LoS) requirement.

Problem & Motivation: The Limitations of "Seeing" Distancing

During the COVID-19 pandemic, two dominant technologies emerged for monitoring social cycles:

  1. Computer Vision: Effective but plagued by occlusion, high computational costs, and significant privacy backlash.
  2. Bluetooth/RSS-based Apps: Limited by the "two-to-tango" problem—they only work if everyone has the app installed and active.

The authors observed a fundamental physical intuition: Virus transmission and sound transmission often share the same physical paths. If an obstruction blocks sound, it likely blocks viral droplets. Therefore, audio serves as a natural proxy for physical proximity risks.

Methodology: From Sound Waves to Distance Metrics

The core of Eaves lies in its specialized signal processing pipeline. Unlike general audio classifiers, Eaves focuses strictly on human-voice-centric components.

1. Feature Extraction (MFCC)

The system filters audio to isolate frequencies between 85-155 Hz (male) and 165-255 Hz (female). It then calculates Mel-frequency cepstral coefficients (MFCC). The logic is that the power and distribution of these coefficients change predictably as the sound source moves away from the microphone.

2. The IoT Architecture

To handle resource-constrained devices, the paper proposes a flexible hierarchy:

  • Independent Mode: High-end devices (laptops/phones) process data locally.
  • Offloading Mode: Low-power IoT nodes stream audio to Fog nodes to minimize network latency compared to centralized Cloud processing.

Overall Signal Processing Flow Figure 1: The information flow showing audio capture, noise removal, and distance prediction.

Experiments & Results: Accuracy vs. Latency

The researchers compared four distinct ML models: Linear SVC, Decision Trees, Random Forest, and CNNs.

Performance Metrics

  • Random Forest emerged as the winner, balancing a high 97.73% accuracy with a rapid 19.6 ms delay.
  • CNNs, while powerful, suffered from higher latency (53.5 ms) due to the complexity of neural layers, making them less ideal for real-time alerting on edge devices.
ModelAccuracy (%)Delay (ms)
Linear SVC61.82917
Decision Tree81.820.47
Random Forest97.7319.6
CNN87.5053.5

Audio Spectrogram Analysis Figure 2: Visualization of Raw, Noisy, and Filtered human voice signals used for training.

Deployment Insights

The confusion matrix revealed that the system is exceptionally reliable at distances of 0.5m, 1.0m, and 2.0m. While accuracy slightly degrades at 2.5m, the system remains perfect for its primary goal: detecting a 2-meter breach.

Critical Analysis & Conclusion

Takeaway

Eaves proves that acoustic sensing is a viable "third way" for public health monitoring. It bypasses the infrastructure requirements of Bluetooth and the privacy/LoS concerns of Vision.

Limitations

  • The "Silent" Problem: The method relies on the presence of human voice. If individuals are silent, the system cannot estimate distance.
  • Battery Drain: Continuous microphone monitoring and FFT processing are energy-intensive for small IoT nodes.

Future Outlook

The authors suggest that future iterations could incorporate "non-human" sounds (like footsteps or clothing rustle) to maintain tracking when speech stops. This work opens the door for Audio-as-a-Service in smart cities, where existing microphone infrastructure (in streetlights or kiosks) could be repurposed for public safety without invasive surveillance.

Find Similar Papers

Try Our Examples

  • Find recent papers that use acoustic signal processing and deep learning for social distancing or occupancy detection in indoor environments.
  • What are the primary theoretical foundations of the Direct-to-Reverberant Energy Ratio (DRR) for speaker distance estimation, and how does the MFCC approach in Eaves compare?
  • Explore how audio-based distance estimation methods have been integrated into multi-modal fusion networks for autonomous mobile robot safety.
Contents
Eaves: Leveraging Acoustic Intelligence for Non-Intrusive Social Distancing
1. TL;DR
2. Problem & Motivation: The Limitations of "Seeing" Distancing
3. Methodology: From Sound Waves to Distance Metrics
3.1. 1. Feature Extraction (MFCC)
3.2. 2. The IoT Architecture
4. Experiments & Results: Accuracy vs. Latency
4.1. Performance Metrics
4.2. Deployment Insights
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook