CNN-Based Indoor Localization: Turning Crowdsourced Wi-Fi into High-Precision "Images"
A shop-level location algorithm based on CNN for crowdsourcing fingerprint
This paper introduces a shop-level indoor positioning algorithm that leverages crowdsourced Wi-Fi fingerprints and a Multi-CNN architecture. By transforming Wi-Fi signal statistics into 2D feature maps, the model achieves a SOTA accuracy of 91% in identifying specific retail stores within large commercial complexes.
TL;DR
Researchers have developed a shop-level indoor positioning system that bypasses expensive site surveys by using crowdsourced transaction data. By treating Wi-Fi signal statistics as 2D "images," a Multi-CNN architecture achieves 91.01% accuracy, outperforming traditional boosters like XGBoost while significantly reducing maintenance costs.
The Localization Dilemma: Accuracy vs. Maintenance
Indoor positioning is the "holy grail" for Location-Based Services (LBS). While GPS fails inside heavy mall structures, Wi-Fi fingerprinting has emerged as a viable alternative. However, the industry faces two major hurdles:
- High Cost: Building a Radio Map through manual surveying is labor-intensive.
- Signal Instability: Wi-Fi RSSI (Received Signal Strength Indication) fluctuates wildly due to human traffic, obstacles, and time of day.
The authors' insight? Use the "crowdsourced" data generated by mobile payments. Every time a user pays via phone, a snapshot of nearby BSSIDs (Wi-Fi access points) is captured.
Methodology: From Signals to Feature Maps
The core innovation lies in how the researchers bridge the gap between abstract signal values and visual-processing power of CNNs.
1. Constructing the Feature Matrix
Instead of feeding raw RSSI values, the team uses a sliding window (from 1 to 23 days) to generate statistical features:
- Ratio Features: The frequency of a BSSID appearing in a specific shop compared to the total mall.
- Deviation Features: Counting BSSIDs within an 8dBm strength fluctuation—a clever trick to account for signal attenuation.
- Distance Features: Calculating the intensity difference between current readings and historical shop averages.

2. Multi-CNN Joint Training
Rather than one giant network, the architecture uses a Joint Training approach. Different feature groups (Ratio, Deviation, Distance) are fed into separate CNN branches. This allows each branch to learn the unique spatial patterns of that feature type before they are flattened and "fused" with manual features like latitude and longitude coordinates.

Experimental Battle: CNN vs. The Boosters
The model was tested against industry-standard classifiers using 2017 CCF BDCI competition data (100+ stores).
| Model | Accuracy (Score) | Training Time |
|---|---|---|
| Logistic Regression | 82.06% | 1h 41min |
| AdaBoost | 89.08% | 4h 22min |
| XGBoost | 90.52% | 1h 49min |
| Proposed CNN | 91.01% | 1h 10min |
The CNN didn't just win on accuracy; it was significantly faster to train than AdaBoost and slightly faster than XGBoost, proving that its "deep" understanding of Wi-Fi correlations is more efficient than shallow tree-splitting.
Preventing Overfitting
One challenge with Wi-Fi data is noise. The authors utilized Early Stopping and Batch Normalization (BN). As seen in the training curves, the validation accuracy begins to fluctuate after 12 epochs, marking the optimal point to stop before the model memorizes the noise (overfitting).

Depth Insight & Conclusion
This paper proves that the "Computer Vision" approach to signal processing is highly effective for LBS. By treating temporal signal statistics as spatial pixels, CNNs can "see" a shop's fingerprint even when individual Wi-Fi signals are missing or fluctuating.
Takeaway: For developers in the LBS space, the future lies in crowdsourced automation. Moving away from manual surveys to self-learning CNN models that utilize transaction metadata can slash deployment costs while keeping precision high enough for hyper-local targeted advertising and mall analytics.
Limitations: The model currently treats each shopping mall as an independent model. Future iterations could benefit from a unified model that transfers knowledge between different mall environments.
