FloorSense: Evolving Indoor Maps from Geometry to Semantics via CRF
FloorSense: a novel crowdsourcing map construction algorithm based on conditional random field
FloorSense is a crowdsourcing-based indoor map construction framework that utilizes Smartphone sensors (IMU, Barometer, Magnetometer) and Conditional Random Fields (CRF) to generate semantic indoor maps. Beyond typical layout reconstruction, it achieves high-accuracy semantic labeling of functional areas like clothing stores and restrooms in complex mall environments.
TL;DR
Researchers have moved beyond simple "skeleton" maps of buildings. FloorSense introduces a novel crowdsourcing framework that doesn't just draw the walls of a mall—it identifies if a room is a clothing store, a restroom, or a cashier. By using Conditional Random Fields (CRF) and smartphone sensor data (IMU/Barometer), it transforms raw walking traces into a rich semantic map with an average accuracy of 2.1 meters.
Problem & Motivation: The "Blind" Grammar Map
Most modern indoor mapping solutions (like those from Google or MazeMap) provide a "grammar map"—a layout showing where walls and corridors are. However, they lack Semantic Intelligence. They don't know the function of the space.
Existing crowdsourcing methods (like Zee or CrowdInside) rely on Hidden Markov Models (HMM) to predict user movement. The problem? HMM assumes that your current activity is only dependent on your current state, ignoring the context of your entire journey. Furthermore, traditional SLAM (Simultaneous Localization and Mapping) requires expensive robots or specialized hardware, making large-scale mall mapping a logistical nightmare.
Methodology: The Three Pillars of FloorSense
The system operates through a sophisticated top-down pipeline that converts kinetic energy (walking) into spatial data.
1. High-Precision PDR (Pedestrian Dead Reckoning)
To track users accurately, FloorSense utilizes a dynamic thresholding formula for step detection: This adaptation allows the system to maintain less than 1% error in distance estimation regardless of the user's walking pace.
2. Geometry Reconstruction (The Grammar Map)
By collecting thousands of "traces" (user paths), the system uses an Alpha-shape algorithm. Imagine a circle rolling along the outer edges of a cloud of points; the lines it traces form the boundary of the reachable space.
Figure: The process of breaking traces into segments and clustering them to define corridors vs. rooms.
3. CRF Semantic Inference (The Core "Brain")
This is where the magic happens. Unlike HMM, the Conditional Random Field (CRF) model allows the system to look at the entire sequence of actions. It uses a 6-dimensional vector for each compartment, including:
- Stop-and-Still (SS) frequency: People stand still longer at cosmetic counters than in restrooms.
- Turning (TN) frequency: Restrooms involve specific, tight turning patterns.
Figure: The procedural integration of grammar maps into a graph-based CRF for semantic labeling.
Experiments & Results: Real-World Mall Testing
The team tested FloorSense in a 10,000 m² shopping mall.
Positioning Accuracy
Compared to state-of-the-art (SOTA) systems like Zee or PiLoc, FloorSense holds its own without needing pre-existing floor plans or expensive Wi-Fi infrastructure.
- Reported Accuracy: 2.1 meters (Average).
- Trace Density: The system reaches peak performance once roughly 300 user traces are collected.
Figure: The correlation between track density and prediction accuracy.
Semantic Precision
| Function | Precision |
|---|---|
| Restroom | 80.00% |
| Clothing Store | 68.03% |
| Cashier | 64.52% |
| Cosmetic Counter | 61.64% |
The high precision for restrooms is a direct result of "continuous turning" patterns being highly distinct in IMU data. Meanwhile, stores are slightly harder to distinguish because "standing and browsing" behavior looks similar across different retail types.
Critical Insight & Conclusion
FloorSense successfully proves that human behavior is a secondary "signal" for mapping. By treating a person's movement as a signature of the room's purpose, we can build maps that are not just blueprints, but functional databases.
Limitations: The system still struggles with "slight turns" that smartphones fail to capture in a pocket, and it assumes users walk naturally. Future work could integrate Wi-Fi RTT (Round Trip Time) or Magnetometer fingerprints to further reduce the 2.1m drift.
Takeaway: The future of LBS (Location-Based Services) isn't just knowing where you are, but what you are doing there.
