Collaid: Decoding the "Who" and "What" of Tabletop Collaboration
Who did what? who said that? Collaid: an environment for capturing traces of collaborative learning at the tabletop
This paper introduces Collaid (Collaborative Learning Aid), a multimodal environment designed to capture high-fidelity traces of collaborative learning at interactive tabletops. By integrating a depth sensor for touch-user mapping and a microphone array for speech attribution, the system enables an automated "Who did what? Who said that?" analysis to support visual feedback and sequential data mining.
TL;DR
Interactive tabletops are great for collaboration, but they are often "blind" to who is actually doing the work. Collaid is a multimodal research environment that uses depth sensors and microphone arrays to track individual touches and speech. It transforms chaotic group interactions into structured data, enabling real-time dashboards and sequence mining to help teachers and learners understand the collaborative process as it happens.
Problem & Motivation: The Identity Crisis of Interactive Surfaces
While digital tabletops encourage face-to-face interaction, they face a fundamental technical hurdle: the lack of user attribution. Most multi-touch hardware treats all contacts as anonymous. If three students are building a concept map, the system sees three fingers, but doesn't know which belongs to the quiet leader and which to the dominant "free-rider."
Furthermore, collaboration isn't just about moving pixels; it's about what is said. Most existing systems ignore verbal tokens, missing the "Why" behind the "What." The authors argue that without capturing both physical actions and verbal participation, we cannot truly model or support the learning process.
Methodology: The Multimodal Sensor Suite
The genius of Collaid lies in its non-intrusive design. Rather than making students wear gloves or use specialized pens, the system uses two overhead "eyes":
- The Depth Sensor (Body Tracking): Positioned above the table, it tracks the skeletons and arm spans of users. When a touch is detected on the LCD, the system uses a greedy search algorithm to map that coordinate back to the closest arm, identifying the "owner" of the action.
- The Microphone Array (Voice Localization): A 7-channel radial array detects the spatial origin of sound. By matching the sound source to the known positions of the players, Collaid knows who is speaking without individual lapel mics.

The software architecture (shown below) funnels this multimodal data into a central repository where it is converted from raw coordinates into semantic logs (e.g., "User A created a link," "User B edited User A's concept").

Experiments & Results: Visualizing the Invisible
To validate the system, the authors conducted a study where groups built concept maps. The captured data was used to generate real-time Collaboration Dashboards.
- Radar of Participation: These blue and red "spider charts" show the symmetry of speech and touch. A balanced triangle indicates egalitarian participation; a skewed shape highlights a dominant user.
- Contribution Charts: Pie charts show exactly how much "knowledge" (concepts/links) each person added to the final product.

The study revealed distinct "group profiles." For instance, some groups were highly communicative (symmetrical blue radars) but physically slow, while others worked efficiently on the screen but in total silence, missing the opportunity for shared sense-making.
Mining the "Patterns of Work"
Beyond mere visualization, the authors applied N-gram sequence mining. They found that successful groups often follow specific behavioral sequences, like "Move Tool -> Add Concept -> Move Concept" (MT-AC-MC), which suggests a structured thinking process rather than random clicking.
Critical Analysis & Conclusion
Takeaway: Collaid moves the field from "Black Box" collaboration to "Glass Box" analytics. By establishing a set of design guidelines—such as distinguishing users and integrating contextual data—the authors provide a roadmap for the next generation of "smart" classrooms.
Limitations:
- Acoustic Noise: In a real classroom, multiple microphone arrays might interfere with each other.
- Occlusion: The depth sensor might struggle if users lean over each other.
- Semantic Depth: While the system knows that someone spoke, the early version described here doesn't perform full Speech-to-Text to analyze the quality of the argument.
Future Outlook: The future of this tech lies in predictive modeling. Imagine a system that detects a "participation slump" in real-time and subtly prompts a quiet student to contribute, or alerts a teacher to a group that has fallen into a "disagreement loop" based on their touch patterns.
