Consolidarity: Mining the "Social DNA" of Your File Directory

Exploring paerns of social commonality among file directories at work

2007-04-29
John Tang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper explores "social commonality" by analyzing file storage patterns among employees in an enterprise setting. It introduces LiveWire, a deduplication-based backup system, and Consolidarity, a framework that leverages shared files (identified via MD5 hashes and "fingerprinting") to infer social networks and expertise without requiring manual user tagging.

TL;DR

What if your hard drive knew more about your professional network than your LinkedIn profile? This paper explores how the files we store at work—often considered isolated personal data—actually reveal a dense web of "social commonality." By analyzing file redundancy, the authors built systems to optimize enterprise backups (LiveWire) and map hidden social connections (Consolidarity) without asking users to do a single lick of extra work.

The "Workplace Intranet" Problem

We’ve all been there: searching the corporate intranet for a specific report and finding nothing, even though you know someone in the building has it. Unlike the World Wide Web, which thrives on PageRank and cross-linking, corporate data is trapped in "personal" silos.

The authors argue that we shouldn't ask employees to manually organize or tag information—they won't. Instead, we should look at what they are already doing: storing and managing files. Every identical PDF, software library, or PowerPoint deck shared between two people is a "social transaction" waiting to be analyzed.

Methodology: Finding the Needle in the Generic Haystack

If you compare two work computers, you'll find thousands of matches. Most of them are boring: kernel32.dll, system drivers, or the standard company wallpaper. To find meaningful commonality, the researchers used a multi-stage filter:

  1. Noise Reduction: They subtracted files found on a "base machine" (a fresh OS install) and used a TF-IDF filter to ignore any file present on more than 70% of machines.
  2. Fingerprinting: To catch "similar" but not identical files (like different versions of a doc), they broke files into 32KB chunks and hashed them. If 20% of the chunks matched, the files were considered "semantically similar."
  3. Privacy-First Hashing: They used MD5 hashes of file content rather than reading the text itself, ensuring researchers didn't actually "read" private documents.

Concept sketch of reflecting commonality patterns back to users Figure: A conceptual UI showing how a user might see "Tag Clouds" of their expertise and the "colored" people they share commonalities with.

Key Insights: We Are More Redundant Than We Think

The study found that the average user has 25% redundancy on their own drive, but across the organization, that jumps to 54%.

More interestingly, they discovered that file commonality often maps to roles rather than just projects. Software engineers share specific libraries; managers share specific templates. This "Consolidarity" system can identify an "information broker"—someone who might not be your direct teammate but has used the exact same obscure developer tool you’re currently struggling with.

Visualization of Productivity and Developer filetypes Figure: The social graph revealed by shared files. Note how some users act as hubs, connecting different clusters through shared digital artifacts.

The "Temporal Pollution" Warning

The researchers encountered a fascinating "meta-data" hurdle. They wanted to use the "Last Accessed" date to see what users were currently interested in. However, they found the data was "polluted." Every time a virus scanner, backup tool, or desktop search indexer touched a file, it reset that date.

Takeaway for Devs: If you're building systems that rely on usage logs, respect the meta-data! If your tool "touches" a file for background maintenance, preserve the original access timestamp, or you'll destroy the signal for future AI/mining tools.

Conclusion and Future Outlook

This 2007 work was a precursor to modern "Graph" technologies (like Microsoft Graph). It proves that the "social DNA" of an organization is already written into its file systems. The challenge isn't getting people to share—it's building the intelligent, privacy-aware filters that can distinguish between a shared operating system and a shared expertise.

Future iterations of this could lead to "transitive discovery"—if you and I share File A, and you also have File B in the same folder, maybe File B is exactly what I'm looking for.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use data deduplication or file fingerprinting techniques for organizational expertise mapping or social network analysis.
  • Who first proposed the use of "passive sensing" or "implicit interaction" to build enterprise knowledge management systems, and how has that field evolved since CHI 2007?
  • What are the modern industry standards or frameworks for balancing employee privacy with "enterprise search" data mining in the age of GDPR?
Contents
Consolidarity: Mining the "Social DNA" of Your File Directory
1. TL;DR
2. The "Workplace Intranet" Problem
3. Methodology: Finding the Needle in the Generic Haystack
4. Key Insights: We Are More Redundant Than We Think
5. The "Temporal Pollution" Warning
6. Conclusion and Future Outlook