Beyond the Email: A Weighted Multi-Attribute Approach to Social Profile Matching
User profile matching in social networks
The paper introduces a comprehensive framework for matching user profiles across heterogeneous social networks by leveraging the FOAF (Friend Of A Friend) vocabulary. It moves beyond simple Inverse Functional Property (IFP) matching, such as emails, by employing a weighted multi-attribute similarity approach combined with Dempster-Shafer theory for final decision-making.
TL;DR
Most profile matching algorithms fail the moment you use a work email on LinkedIn and a personal one on Facebook. This paper proposes a robust framework that doesn't just look for exact matches in email addresses. Instead, it transforms profiles into a unified FOAF format, applies specialized similarity metrics (syntactic and semantic) to every attribute, and uses Dempster-Shafer evidence theory to decide if two profiles belong to the same human, even when the data is messy or contradictory.
The Motivating Crisis: The "Data Isolated Island"
In the modern web, "Bob the Developer" is a fragmented entity. On Facebook, he is connected to high-school friends; on LinkedIn, he is a software engineer. If an application wants to merge these two "Bobs" to enrich his profile or find common friends, it usually looks for a shared Inverse Functional Property (IFP)—typically a hashed email address.
The Pain Point: Users don't use the same email everywhere. They change addresses, use aliases, or even share proxies. When the IFP fails, the link breaks, creating "Data Isolated Islands." Current SOTA (at the time) was too restrictive, leading to low recall and lost connections.
Methodology: How to Match Without a Key
The authors argue that matching should be global. If two profiles share a name, a similar profile picture, and overlapping interests, it is highly likely they are the same person, even if the emails differ.
1. The FOAF Normalization
To handle heterogeneous schemas (where Facebook calls it "School" and LinkedIn calls it "Education"), the framework uses a FOAF Middleware to map everything to the Friend Of A Friend semantic vocabulary.
2. Specialized Similarity Functions
One size does not fit all. The framework maps attributes to appropriate mathematical metrics:
- Senseless One-term (Names): Jaro metric.
- Multi-terms (Biographies): SoftTFIDF.
- Semantic (Interests): Explicit Semantic Analysis (ESA) using Wikipedia to see if "Coding" and "Software Engineering" are related.
- Numeric/URIs: Edit Distance.

3. Attribute Weighting & Decision Making
Not all attributes are equal. A shared last name is less "identifying" than a shared homepage. The system calculates weights either manually or through an automated algorithm that measures the prevalence of attributes across social sets.
Finally, they employ Dempster-Shafer (DS) theory. Unlike Bayesian Networks, DS theory is excellent at handling "lack of evidence." If an attribute is missing, DS doesn't assume it's a "mismatch"—it correctly identifies it as "uncertainty," leading to much higher precision.
Performance & SOTA Comparison
The authors tested three datasets, varying the percentage of "different" attributes between identical users.
Key Findings:
- Recall Power: When users had different IFPs, the proposed method successfully linked profiles where traditional IFP-based methods found 0 results.
- The Power of Weights: Without weighting, the system produced massive amounts of False Positives. Once weights were applied, precision stabilized (see Figure 6).
- Algorithm Superiority: DS Theory outperformed Bayesian Networks (BN) and simple Averages (Avg) in both precision and recall.

Critical Insight: Why it Works
The brilliance of this work lies in its flexibility regarding uncertainty. By using a middleware (FOAF) and a probabilistic decision engine (DS Theory), the system acknowledges that social data is inherently incomplete. The shift from "Hard Matching" (Email == Email) to "Soft Evidence Fusion" represents the transition from brittle database joins to intelligent identity resolution.
Limitations
- Scalability: Crawling diverse social networks is computationally expensive and often blocked by APIs.
- Privacy: The paper focuses on the utility of matching but doesn't deeply address the privacy implications of linking "isolated" accounts without explicit user consent.
Conclusion
This framework provides a template for data enrichment and cross-network search. By treating identity as a collection of weighted evidence rather than a single unique key, the researchers have built a bridge between the "Data Islands" of the social web.
