Redundancy Avoidance in Entity Resolution: A Social Network Perspective
Redundancy Avoidance in Entity Resolution Based On Social Networks Paradigm
The paper introduces a seven-phase framework for Entity Resolution (ER) that reimagines datasets as Social Network Analysis (SNA) graphs, treating records as nodes and relationships as edges. By leveraging SNA centrality measures alongside traditional deterministic matching algorithms like Levenshtein and Jaro-Winkler, the method effectively identifies and merges redundant real-world entities.
TL;DR
Entity Resolution (ER) is the process of identifying duplicate records that refer to the same real-world entity. This paper proposes a novel framework that transforms standard tabular data into a Social Network graph, applying Social Network Analysis (SNA) centrality measures to improve matching accuracy and handle redundancy. By benchmarking algorithms like Levenshtein and Jaro-Winkler within this paradigm, the authors prove that relational context is a powerful tool for Big Data integration.
Problem & Motivation: The Identity Crisis in Big Data
In an ideal world, every data record would have a unique, universal ID (like an SSN). In reality, datasets are riddled with "dirty" data: misspelled names, transposed birthdates, and missing fields.
Existing Deterministic (rule-based) and Probabilistic approaches often treat records as isolated rows. The authors argue that this ignores the Relational Similarity. For instance, if "Jane Aha" and "Jane Lindsay" share the same household or parents, they are likely the same person despite name variations. The limitation of prior work lies in the high computational cost of "Collective Entity Resolution" and a failure to visualize the latent structure of record relationships.
Methodology: From Rows to Relationships
The core innovation is a seven-phase framework that converts data into a graph of nodes and edges.
1. The SNA Transformation
Records are treated as Actors (Nodes) and their attributes or known associations (like family ties) as Links (Edges). By calculating Centrality Measures (Degree, Betweenness, Closeness), the framework identifies "influential" nodes that act as anchors for merging duplicates.
2. The Stepwise Deterministic Strategy (SDS)
Instead of a single-pass comparison, the framework uses SDS. If a pair doesn't meet the first set of strict criteria, it is passed to a secondary round with different identifiers. This ensures that even records with significant noise are eventually captured.

Experiments & Results: Precision vs. Performance
The framework was tested on the RLdata10000 dataset across four core algorithms: Levenshtein, Jaro-Winkler, Soundex, and K-means clustering.
Key Findings:
- The Accuracy King: Both Levenshtein and Soundex achieved 100% precision.
- The Scalability Winner: Jaro-Winkler proved to be the most efficient for large-scale data (2 million records), maintaining 99.999% precision while being significantly faster than Levenshtein.
- The K-means Failure: K-means performed the worst (91.33% precision), suggesting that simple clustering without string-similarity nuance is insufficient for ER.

Critical Insight & Conclusion
The study highlights a vital trade-off in technical implementations: Accuracy vs. Latency. While Levenshtein provides an "optimal" result, the near-optimal performance of Jaro-Winkler makes it the superior choice for "Massive Big Data" where execution time is a bottleneck.
Future Outlook: The transition from flat-file processing to graph-based analysis (SNA) is a significant shift. By treating data as a network, we move closer to "Human-like" reasoning—using the context of who a person knows or where they live to confirm their identity, rather than just how their name is spelled.
Takeaways for Engineers:
- Context Matters: If your ER task is failing, consider building a relationship graph between your records.
- Algorithmic Choice: Use Jaro-Winkler for high-volume streaming data; reserve Levenshtein for high-stakes, lower-volume master data management.
