Bridging the Gap: Transforming Relational Data into Enterprise Social Networks
Towards Social Network Extraction Using a Graph Database
The paper proposes a novel framework for extracting Social Networks (SN) from enterprise Relational Databases (RD) by converting them into Graph Databases (GD). Specifically, it utilizes an extended Hypernode Model to represent complex administrative relationships, facilitating high-level graph-based querying for social structure discovery.
TL;DR
Enterprises sit on goldmines of social data locked within Relational Databases (RD). This paper introduces a robust methodology to unlock this potential by migrating RD structures into a Hypernode-based Graph Database (GD). This transition allows for more natural social network modeling, enabling businesses to identify expertise, leadership, and hidden professional connections that traditional SQL queries struggle to reveal.
Background & Motivation: Why SQL Fails Social Analysis
In the business context, crucial information—who knows whom, who is an expert in a specific domain, and how departments interact—is typically buried in sales records, project tables, and contact lists.
The authors argue that Relational Databases are fundamentally "graph-blind." Representing evolving social structures requires expensive schema renormalization and complex joins. Furthermore, while web-based extraction (like LinkedIn or FOAF) is popular, it often lacks the high-fidelity, trusted data found within an organization's internal SQL servers. The paper identifies a missing link: a formal way to move from the rigid rows of an RD to the fluid nodes of a Social Network.
Methodology: The Hypernode Advantage
The core innovation lies in choosing the Hypernode Model over simple node-edge graphs. A hypernode is a directed graph where nodes can themselves be graphs, allowing for the encapsulation of complex objects (e.g., a "Person" node containing "Education" and "Project" sub-nodes).
1. Schema Translation
The system extracts metadata from the RD to identify Primary Keys (PK) and Foreign Keys (FK). It then applies mapping rules:
- Tables become Hypernodes.
- FK relationships are turned into directed edges.
- Specialized mappings: "IS-A" (inheritance) and "Part-Of" (composition) relationships are inferred from composite keys and shared primary keys.
Figure 1: The resulting Hypernode Database Schema showing the mapping from PhD thesis data.
2. Data Conversion
The second phase unloads relational tuples and restructures them to populate the graph. This ensures that every entry in the RD is represented as a specific instance within the Hypernode framework, facilitating high-speed traversal without the overhead of massive table joins.
Experiments & Validation
The authors developed a prototype using Java and PostgreSQL, visualizing results with JGraph. They tested the system on the ADEME database (a repository of PhD student data).
To prove the method's correctness, they executed parallel queries on both the source RD and the new HD.
- Query Example: Finding the names of students who finished their thesis.
- Result: The HD returned identical results to the RD, confirming that the transformation is lossless and non-redundant.
Figure 2: The prototype interface used to validate the transformation logic.
Critical Insight: The Road to a Living Social Network
By moving data into a graph format, the "social" nature of the data becomes explicit. For instance, the system can automatically generate a "works_on_same_topic" relation between researchers in different labs who were previously isolated in separate rows of a relational table.
Figure 3: The final social network layer, where nodes represent people and edges represent extracted professional relationships.
Summary of Takeaways
- Structural Flexibility: Hypernodes handle the "messy" and hierarchical nature of human relationships better than flat graph models (like GOOD or GMOD).
- Integrity: The mapping maintains 1NF (First Normal Form) integrity during the transition.
- Future Potential: This framework sets the stage for advanced "Expertise Search" engines within corporations.
Conclusion & Extensions
The paper successfully demonstrates that graph conversion is not just a format change but a semantic upgrade. While the prototype is solid, future work could integrate Natural Language Processing (NLP) to disambiguate identical names and pull even more context from unstructured fields within the database. For any organization looking to implement a "Knowledge Graph," this methodology provides a clear, mathematically grounded roadmap.
