Structuring the Chaos: Multi-granular Fusion for Social Intelligence
Multi-granular fusion for social data analysis for a decision and intelligence application
The paper presents a multi-granular fusion framework based on Conceptual Graphs to structure and homogenize "soft data" from social networks. By leveraging semantic technologies and specialized fusion strategies for persons and organizational hierarchies, the method transforms noisy, unstructured social information into reliable data for Social Network Analysis (SNA).
TL;DR
Social media data is a goldmine for decision support, but it is notoriously messy. This paper by Thales Research & Technologies introduces a graph-based fusion framework that cleans and structures "soft data" (human-generated text and links). By treating entities as nodes in a Conceptual Graph and applying specialized reconciliation strategies for people and organizations, the authors transform fragmented social noise into a structured network suitable for high-level Social Network Analysis (SNA).
Background: Hard Data vs. Soft Data
In the world of information fusion, "hard data" comes from physical sensors (radar, cameras) and is typically quantitative. "Soft data," however, is produced by humans. It is qualitative, subjective, and full of "imperfections"—think of the difference between a GPS coordinate and a tweet saying "near the main gate."
The authors argue that traditional SNA collapses when faced with the nicknames, misspellings, and lack of semantic annotation common in social software. To bridge this gap, they move beyond simple data cleaning toward Semantic Coordination.
Methodology: The Power of Conceptual Graphs
The core of the methodology lies in representing social data as Conceptual Graphs (CGs). A CG is a bipartite graph consisting of:
- Concept Nodes: Representing entities (e.g., a Company, a Person).
- Relation Nodes: Describing the links between concepts (e.g., "collaborates with," "is member of").
The Fusion Strategy
Fusion isn't just about merging identical items; it’s about finding isomorphisms between graphs. The authors define a compatibility threshold: two nodes can be fused if their types share a common subtype in a hierarchy and their values pass a similarity test (like Levenshtein distance for strings).

Managing Hierarchical "Partners"
One of the most innovative parts of the paper is the Partner Fusion Strategy. In a social network of organizations, information is often partial. One user might list Thales, while another lists Thales/TRT/Reasoning Lab.
The algorithm treats these as partial hierarchies. It calculates similarity based on:
- Acronym Detection: Recognizing that "MG" might mean "MyGroup."
- Level Inclusion: Checking if the most specific sub-entity of one hierarchy exists in the other.
- Structure Merging: Combining
Company/Group/TeamandCompany/Division/Teaminto a single, enrichedCompany/Group/Division/Teamstructure.
Experiments and Performance
The system was implemented in the InSyTo platform and tested on a real-world scientific intelligence dataset containing over 11,000 nodes.

Key Results:
- Network Cleaning: 10 minutes for 11k+ nodes (efficient for offline bootstrapping).
- Pattern Matching: Finding similar entities took only 6 seconds.
- Redundancy Check: Adding new, partially redundant collaborations took 7 seconds.
By reducing the sheer number of redundant nodes, the fusion platform allows decision-making heuristics to run faster and on higher-quality data.
Critical Analysis & Conclusion
While the paper provides a robust framework for structural homogenization, it explicitly leaves out "deliberately misleading information" (disinformation). In today's social media landscape, distinguishing between a "misspelling" and a "bot-driven falsehood" is a critical challenge.
Future Outlook: The authors suggest incorporating Uncertainty Management using Transferable Belief Model Theory. This is a vital next step—if the system can not only fuse data but also provide a "confidence score" for the fused entity, it becomes infinitely more useful for sensitive applications like counter-terrorism or scientific intelligence.
Takeaway: Effective social data analysis requires moving from viewing data as "strings" to viewing it as "evolving semantic structures." This graph-based approach provides the roadmap for that transition.
