Uncovering the Shadow Organization: Mining Social Networks from Process Data
Extraction and analysis social networks from process data
The paper presents a framework for extracting social networks from enterprise process logs (specifically SAP ERP/DMS) and applying local community detection algorithms to uncover hidden organizational structures. By transforming process data into an affiliation network and then a social network, the authors successfully identified functional teams and cooperation patterns that are not explicitly defined in official organizational charts.
TL;DR
While Business Process Management (BPM) usually tells us what activity happens and when, it rarely explains who really works with whom. This paper bridges that gap by extracting social networks from SAP enterprise logs and using local expansion algorithms to find functional communities. The result? A clear map of the informal collaborative structures that actually drive business efficiency.
Background: Beyond the Flowchart
In the world of Process Mining, we are great at drawing flowcharts. We know the path an invoice takes from creation to posting. However, the Human Factor remains a black box. Traditional organizational charts are static and often disconnected from the daily "hustle" of cross-departmental collaboration. The authors argue that by looking at "joint cases"—i.e., people who frequently work on the same objects—we can uncover a "shadow organization" that is vital for process optimization and knowledge management.
Motivation: The Hidden Connections
The core problem identified is that quantitative parameters (like resource load) are easy to extract, but relational parameters (who is the "bridge" between two plants?) are hidden. The authors believe that if two people consistently process the same invoices in an SAP system, they share a functional bond that should be recognized when planning for capacity or system updates.
Methodology: From Logs to Communities
The authors propose a structured pipeline to turn raw transaction logs into actionable social insights:
1. The Two-Mode Transformation
The process starts with an Affiliation Network (Two-mode).
- Vertices: Actors (People) and Entities (Invoices).
- Edges: A link exists if an actor processed a specific invoice.
- Weighting: Determined by the frequency of activities performed by an actor on that object.
2. Social Network Projection
To focus on human relationships, the two-mode network is projected into a One-mode Social Network. The strength of the tie between two people is calculated using Cosine Similarity of their activity vectors across the set of all invoices.
Figure 1: Transition from a two-mode affiliation network (left) to a social network (right).
3. Local Community Detection
Instead of using global clustering, the authors use a Local Expansion Algorithm. This is a sophisticated choice because real-world social groups are often overlapping and nested. A person can be part of a "Logistics" community while also being the bridge to the "Accounting" community.
Experiments: The SAP Invoice Use Case
The researchers tested their approach on an invoice verification process within a real SAP ERP environment.
- Data Scale: ~15,000 invoices, 130 users, and over 70,000 activities.
- Clustering: They identified 16 communities.
Figure 2: Visualization of the 16 detected communities. Circle size represents activity frequency, while edge thickness indicates tie strength.
Key Findings from the Communities:
- The Power Users: Large nodes in the center identified key accountants who acted as the "heart" of the invoice process.
- Functional Pairs (A1): Identified an accountant and a buyer in a specific small plant who worked almost exclusively together.
- Articulate Actors (A2/B3): Identified individuals who acted as "bridges" between different company codes or logistics managers.
- Resundancy Discovery (A4): Found a community where one member had exactly the same authorization and behavior as the others, confirming their role as a "spare" or backup.
Critical Analysis & Conclusion
Takeaway
This paper provides a robust mathematical framework ( and Cosine Similarity) to turn dry IT logs into a vibrant map of human interaction. For an enterprise, this means knowing which employees are critical "nodes" whose absence would break the process chain.
Limitations
The current approach relies heavily on the choice of the similarity function () and weight function (). Different business contexts (e.g., creative collaborative work vs. repetitive transactional work) might require different mathematical definitions of "similarity" to be accurate.
Future Outlook
The authors plan to generalize these community "archetypes" across different industries and refine the functions to suit varying process contexts. In the future, this could lead to AI-driven organizational design where systems automatically suggest the most efficient team groupings based on real-time log data.
Summary Table of Community Interpretations
| Community ID | Real-world Interpretation |
|---|---|
| A1 | Small plant "conjoint" processing duo |
| A2 | Articulate actor bridging two separate plants |
| A4/B1 | Redundancy/Back-up role identification |
| B3 | Multi-plant logistical bridge |
