iOLAP: Bridging the Gap Between Search and Multi-Dimensional Business Intelligence

iOLAP: A Framework for Analyzing the Internet, Social Networks, and Other Networked Data

2009-03-19
Yun Chi, Shenghuo Zhu, Koji Hino, Yihong Gong, Yi Zhang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces iOLAP, a systematic framework for analyzing multi-dimensional networked data (Internet, social networks, citations) using a Polyadic Factor Model. By employing Non-negative Tensor Factorization (NTF), the framework decomposes data cubes into meaningful, correlated factors across dimensions like People, Relation, Content, and Time.

TL;DR

The iOLAP framework transforms the "wild west" of Internet networked data—blogs, citations, and social links—into a structured multi-dimensional "Data Cube." By utilizing a Polyadic Factor Model based on Non-negative Tensor Factorization, iOLAP allows users to perform "Roll-up" and "Drill-down" operations on non-numerical data like social relations and content topics, significantly outperforming traditional recommendation baselines.

Problem & Motivation: Beyond the Power-Law Search

Current Internet technologies are dominated by search engines optimized for retrieval. However, search engines inherently favor the "head" of the power-law distribution—the most popular and authoritative sites. This leaves the "Long Tail"—grassroots opinions, niche communities, and emerging trends—largely unanalyzed.

Traditional Online Analytical Processing (OLAP) works wonders for structured databases but fails here because social data isn't just numbers; it's a messy mix of:

  • People: Diverse actors (bloggers, authors).
  • Relation: Links (citations, friendships).
  • Content: Unstructured text (blogs, abstracts).
  • Time: Temporal dynamics (timestamps).

Previous attempts to analyze these dimensions often did so in pairs (e.g., text-to-links), missing the joint influence these dimensions have on one another.

Methodology: The Polyadic Factor Model

The core innovation is treating the dataset as a Tensor (a multi-dimensional array). For a paper citation network, a single "event" is a triple: .

1. Probabilistic Latent Factors

The model assumes that every author, keyword, and reference belongs to a set of "latent factors" (hidden groups). Instead of a rigid mapping, a Core Tensor () captures the many-to-many correlations between these factors.

2. The NTF Approach (Tucker Decomposition)

The authors use Non-negative Tensor Factorization (NTF). The goal is to minimize the KL-divergence between the observed data and a reconstructed tensor derived from factor matrices and core tensor :

Model Architecture: Core Tensor Interaction Figure 1: The Base Transform representing the interaction between core tensor and factor matrices.

3. Scalability via "Lazy" Computation

Handling 10,000+ dimensions would normally crash a system due to memory usage. iOLAP employs a sparse-aware implementation:

  • It only computes entries for non-zero data records.
  • It reuses intermediate matrix products to minimize redundant calculations.
  • The time complexity is linear relative to the number of data records, making it production-ready.

Experiments & Results: The Power of Personalization

The authors tested iOLAP on two major datasets: a 400-blog social network and the CiteSeer citation database.

Multi-Dimensional Drilling

In the blogogram analysis, the model didn't just find "Politics" as a topic. By drilling down into the "Relation" and "Content" dimensions simultaneously, it identified specific sub-groups (e.g., bloggers focusing specifically on terrorism vs. general elections), as shown in Figure 4 of the paper.

Experimental Insights: Community and Topic Correlation Figure 2: The Core Tensor rolled-up, showcasing how different blog groups (x-axis) correlate with content topics (y-axis).

Superior Recommendation Performance

In the CiteSeer task (recommending references for a specific author and keyword), iOLAP crushed the baselines:

  • NDCG@10 Score: iOLAP (0.120) vs. Popularity Baseline (0.076).
  • Why? Because iOLAP understands the context. It doesn't just recommend the most popular paper on "Databases"; it recommends a paper that fits the specific author's research background and the keyword's nuance.

Critical Analysis & Conclusion

Takeaway

iOLAP successfully translates the mathematical rigor of tensor factorization into a functional tool for social media intelligence. It moves us from "Retrieval" (finding a document) to "Analysis" (understanding the landscape).

Limitations & Future Work

The current version of iOLAP captures the intensity of communities over time but does not yet handle structural evolution (e.g., when a community splits into two or merges). Future iterations would benefit from integrating "FacetNet" style dynamic community detection.

For practitioners in Business Intelligence and RecSys, iOLAP provides a blueprint for handling heterogeneous linked data at scale without sacrificing the rich, multi-way relationships that define human networks.

Find Similar Papers

Try Our Examples

  • Search for recent papers that extend the iOLAP framework or Non-negative Tensor Factorization to handle dynamic graph evolution where community structures change over time.
  • Which seminal papers first established the equivalence between Probabilistic Latent Semantic Analysis (PLSA) and Non-negative Matrix Factorization (NMF), and how does this paper generalize that link to tensors?
  • Find studies that apply Polyadic Factor Models or similar multi-way clustering techniques to multi-modal recommendation systems involving images and audio metadata.
Contents
iOLAP: Bridging the Gap Between Search and Multi-Dimensional Business Intelligence
1. TL;DR
2. Problem & Motivation: Beyond the Power-Law Search
3. Methodology: The Polyadic Factor Model
3.1. 1. Probabilistic Latent Factors
3.2. 2. The NTF Approach (Tucker Decomposition)
3.3. 3. Scalability via "Lazy" Computation
4. Experiments & Results: The Power of Personalization
4.1. Multi-Dimensional Drilling
4.2. Superior Recommendation Performance
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work