SocialRank: Leveraging the Social Cloud to Redefine Web Search Relevance
Exploring Social Networks and Improving Hypertext Results for Cloud Solutions
This paper introduces SocialRank, a novel information retrieval ranking algorithm that leverages social network references (Facebook, Twitter, Google+, Delicious) using a cloud-based crawler architecture. Deployed on Amazon EC2, the system aims to improve search relevance by moving beyond traditional hyperlink visibility to incorporate real-time human-driven content sharing.
TL;DR
In an era where "likes" and "retweets" define content value more than hidden hyperlinks, SocialRank proposes a shift from static link analysis to dynamic social data mining. By deploying a specialized crawler on Amazon EC2, the authors developed a ranking mechanism that synthesizes data from Facebook, Twitter, and other platforms to deliver search results that reflect human interest rather than just web structure.
The "Authority" Crisis in Search
For decades, PageRank has been the gold standard, treating a hyperlink as a "vote" of confidence. However, this creates a vulnerability: webmasters can manufacture authority through link farms. Moreover, PageRank is often "blind" to the immediate social relevance of a document.
The authors argue that the quality of a search method requires human evaluation. Social networks provide this evaluation in real-time. The challenge? Social data is vast, volatile, and fragmented across different APIs.
Methodology: Cloud Crawling and the SocialRank Formula
1. The Cloud-Based Crawler
To handle the scalability requirements of social data, the authors moved the crawling process to the cloud (IaaS/PaaS). Unlike traditional crawlers that follow links blindly, this Social Crawler uses REST APIs to target specific keywords and extract metadata like shares, likes, and comments.

2. The SocialRank Algorithm
The core innovation is SocialRank (). A simple count of shares is insufficient because one platform (like Twitter) might over-represent a specific link due to platform-specific trends. To counter this, the authors introduced Standard Deviation () as a smoothing factor.
The formula is defined as: Where:
- = number of references/shares per social network.
- = standard deviation across platforms.
- = total number of social platforms supported.
By subtracting the standard deviation, the algorithm penalizes resources that are only popular on a single network (likely due to bots or "pay-to-share" schemes) and rewards content with broad, cross-platform appeal.
Experimental Insights: SocialRank vs. Google
The researchers compared SocialRank against Google and Bing using technical queries like "Linux."

Key Findings:
- Contextual Difference: Google's top results are often Wikipedia pages or official homepages (structural authority). SocialRank's top results included "BackTrack Linux" and "Linux Mint," reflecting what the community is actually discussing and using right now (utility authority).
- Distribution: The study found that 84.54% of social references are uniformly distributed, validating that SocialRank’s smoothing approach targets the 15% of "anomalous" viral content.
Critical Analysis & Conclusion
SocialRank represents a meaningful evolution in Information Retrieval (IR). By treating social platforms as a distributed "human sensor network," it surfaces content that is currently valuable to users.
Limitations: The current database is small (Google indexes 45+ billion pages). Furthermore, social networks are "walled gardens"; changing API policies can make data extraction difficult.
Future Outlook: The next step for this technology is Personalization. By incorporating a user's own social graph—who their friends are and what they like—Search Engines could move beyond "Broad Relevance" to "Individual Significance." The integration of gender and geographic data, as suggested by the authors, could further refine this social filter.
