From Tweets to Ontologies: Automating User Profiling Through Shared URLs
Automatic Ontology User Profiling for Social Networks from URLs Shared
The paper introduces a scalable, automated framework for building semantic user profiles from URLs shared on Twitter. By leveraging collective knowledge bases like OpenDNS and DBpedia, the system categorizes user interests and purchase intentions into a formal ontology (OWL/RDF) to enable personalized services like targeted advertising.
TL;DR
Researchers have developed a scalable system that transforms the URLs shared in tweets into a rich, semantic map of a user's world. By combining NoSQL scalability with the reasoning power of Ontologies, the system can infer not just what a user likes, but what they intend to buy—moving beyond simple keyword matching to true semantic understanding.
Background: The Problem with Surface-Level Profiling
In the era of personalized advertising, knowing your user is everything. However, most platforms struggle because:
- Data is Sparse: Twitter bios are often empty or sarcastic.
- Context is Missing: A tweet containing the word "Apple" could refer to a fruit, a tech giant, or a record label.
- Volume is Overwhelming: Tracking millions of users in real-time requires more than just a relational database; it requires a specialized architecture.
The authors of this paper argue that URLs are the key. When a user shares a link, they provide a curated piece of their interest. By "de-referencing" these links through collective knowledge bases, we can build a much more accurate profile than by analyzing text alone.
Methodology: The "Moriarty" Architecture
The proposed system, part of the Novared project, follows a sophisticated pipeline to turn raw tweets into ontological assertions.
1. Data Ingestion & Storage
The system uses Apache Cassandra to store raw tweets. Its column-oriented nature allows it to handle the massive write-speeds required for Twitter streams.
2. Semantic Analysis
The core logic resides in a library called Moriarty. This component:
- Extracts URLs from tweets.
- Filters out "noise" (search engines, shorteners).
- Queries OpenDNS and DBpedia to find out what those URLs represent.
3. Ontological Reasoning
The true "magic" happens in the Virtuoso RDF database. The system doesn't just store "User A likes Cars." It uses the Pellet Reasoner to apply logic:
- If a URL is about "Avis", and "Avis" is a "Car Rental Company", and "Car Rental" is an "Intention to Travel", then the User has a Travel Intention.
Figure 1: High-level architecture showing the flow from Twitter to the Reasoner and NoSQL storage.
The Power of the Ontology
The researchers built a custom ontology that merges several standards:
- FOAF (Friend of a Friend): For basic user relationships.
- Advertisement Taxonomies: To define specific categories like "Tech Enthusiasts" or "Auto Buyers."
By using classes and subclasses, the profile "evolves" over time. Every new tweet is not just a data point; it's a new "individual" added to a living knowledge base.
Figure 2: Visualization of inferred relationships within the interest and intention hierarchy.
Experimental Insights
In a test of 8,000 users, the system proved remarkably adept at pinpointing intentions.
- Specific Inference: A user sharing a link to
avis.comwas correctly mapped to categories like "Companies based in Detroit" and "Transportation companies," finally landing in the "Travel Intention" bucket. - Scalability: By utilizing Virtuoso and Cassandra, the system managed billions of relationships, proving it could lead to industrial-scale applications.
Critical Analysis & Conclusion
Takeaway
The shift from Bag-of-Words (BoW) to Ontologies is a shift from "guessing" to "reasoning." This paper demonstrates that by using shared URLs as a semantic anchor, we can bypass the noise of social media language and get straight to the user's intent.
Limitations
While powerful, the system is heavily dependent on the coverage of OpenDNS and DBpedia. If a user shares a very new or obscure blog, the system might categorize it as UnknownCategory. Furthermore, the approach currently focuses on the "what" (the URL) but largely ignores the "how" (the sentiment or opinion expressed about that URL).
Future Directions
The authors suggest that the next step is integrating sentiment analysis of the tweet text itself. Imagine a system that knows you are interested in "Electric Cars" (from a URL) but also knows you have a "Negative Sentiment" towards a specific brand (from the text). That is the future of hyper-personalized internet services.
Key References
- Malhotra, A. et al. (2012). Studying User Footprints in Different Online Social Networks.
- Pennacchiotti, M., Popescu, A. (2011). A Machine Learning Approach to Twitter User Classification.
- Sakaki, T. et al. (2010). Earthquake Shakes Twitter Users: Real-time Event Detection by Social Sensors.
