Moveek: Bridging Social Networking and the Semantic Web via Service Choreography
Towards a semantic social network
This paper introduces "Moveek," a semantic social network designed for academic knowledge sharing that leverages Web Service Composition to integrate disparate technologies. The core contribution is an automated semantic annotation and query system built using the GATE and Protégé APIs to transform plain text posts into ontology-linked data resources.
TL;DR
Moveek is a semantic social network that automates the transition from "social posts" to "knowledge resources." By implementing a Choreography-based Web Service architecture, it integrates Java-based NLP tools with PHP/.Net frontends to automatically annotate scientific text and enable conceptual (rather than keyword) search.
Problem & Motivation: The "Static" Nature of Social Data
Standard social platforms treat text as "opaque" strings. In an academic/scientific context, this is highly inefficient. If a researcher posts about "DNA," the system should understand it not just as a three-letter word, but as a biological entity with relationships to other concepts.
The technical challenge lies in interoperability. Implementing advanced semantic tasks (Information Extraction, Ontology Reasoning) typically requires specific Java environments (like GATE or Protégé), while modern web interfaces often prefer .Net or PHP. Bridging these while maintaining a seamless user experience is the core hurdle this paper addresses.
Methodology: Automated Annotation and Service Composition
The authors propose a "Service Composition" strategy, opting for Choreography over Orchestration. Unlike a centralized orchestrator, choreography allows services to interact through public message exchange, which is ideal for a distributed social network architecture.
The Semantic Pipeline
- Semantic Annotation: When a user publishes a post, it is sent to an Information Extraction (IE) service. Using the GATE API, the system identifies scientific terms modeled in an underlying ontology.
- XML Mapping: The engine produces an XML output containing the
classURIList(identifying the category) and theindividual identification(identifying the specific instance). - Hyperlink Transformation: The final post is rendered with hyperlinks. Clicking these links triggers a semantic query.
Figure 1: The pipeline from plain text to semantic annotation.
The Query Engine
Instead of performing a traditional SELECT * LIKE '%keyword%' database query, Moveek uses the Protégé API. It receives a Class URI and uses an object property within the ontology to retrieve all publication identifiers related to that concept, ensuring that users find semantically relevant content even if the exact words differ.
Figure 2: The logic of the Semantic Query Service.
Experiments & Results: Real-world Performance
The system was deployed at the Benemérita Universidad Autónoma de Puebla. The authors analyzed 142 posts, finding that 53 were correctly annotated.
| Metric | Value |
|---|---|
| Total Posts | 142 |
| Annotated Posts | 53 |
| Success Logic | Ontology-dependent (Terms must exist in scientific thesaurus) |
A significant result revealed in the experiments is the semantic precision. For instance, a query for "Guerra" (War) can return publications that share the same context or meaning defined in the ontology, rather than just matching the literal string.
Figure 3: The interaction between Java-based semantic services and .Net data layers.
Critical Analysis & Conclusion
Takeaway
The paper successfully demonstrates that ontologies can be integrated into social platforms without overwhelming the user. The use of Service Choreography is a smart architectural choice for handling the diverse tech stacks (Java, .Net, PHP) required for such a system.
Limitations
- Ontology Coverage: The system currently only annotates ~37% of posts. This is a "Cold Start" problem for ontologies; if the scientific term isn't in the database, the system reverts to a standard post.
- Language Support: The authors noted a lack of comprehensive Spanish scientific thesauri, which limits the extraction engine's effectiveness in its local context.
Future Outlook
As Natural Language Processing moves toward Generative AI and Large Language Models, the "hard-coded" ontology approach in Moveek could be evolved into a hybrid system. LLMs could generate "on-the-fly" ontology suggestions, which Moveek's choreography could then validate and store, significantly increasing the annotation rate.
