Detecting Plagiarism in Micro-Blogs: A Semantic Approach to Short-Text Integrity
Detecting plagiarism in micro-blogging social networks
The paper presents a plagiarism detection tool tailored for micro-blogging social networks in educational settings, integrated into the "Bolotweet" platform. It utilizes a semantic analysis approach based on a multilingual thesaurus (MCR) to identify rewritten content beyond simple cut-and-paste.
TL;DR
As social media-style "micro-annotations" become popular in classrooms (the "Write-to-Learn" model), academic integrity faces a new challenge: short-text plagiarism. This paper introduces a plugin for the Bolotweet network that uses semantic synset matching rather than simple text comparison to catch students who "reword" rather than "research," all while maintaining a low computational footprint through optimized SQL queries.
Background: The Micro-Blogging Dilemma
In modern pedagogy, students are often asked to summarize concepts in 140-character snippets. While this encourages concise thinking, it also makes cheating easier. Standard tools like Turnitin are overkill—they are expensive, slow, and designed for long essays. On the other hand, simple keyword matching is easily defeated by a student with a dictionary. This paper situates itself as a practical, self-hosted middle ground for educators.
Why Literal Comparison Fails
The author highlights that plagiarism in micro-blogs isn't just "copy-paste." Students often:
- Change word order.
- Substitute words with synonyms.
- Slightly rephrase a peer's successful summary to gain similar points.
Traditional n-gram analysis lacks the "depth" to see that two different sentences are expressing the exact same concept.
Methodology: Semantic Synsets and SQL Efficiency
The core innovation lies in the use of the Multilingual Central Repository (MCR). Instead of comparing strings like "Search" and "Hunt," the system maps both to a shared Synset ID.
The Workflow:
- Preprocessing: Remove stop words and stem the remaining terms.
- Expansion: Every word is mapped to its possible synsets (meanings) and their synonyms.
- Database Integration: These synsets are stored in a relational database.
- The Query: When a new annotation arrives, a single SQL query calculates the intersection of its synsets with all previous entries.

The similarity is a simple ratio:
Experimental Use Case
The system was tested in an Artificial Intelligence course where Spanish-speaking students submitted summaries of lectures. The tool provided a real-time dashboard for professors.
Figure 1: The interface showing a micro-annotation review where "?? " indicates low word count, necessitating human intervention.
In one instance (Figure 3 and 4), the system flagged a student's post about search algorithms with a 0.71 similarity score. The professor could immediately see that the student had mirrored a peer's post from the previous day, allowing for a more informed (and perhaps lower) originality score.
Figure 2: A student's post flagged with 0.71 similarity, allowing the professor to identify the source.
Critical Insight & Conclusion
The beauty of this approach is its computational efficiency. While BERT-based or modern LLM embeddings might offer higher nuanced precision today, the author’s SQL-based synset approach allows for near-instantaneous results on standard server hardware without the need for GPUs or expensive APIs.
Limitations: The system relies heavily on the quality of the thesaurus (MCR). If a student uses highly technical slang or "circumlocution" (describing a concept without using its name), the system may produce a false negative. However, as an "assistant" tool rather than an "automated judge," it significantly reduces the professor's mental load in identifying potential academic dishonesty.
