Syntax-Based Plagiarism Classification: Efficiency Meets Accuracy in NLP

Expert Systems With Applications

2025-01-01
Som Gupta
Summary
Problem
Method
Results
Takeaways
Abstract

The paper proposes a binary text plagiarism classification framework using shallow syntactic linguistic features (POS tags and phrase chunks). By employing a two-phase feature selection approach and machine learning classifiers (NB, SVM, DT), the method achieves SOTA performance on the PSA and PAN corpora, reaching up to 97.89% accuracy.

TL;DR

Plagiarism detection is often a battle between speed and depth. This paper introduces a highly efficient intermediate classification stage that uses shallow syntactic features (POS tags and Chunks) to filter suspicious documents. By reducing the problem to just three core linguistic features—Nouns, Verbs, and Adjectives—the authors achieved nearly 98% accuracy, outperforming complex baselines while significantly lowering computational overhead.

The Bottleneck in Current Detection Systems

Most plagiarism detection workflows follow a three-step process: Candidate Retrieval, Detailed Analysis, and Post-processing. The current crisis in the field lies in the "Detailed Analysis" phase. While N-gram overlaps work for "copy-paste," they fail at "paraphrasing." Conversely, deep semantic analysis (like word embeddings) is too slow to run on thousands of candidate pairs.

The authors identify a missing link: a high-speed, high-precision Intermediate Classification Stage that can prune the search space before expensive passage-level alignment begins.

Methodology: The Power of Shallow Syntax

The core insight of this work is that even when a sentence is paraphrased, its "syntactic skeleton" often leaves a trail. Instead of looking at every word, the authors focus on:

  1. POS Tags: Identifying the distribution of functional categories.
  2. Chunks: Analyzing phrases (Noun Phrases, Verb Phrases) of varying lengths (>=1 and >=2).

Feature Engineering & Selection

The authors started with 14 features but realized that "more is not always better." They implemented a Two-Phase Feature Selection process:

  • Phase 1 (Filter): Using Pearson’s Correlation to rank features and remove the "noise."
  • Phase 2 (Wrapper): Using Correlation-based Subset Selection (CFS) with a Best-First search to find the synergistic "Goldilocks" set of features.

Flowchart of the Proposed Methodology

Experiments and Results

The model was tested against two major benchmarks: the PSA (Plagiarized Short Answers) corpus and the PAN competition sets (including artificial and realistic/simulated cases).

Key Performance Metrics

  • Dimensionality Reduction: The system successfully compressed 14 features down to just 3 (Noun, Verb, Adjective frequencies) without losing signal.
  • Superior Accuracy: On the PSA corpus, the Decision Tree (DT) classifier reached 97.89% accuracy.
  • The Baseline Killer: In a multi-class comparison, the proposed method crushed existing SOTA models by over 7 percentage points in accuracy while using half the number of features.
MetricChong et al. (2010)Sanchez-Vega et al. (2013)Proposed (NB/DT)
Accuracy85.26%90.63%97.89%
Feature Count77+3

ROC and Entropy Analysis

To avoid the "Accuracy Paradox" (where a model looks good just by guessing the majority class), the authors used Entropy Triangles and Normalized Information Transfer (NIT). This proved that the model wasn't just lucky—it was actually capturing mutual information between the source and the suspicious text.

ROC Curve Analysis

Critical Insight: Why Does This Work?

The paper offers a fascinating ablation study on "Plagiarism Complexity."

  • For "Near Copy": Both POS and Chunks work perfectly because the structure is identical.
  • For "Heavy Revision": Chunks fail because the sentence structure is rebuilt. However, the density of specific POS tags (Noun/Verb/Adjective) remains surprisingly stable, acting as a "linguistic fingerprint" that survives even deep paraphrasing.
  • The "Artificial" Limitation: Interestingly, the model performs slightly worse on "Artificial" plagiarism (randomly shuffled words). Why? Because random shuffling destroys the grammatical structure that POS tagging relies on. This highlights that this method is best suited for real-world, human-generated plagiarism.

Conclusion

This research demonstrates that we don't always need the "heavy machinery" of Transformers or deep semantics for every stage of detection. By focusing on the shallow syntactic layer, we can build filters that are not only faster but more robust against the cleverest human paraphrasers.

Future Outlook: Integrating these syntactic features with citation analysis could be the final frontier in catching high-level "idea plagiarism" in scientific publishing.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize deep learning-based dependency parsing or graph neural networks for structural plagiarism detection.
  • What are the current state-of-the-art methods for "Intrinsic Plagiarism Detection" where no reference source corpus is available?
  • Explore research that applies Cross-Language Word Embeddings (CLWE) to solve translation-based plagiarism in academic publications.
Contents
Syntax-Based Plagiarism Classification: Efficiency Meets Accuracy in NLP
1. TL;DR
2. The Bottleneck in Current Detection Systems
3. Methodology: The Power of Shallow Syntax
3.1. Feature Engineering & Selection
4. Experiments and Results
4.1. Key Performance Metrics
4.2. ROC and Entropy Analysis
5. Critical Insight: Why Does This Work?
6. Conclusion