Forensic Translation Analysis: Identifying the Human in the Machine

A Classifier to Determine Whether a Document is Professionally or Machine Translated

2016-01-01
Michael Luckert, Mortiz Schaefer-Kehnert, Welf Löwe, Morgan Ericsson, Anna Wingkvist
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a binary classification method based on Decision Trees (C4.5) to distinguish between professional and machine-translated technical documentation. By leveraging the METEOR and BLEU metrics as features alongside pseudo-references, the system identifies translation origins with up to 70.48% accuracy at the sentence level and achieved 100% document-level accuracy for texts exceeding 100 sentences.

TL;DR

In an era of ubiquitous AI, how do you know if the technical manual you paid for was translated by an expert or just fed into a neural engine? This paper presents a machine learning approach using Decision Trees and pseudo-references to detect machine-translated (MT) content. While sentence-level detection hovers around 70% accuracy, the authors prove that for any document over 100-250 sentences, they can identify its origin with near 100% certainty.

Background: The High Stakes of Technical Translation

For international companies, translation isn't just about communication—it's about liability. Technical documentation must meet strict legal regulations (like EU directives). If an MT engine hallucinates a critical safety instruction, the cost isn't just a linguistic error; it's a legal catastrophe. The "Professional vs. Machine" classifier acts as a gatekeeper, ensuring that "professional" translations aren't just "post-edited" machine outputs in disguise.

The Core Challenge: Missing References

Traditional metrics like BLEU or METEOR require a "Gold Standard" human translation to compare against. But if the goal is to determine if a translation is human to begin with, you usually don't have that reference.

The authors solve this through two clever workarounds:

  1. Pseudo-References: If you have the source text, translate it using three different MT engines (Google, Bing, Freetranslation). A machine-translated candidate will naturally gravitate toward these automated references in terms of n-gram overlap and structure.
  2. Round-Trip Translation (RTT): If you don't even have the source document, translate the candidate into another language and back. The "identity" loss in this process varies significantly between human-written and machine-generated text.

Methodology: Decision Trees and Feature Engineering

The researchers extracted 32 features, prioritizing transparency. They chose C4.5 Decision Trees specifically because they allow us to see why a decision was made.

Key Features Include:

  • METEOR & BLEU Scores: Measuring unigram/n-gram similarity.
  • Translation Edit Rate (TER): How much effort it takes to turn the candidate into a reference.
  • Linguistic Indicators: Flesch Reading Ease (readability) and Part-of-Speech (POS) patterns.
  • Mistake Count: Using tools like LanguageTool to find automated-looking grammatical slip-ups.

Model Decision Tree Figure 1: An optimized Decision Tree showing METEOR scores at the root, indicating it is the most influential factor in detection.

Experimental Results

The study used a dataset of 22,327 sentences. The findings reveal a "Law of Large Numbers" effect:

  • Sentence Level: Hard. Accuracy is roughly 70.48%. There is too much noise in single sentences to be perfectly sure.
  • Document Level: Extremely reliable. By using a majority-voting system (classifying every sentence and taking the winner), the error rate vanishes as document length increases.

Performance Comparison Table 5: Note how the "Misclassification" count drops to 0 once document length hits 100 sentences.

Critical Insight: The Advantage of Technical Domains

The authors limited their scope to German-English technical documentation. This was a strategic choice. Technical language has a limited vocabulary and restricted syntax. In this "bounded" environment, machine learning finds it easier to spot the rigid, often over-generalized patterns of MT engines compared to the more varied (yet precise) terminology of a human expert.

Conclusion & Future Look

The paper effectively turns MT evaluation metrics on their head—using them not to measure "goodness," but to find "similarity to other machines."

Limitations: The study relies on 2016-era MT engines. Modern LLMs (like GPT-4) produce much more "human-like" syntactic patterns, which might defeat a classifier relying purely on BLEU/METEOR scores.

Future Work: The transition from Decision Trees to Deep Learning architectures (like Transformers) and the inclusion of semantic consistency checks (rather than just syntactic) will be the next frontier in the "Human vs. Machine" translation war.

Find Similar Papers

Try Our Examples

  • Find recent studies on "Machine Translation Detection" that use deep learning or Large Language Model (LLM) embeddings instead of traditional n-gram metrics like BLEU.
  • Which paper originally established the effectiveness of "pseudo-references" in MT evaluation, and how does the current work's use of them for classification differ from their use in quality estimation (QE)?
  • Explore research that applies similar document-origin classification techniques to detect AI-generated content (LLM detection) in specialized technical or medical domains.
Contents
Forensic Translation Analysis: Identifying the Human in the Machine
1. TL;DR
2. Background: The High Stakes of Technical Translation
3. The Core Challenge: Missing References
4. Methodology: Decision Trees and Feature Engineering
4.1. Key Features Include:
5. Experimental Results
6. Critical Insight: The Advantage of Technical Domains
7. Conclusion & Future Look