Beyond Accuracy: Building Trust in Historical Genre Classification Through Transparency
Utilizing a Transparency-Driven Environment Toward Trusted Automatic Genre Classification: A Case Study in Journalism History
This paper introduces a transparency-driven environment designed for historical journalism research, focusing on the automatic genre classification of newspaper articles. By integrating interpretability tools like LIME and feature importance visualizations, the authors enable non-computer scientists to move beyond raw accuracy and select models based on expert alignment and trust.
TL;DR
In the world of Digital Humanities, a model that is 90% accurate but looks at the "wrong things" is more dangerous than a 60% accurate model that follows human logic. This paper presents a transparency-driven environment that empowers journalism historians to "peek under the hood" of AI. By using local and global explanations, researchers discovered that high-performing models were often just "topic classifiers" in disguise, leading them to choose theoretically sound models over high-scoring ones.
Background: The Hidden Danger of Black-Box Accuracy
In journalism history, the genre of an article (e.g., a "News Report" vs. an "Op-ed") is a crucial indicator of shifts from opinion-oriented to fact-centered reporting. However, manual labeling of millions of digital archive entries is impossible. While machine learning offers a solution, the "Black Box" problem looms large: how do we know the AI isn't just labeling an article as "Sports" because it sees the word "ball," rather than identifying the structural style of a "Report"?
The Insight: "All Models are Wrong, but Some are Transparent"
The authors argue that for a historian to trust a machine-generated result as the basis for a scientific claim, they must understand the Why behind a prediction. They developed an environment that moves beyond the Accuracy/F1-score fetishism, focusing instead on decomposability and interpretability.
Methodology: A 6-Step Workflow for Trust
The system guides users through a pipeline—from data preprocessing to hypothesis testing—integrating tools that visualize what the model "thinks" is important.
1. The Architecture of Understanding
The environment supports multiple document representations, notably comparing TF-IDF (which focuses on word importance) against NLP-curated features (which look at POS tags, sentence length, and subjectivity).
Fig 1. The workflow turning raw data into robust historical claims via transparent ML pipelines.
2. Global vs. Local Explanations
- Global Explanations: Uses feature importance plots to show what a model values across the entire dataset. For instance, the use of "yesterday" (gisteren) was correctly identified as a key feature for News and Reports.
- Local Explanations (LIME): Provides a per-article breakdown. This proved critical when historians found that a Support Vector Machine (SVC) was misclassifying interviews as "Op-eds" because the topic was political, whereas a Random Forest (RF) model correctly identified the "Interview" based on the frequency of quotation marks and pronouns.
Experimental Battle: Accuracy vs. Theory
The researchers compared several pipelines. Interestingly, the model with the highest accuracy (SVC using TF-IDF) was eventually sidelined.
The "Smoking Gun" in the Confusion Matrix
By examining the confusion matrices, the historians noticed that while some models scored high, they were completely failing to identify the "Feature" (Reportage) genre accurately, often confusing it with "Background" due to topic overlap.
Fig 2. Confusion matrices showing the performance of different algorithms on unseen data.
Lessons Learned: The Value of "Lower" Performance
The most striking takeaway is that the domain experts ultimately specialized in a Random Forest model using FROG scaling. Why? Because its explanations were "closer to the journalism historians' abstract understanding."
- The Topic Bias: TF-IDF models often fall into the trap of "Topic Classification." If all "Reviews" in the training set are about books, the model learns that "book" = "Review," which fails when it encounters a review of a movie.
- The Linguistic Signal: The NLP-based features caught "subjectivity" and "adjectives," which are the actual stylistic markers of a "Feature" genre, regardless of whether the topic is politics or gardening.
Critical Analysis & Conclusion
Takeaway
This work marks a shift in eScience from "performance-first" to "transparency-first." It proves that for multi-disciplinary research, a common language—provided here by LIME and feature plots—is necessary to bridge the gap between algorithmic math and humanistic theory.
Limitations & Future Work
While the environment provides transparency, it still relies on "Gold Standard" data that may have its own selection biases. Future iterations aim to look at "Bias in the Pipeline," examining how the choice of NLP tools (like Frog vs. SpaCy) impacts the final historical interpretation.
Ultimately, this study confirms a vital rule for the future of AI in research: Confidence in a model should come from understanding its logic, not just staring at its scoreboard.
