Gender through the Lens of Vectors: Decoding 19th Century Fiction
Exploring the Role of Gender in 19th Century Fiction Through the Lens of Word Embeddings
This paper explores how author gender influences narrative word choice in 19th-century fiction using word embeddings. By applying Word2Vec and t-SNE to a curated corpus of 48 novels, the authors quantify semantic differences in gendered terms (e.g., "she", "husband", "gentleman") between male and female writers.
TL;DR
Does a writer's gender fundamentally change the "semantic neighborhood" of the words they use? This paper uses Word2Vec and t-SNE to analyze 48 British and Irish novels from the 1800s. The authors discover that while some words like "gentleman" are used similarly across genders, pronouns and domestic terms like "husband" reveal vast differences in how male and female authors perceive their social worlds.
Problem & Motivation: Beyond Close Reading
In the Digital Humanities, scholars are moving from "close reading" (analyzing single texts) to "distant reading"—understanding literary history at a macro scale. The researchers noticed that while Word2Vec is often used to remove bias (debiasing), it is a powerful tool to expose and quantify historical bias. They set out to see if the "distributional hypothesis" (words are defined by the company they keep) could reveal the hidden gendered textures of 19th-century life.
Methodology: Mapping the Victorian Mind
To capture these nuances, the team didn't just train a model on raw text. They:
- Annotated the Corpus: Manually identified character names and gendered unigrams.
- Gender-Encoding: They tagged words based on the author; for example, every "she" in a Jane Austen novel became
she_female. - Embeddings: Used a Skip-gram model (300D) to turn these words into vectors.
- Spatial Analysis: Used Cosine Similarity to compare how "close" the male and female versions of the same word were in the vector space.

Key Insights: Emotional "She" vs. Structural "He"
The results provide a fascinating quantitative look at Victorian social norms.
1. The Proximity of Pronouns
The study found that for female authors, the pronoun "her" lives in a neighborhood of emotional and physical vulnerability: words like “shaking,” “sobs,” and “trembling”. In contrast, in male-authored texts, "her" is surrounded by other functional pronouns like "she" and "him," suggesting a more externalized or structural use of the character.
2. The Occupational Bridge
Interestingly, "gentleman" and "lady" often occupied similar semantic pockets regardless of the author's gender. Both genders associated "gentleman" with professional titles like farmer, clergyman, and lawyer. This suggests that the public, professional sphere had a relatively standardized linguistic representation.

3. Domestic Dissimilarity
The word "husband" had one of the lowest cosine similarities between male and female authors. This implies that the concept of a husband was framed very differently depending on the gender of the person writing the narrative—likely reflecting the starkly different lived experiences of men and women in the 19th-century domestic sphere.

Critical Analysis & Conclusion
This paper successfully bridges the gap between Computational Linguistics and Literary Criticism.
Takeaway: The study proves that "style" is not just about word frequency, but about semantic association. Male and female authors lived in the same century but often wrote in different "semantic universes."
Limitations: The corpus of 48 novels is relatively small for deep-learning models. The authors acknowledge that a "diachronic" (through-time) analysis is needed to see how these gendered associations shifted as the century progressed into the Victorian fin de siècle.
Future Outlook: By scaling this to thousands of books, we could create a "heatmap" of social change, identifying exactly when professional roles or domestic descriptions began to desegregate in the literary imagination.
