Garbage, Glitter, or Gold: Decoding Quality in Social Media Seeds for Web Archiving
Garbage, Glitter, or Gold: Assigning Multi-Dimensional Quality Scores to Social Media Seeds for Web Archive Collections
This paper introduces the Quality Proxies (QP) framework, a multi-dimensional scoring system designed to evaluate and rank seed URLs extracted from social media for Web archive collections. By combining 10 major dimensions including popularity, proximity (geographical/temporal), and reputation, the framework achieved a precision increase of approximately 0.13 compared to non-scored baselines across topics like COVID-19 and the Flint Water Crisis.
TL;DR
The Web is ephemeral; roughly 11% of resources shared on social media disappear within a year. To combat this "reference rot," archivists use social media to find "seeds" (URLs) for preservation. However, social media is a mix of gold and garbage. The Quality Proxies (QP) framework offers a robust, 14-dimensional scoring system—incorporating popularity, geography, and reputation—to help curators automatically identify and prioritize high-quality URLs for archive collections, improving selection precision by 13%.
The Archival Dilemma: Popularity $
eq$ Quality When a major event like the COVID-19 pandemic or the Flint Water Crisis occurs, archivists race against the clock. While social media is a treasure trove of real-time information, it suffers from two major issues:
- The Noise Floor: High-volume social media data is filled with "glitter" (shallow popularity) and "garbage" (misinformation).
- The Popularity Bias: If we only archive what is popular, we miss the "local gold"—the small-town journalist or the local doctor whose reports are expert and timely but lacks a global following.
Existing tools often rely on simple relevance or raw "likes." The authors argue that quality is multi-dimensional. A seed might be valuable because it’s reputable (cited by Wikipedia), proximal (written by someone at the epicenter), or scarce (not found on Google).
Methodology: The Quality Proxies (QP) Framework
The core innovation is viewing quality as a vector rather than a single number. The framework identifies three main classes of proxies:
1. Popularity Proxies
These track engagement across the Post (retweets, likes), the Author (follower count), and the Domain (associated social media influence).
2. Proximity Proxies (The "Local" Advantage)
- Geographical (gea, ged): Uses the Haversine formula to reward sources physically close to the event.
- Temporal (tp): Rewards "early" reporting.
3. Reputation & Expertise Proxies
- Reputation (reb, ren): A clever insight where a domain’s quality is approximated by how often it is cited in topic-specific Wikipedia "gold standard" articles.
- Retrievability (rt): Measures how hard a URL is to find on Google; higher scores for "hard-to-find" but relevant content increase the novelty of the archive.

The Scoring Function
The framework calculates the final score using the 2-norm of the proxy vector: This allows curators to "flip" or weight dimensions. For example, if you want to amplify obscure voices, you can penalize popularity while rewarding local proximity.
Experimental Results: Proving the "Gold"
The authors evaluated the framework against human-curated Expert collections (Archive-It) and Google search results across topics like the 2018 World Cup and the Ebola outbreak.
- Precision Boost: Using QP scores to rank Twitter seeds improved the overlap with expert-selected "gold" seeds significantly. Precision at K (P@K) increased by up to 0.17.
- The Novelty Test: Perhaps most importantly, even when seeds had zero overlap with Google (i.e., they were completely novel), the QP framework was still able to identify relevant, high-quality content that search engines missed.

Deep Insight: Why This Matters
Social media research often focuses on What is being said (sentiment, text). The QP framework focuses on Where that information leads (the URL).
By grounding reputation in Wikipedia citations and geography in physical coordinates, the authors move away from "black-box" AI models toward an explainable curation system. If a seed is ranked highly, a curator can see exactly why—perhaps it was a local news outlet in Michigan during the Water Crisis that was cited by Wikipedia health editors.
Limitations and Future Outlook
While powerful, the framework relies on existing data density. For "esoteric or obscure" stories with very few social media posts, the proxies may struggle due to sparse data. Additionally, the current relevance metric is limited by long-tail text analysis.
However, the QP framework serves as a vital blueprint for the next generation of "Autonomous Archivists," ensuring that our digital history isn't just a collection of what was loudest, but what was truest.
Senior Editor's Take: This work bridges the gap between Social Media Analytics and Library Science. Its biggest strength is its modularity; as new platforms emerge, the 'Proxies' can be swapped, but the underlying logic of multi-dimensional quality remains a gold standard for digital preservation.
