Subjective KBC: Bridging the Gap Between Hard Facts and Human Opinions
Subjective Knowledge Base Construction Powered By Crowdsourcing and Knowledge Base
The paper introduces a two-stage framework for Subjective Knowledge Base Construction (SKBC) that integrates Probase for core structure and DBpedia for enrichment. It utilizes crowdsourcing and partial-order inference to acquire subjective facts like "(New York, big City, yes)" at scale and with high quality.
TL;DR
While modern Knowledge Bases (KBs) are masters of objective facts (e.g., "Paris is in France"), they are largely "opinion-blind." This paper presents a framework to build Subjective Knowledge Bases by combining the structured data of DBpedia/Probase with the wisdom of the crowd. By using Partial-Order Relationships, the system can infer subjective labels (like "big," "famous," or "competitive") for thousands of entities while only paying for a fraction of them to be manually annotated.
The "Opinion Gap" in Modern AI
Knowledge Graph construction has historically focused on the observable and the factual. However, a significant portion of user intent is subjective. When a user searches for a "large company," they aren't just looking for a revenue number; they are looking for entities that satisfy the human consensus of "large."
The challenge? Subjective knowledge has no ground truth in some database; the truth lives in the collective mind of the crowd. Traditional KBC methods fail here because they lack a mechanism to aggregate "dominant opinions" cost-effectively.
Methodology: High-Logic Crowdsourcing
The authors don't just ask the crowd to label everything. They use a two-step process:
1. Core Construction & Property Mapping
They mine Probase to identify "ST pairs" (Subjective Property-Type pairs) like <big, City>. They then map these to DBpedia to pull in objective features (Area, Population, etc.) that might correlate with the subjective label.
2. Inferring via Partial-Order
This is the "secret sauce." The authors define a semantic order. If we are looking for "big cities," and we know:
- City A is "big" (annotated by crowd).
- City B has a larger area and more people than City A.
- Therefore, City B is also "big" (inferred).

Cost-Aware Strategy: Adaptive vs. Batch
To minimize the "crowd budget," the paper treats instance selection as a mathematical optimization problem. They prove that their utility function is Adaptive Submodular, meaning we get "diminishing returns" for each new annotation.
- Adaptive Annotation: Selects instances one by one, reacting to the last worker's answer. This is the most cost-efficient.
- Batch-Mode: Selects groups of instances to reduce latency (waiting for workers on Amazon Mechanical Turk).
Experimental Proof
The researchers tested this on 177 gradable adjectives. The results showed that by using the Partial-Order logic, the system maintained high accuracy while drastically reducing the number of human tasks needed.

As shown in the cost comparison, the Adaptive and Batch strategies significantly lowered the "HIT cost" (Human Intelligence Tasks) required to fully populate the KB compared to random sampling.
Performance vs. Machine Learning
Interestingly, the framework's inference rules outperformed pure machine learning models like SVMs and Decision Trees (83.7% vs 70.5-79.6% accuracy). This suggests that human-derived logic (partial ordering) is a stronger inductive bias for subjective tasks than standard statistical features.
Critical Insight & Future Outlook
This paper effectively turns "subjectivity" into a structured, computable domain. By linking objective metrics (area, employees, endowment) to subjective adjectives (big, large, famous), it creates a bridge between quantitative data and qualitative human language.
Limitations: The model relies on "gradable" adjectives. It might struggle with highly polarized or culturally dependent subjectivity (e.g., "beautiful") where a "dominant opinion" may not exist or may shift rapidly over time.
Takeaway: The future of KBs is not just about what is true, but what is perceived. This framework is a vital step toward search engines that truly "understand" human adjectives.
