Text Knowledge Mining: Shifting from Pattern Finding to Logical Discovery
Text Knowledge Mining: An Alternative to Text Data Mining
This paper introduces Text Knowledge Mining (TKM) as a distinct paradigm from traditional Text Data Mining (TDM). While TDM relies on inductive inference to find patterns in structured "intermediate forms," TKM utilizes deductive and abductive inference to derive new, non-trivial knowledge directly from the semantic content of natural language.
TL;DR
Is text just "unstructured data," or is it a "Knowledge Base" waiting to be queried? This paper argues the latter. By introducing Text Knowledge Mining (TKM), the authors distinguish between simply finding statistical regularities (Induction) and deriving new logical truths through reasoning (Deduction and Abduction). This paradigm shift allows for the discovery of hypotheses and contradictions that traditional "bag-of-words" models completely miss.
Problem & Motivation: The Poverty of Induction
Most text mining today follows the Text Data Mining (TDM) pipeline:
- Flatten text into a "Bag-of-Words" or N-grams.
- Apply inductive algorithms (Clustering, Association Rules).
- Find patterns based on frequency.
The authors argue that this is fundamentally flawed because natural language is not data—it is knowledge. When we flatten a sentence into a frequency vector, we lose the "Why" and the "How." We lose the causal links and the logical constraints. For example, induction can tell you that "migraine" and "magnesium" are statistically related, but it cannot easily hypothesize why if they never appear together in the same document.
Methodology: The Core of TKM
The authors propose that TKM should treat text as a representation of general knowledge. The distinction lies in the type of Inference:
- Inductive (TDM): Starting from cases to find general rules.
- Deductive (TKM): Starting from known facts/rules in text to derive new consequences or identify contradictions.
- Abductive (TKM): Reasoning from effects to possible causes (Hypothesis generation).
Architecture of Representation
To perform reasoning, TKM requires more sophisticated "Intermediate Forms" than simple word counts. The paper highlights several structures:

As shown in the table, moving from "Words" to "Concepts" and "Paragraphs" allows for Conceptual Graphs and Semantic Graphs. These structures preserve the relationships necessary for formal logic.
The Discovery of Contradictions
One of the paper's specific contributions is a TKM algorithm for finding contradictions. It translates text into First-Order Logic (FOL) clausal forms. If a collection of texts leads to a logical "False" through resolution, a contradiction is discovered. This is "mining" because the contradiction is a non-trivial piece of new knowledge about the consistency of the repository.
Experiments & Results: Real-World Discovery
The paper cites two major "battle-tested" examples of TKM in action:
- The Magnesium-Migraine Hypothesis (Abduction): Using the Arrowsmith project, researchers extracted causal clues from MEDLINE titles. By linking and from disparate papers, they hypothesized (Magnesium deficiency causes migraines). This hypothesis was medically confirmed after the mining process.
- Logic-Based Consistency Checking (Deduction): Using the UNL (Universal Networking Language) as an interlingua, the authors demonstrate an algorithm that can prune consistent text sets to isolate specifically where different authors or reports conflict.
Algorithmic Approach to Contradiction Discovery
The author proposes a levelwise exploration of text subsets (similar to the Apriori algorithm) but uses Resolution instead of Support/Confidence:
(Note: This algorithm iterates through subsets of text, assessing consistency at each step to find the minimal set of contradicting documents.)
Critical Analysis & Conclusion
Takeaway
The value of TKM is its ability to create "Novel Investigation." While TDM finds what is already there (just hidden), TKM creates hypotheses that might not exist in any single document.
Limitations
- Knowledge Acquisition Bottleneck: Translating natural language into First-Order Logic remains computationally expensive and manually intensive (semi-automatic).
- Monotonicity: FOL struggles with the "reliability" of text. If one document is a lie, the whole deductive chain might collapse.
Future Outlook
With the advent of modern Knowledge Graphs and LLMs, the "TKM" vision is more relevant than ever. The industry is moving away from simple keyword search toward RAG (Retrieval-Augmented Generation) and Reasoning Engines, which are the spiritual successors to the TKM framework proposed in this paper.
