Active Knowledge Mining: Beyond Simple Clicks to Intelligent Web Management
Active Knowledge Mining for Intelligent Web Page Management
The paper introduces "Active Knowledge Mining," a framework for intelligent Web page management using the "System L-R." It leverages Web usage mining to construct user models via transition probabilities and clustering, achieving significant improvements in user navigation efficiency.
TL;DR
In the early 2000s, as the Web exploded in complexity, users struggled to find information while administrators lacked tools to diagnose poor site architecture. This paper introduces Active Knowledge Mining via System L-R, a framework that uses transition probabilities and structural weighting to recommend pages and identify "ill-formed" site designs. By filtering out "back-button" noise and clustering user paths, the system significantly lowers the navigation hurdle for novice users.
Problem & Motivation: The Gap Between Design and Reality
The core challenge in Web management is the misalignment between how a site is intended to be used and how it is actually navigated. Existing methods often viewed Web logs as a flat sequence of URLs, failing to:
- Distinguish between intentional forward navigation and corrective "back-button" sequences.
- Recognize that some pages are structurally more "authoritative" (hubs) regardless of immediate transition density.
- Identify when users are forced into "anomalous transitions" due to missing or empty information (e.g., checking next week's schedule before it's posted).
The authors' insight was that Web Usage Mining should be "active"—meaning it shouldn't just predict the next click, but should actively inform the administrator on how to restructure the site itself.
Methodology: Mining the Logic of Navigation
The researchers developed System L-R, which operates on two distinct but complementary layers:
1. Transition Probability Modeling
The system calculates the probability as the ratio of transitions from up to against all exits from . Critically, they preprocess the data to delete robot logs and "back-button" events to ensure they are modeling human intent rather than browser mechanics. They then introduced Weighted Probability, where the likelihood of a recommendation is boosted by the "Number of References"—the more pages that link to a target, the higher its "authority" weight.
2. User Path Clustering
To detect design flaws, they used Ward’s Method for clustering. By representing sessions as high-dimensional transition vectors (e.g., dimensions for a site with 138 pages), they could find groups of users exhibiting similar behaviors.

Experiments & Results: Bridging the Knowledge Gap
The system was tested on both a university department portal and a commercial cinema site.
- Quantitative Gains: All recommendation methods (2-5 in the chart below) outperformed the baseline "No Recommendation" (1). The Weighted by References method proved superior, especially in reducing the time required for complex tasks.
- The "Expert" vs. "Novice" Effect: Interestingly, the system was most effective for "out-campus" students (novices), significantly reducing the performance gap between them and "in-campus" experts who already knew the site structure.

- Identifying Design Flaws: Using collection-typed transition clustering on the cinema site, the authors discovered a cluster where users frequently jumped back and forth between "Next Road Show" and "Current Road Show." This revealed a flaw: users expected next week's schedule on Sunday, but the site didn't update until Tuesday, forcing users into a loop of confusion.
Critical Analysis & Conclusion
The value of this work lies in its holistic view of Web management. It treats user logs not just as a data stream for a black-box recommender, but as a diagnostic signal for human architects.
Takeaways:
- Noise Matters: Cleaning "back-button" data is essential for accurate intent modeling.
- Structure + Usage: Behavioral probability is more powerful when weighted by the site’s own link topology.
- Diagnostic Mining: Clustering is a powerful way to find "hidden" user frustration that aggregate statistics might miss.
Limitations & Future Work:
While groundbreaking for its time, the approach relies on relatively static HTML structures. Modern Web applications with dynamic content (AJAX/React) require a transition from URL-based mining to event-based mining. The authors began addressing this by discussing "Dynamic Page Recommendation," hinting at the future of real-time, adaptive interfaces.
