ELM: Bridging the Gap in Knowledge Mining through Aspect-Oriented Instrumentation
System for Knowledge Mining in Data from Interactions between User and Application
The paper introduces ELM (Event Logger Manager), a flexible Web Usage Mining (WUM) system designed to register and analyze user interaction data from Java applications. By leveraging Aspect-Oriented Programming (AspectJ) and the Weka library, ELM provides a modular framework for knowledge discovery tasks such as service personalization and behavior modeling.
TL;DR
The Event Logger Manager (ELM) is a sophisticated Web Usage Mining (WUM) framework that simplifies the extraction of user behavior patterns from Java-based applications. By employing Aspect-Oriented Programming (AOP), it allows developers to "inject" logging capabilities into compiled code, subsequently feeding this data into a suite of powerful mining algorithms including Apriori, ID3, and C4.5.
The Evolution of Web Usage Mining (WUM)
In the digital economy, understanding why and how a user interacts with an application is as critical as the application's core logic. Web Usage Mining (WUM) traditionally relies on analyzing server logs. However, these logs are often noisy, incomplete, or lack the "semantic context" of specific user actions.
Existing systems like Analog and SUGGEST pioneered this field but faced a common bottleneck: they were either hard-wired to specific server environments or lacked the flexibility to adapt to custom logical events. This is where ELM introduces a paradigm shift.
Architectural Innovation: Non-Invasive Logging
The core "secret sauce" of ELM is its use of AspectJ. Typically, to monitor specific user events, a developer would need to manually insert logging code across thousands of lines of source code—an error-prone and maintenance-heavy task.
ELM bypasses this by using Aspect Modification. This technique allows the system to:
- Intercept method calls at runtime or compile-time.
- Define Pointcuts that specify exactly which interactions (e.g., button clicks, transaction starts) should be recorded.
- Store events in a centralized relational database (eventDB) via a streamlined JDBC API.

The Mining Pipeline: From Raw Events to Insight
ELM is not just a logger; it is a full-stack discovery environment. The process follows a rigorous four-stage pipeline:
1. Data Acquisition (EventLogger)
Unlike systems that only look at one side of the interaction, ELM can handle data registered at the server level and user-defined logical events.
2. Preprocessing (ArffBrowser)
Data Mining algorithms are "picky" about their inputs. ELM transforms relational data into ARFF (Attribute Relation File Format), the standard for the Weka machine learning library. The ArffBrowser module permits researchers to group rows, merge parameters, and filter out noise before training begins.
3. Pattern Discovery (AlgorithmRunner)
ELM integrates several classic yet robust algorithms:
- Apriori: For finding association rules (e.g., "Users who visited Page A also clicked Button B").
- ID3 & C4.5: For building decision trees to classify user sessions into categories (e.g., "High-Value Customer" vs "Window Shopper").

Experimental Workflow & Visualization
The system provides a GUI-driven experience for managing experiments. The EventBrowser allows analysts to filter logs by time and session, while the AlgorithmMonitor tracks the execution of long-running mining tasks.

By comparing results across different periods, experts can identify trends—such as a sudden decrease in site structure efficiency or an improvement in conversion rates after a UI change.
Critical Analysis & Future Outlook
ELM’s greatest strength is its modularity. By decoupling the logger from the manager, it enables a "plug-and-play" ecosystem for algorithms.
Limitations:
- Currently, the system is primarily optimized for the Java ecosystem through AspectJ.
- The use of traditional relational databases like eventDB might face performance bottlenecks when dealing with "Big Data" scales (e.g., millions of events per second), where NoSQL or stream-processing frameworks would be more appropriate.
Final Takeaway: ELM represents a robust bridge between software engineering (AOP) and data science (WUM). It proves that the most valuable data often lies in the "logical events" defined by developers, rather than just the "physical logs" produced by servers.
