Attribute Transformations: Why Your Raw Database is Useless for Data Mining
Attribute Transformations on Numerical Databases. Applications to Stock Market and Economic Data
This paper explores the systematic use of attribute transformations—linear, quadratic, and non-linear—to enhance data mining on numerical databases. Using Rough Set Theory as the analytical framework, the authors demonstrate that transforming raw "record-keeping" attributes into predictive features enables the extraction of actionable rules for stock market and economic forecasting.
TL;DR
Data mining often fails because database schemas are designed for storage, not discovery. This seminal work argues that Attribute Transformations—ranging from simple moving averages to complex quadratic discriminants—are the "missing link" in extracting knowledge. By applying Rough Set Theory to transformed stock market data, the authors demonstrate how to turn raw price logs into high-probability predictive rules.
Context: The Representation Gap
In AI and Data Mining, the success of an algorithm is rarely about the model itself, but about the features it consumes. The authors point out a fundamental divide:
- Scientific Research: Attributes are often derived or unobservable (e.g., forces, latent variables).
- Business Databases: Attributes are chosen for "record-keeping" (e.g., closing price, date).
To bridge this gap, the paper introduces a systematic way to transform "record-keeping" data into "knowledge-discovery" data.
Methodology: The Geometry of Data
The authors define a numerical -relation as a set of points lying on a "hyper-surface" or manifold in Eucledian space. The goal of attribute transformation is essentially a change of coordinate system to find one where the patterns become linear or logically separable.
1. Types of Transformations
- Linear (Averages/Sums): Essential for smoothing noise in time-series data.
- Quadratic (The Discriminant): The authors use the example of Conic Sections. You cannot classify an ellipse vs. a hyperbola by looking at raw coefficients. You need the transformation .
- Geometric (Distance Formulas): Converting coordinates into side lengths to classify triangles.
2. The Power of Polynomials
A key theoretical contribution is the Polynomial Approximation Theorem: Since any information table is finite, any possible transformation can be represented as a polynomial. This provides a "brutal force" pathway for discovery: if a domain expert doesn't know the rule, a computer can search through polynomial degrees () until a pattern emerges.
The mathematical representation of a tuple in an information table, allowing for non-unique object descriptions—a core tenet of Rough Set Theory.
Experiments: Hacking the Stock Market
The authors tested their hypothesis on 6 years of stock data (1990-1996), specifically targeting Applied Materials.
They applied three specific programs:
- Delay: Shifting data to create time-lagged features.
- Average: N-day moving averages to capture trends.
- Sum: Cumulative changes over specific windows.
Key Predictive Rules Found:
- Falling Rule: If the 5-day average of the Equipment Index is and Texas Instruments (TI) is , the stock likely falls (62.2% support).
- Rising Rule (The High Performer): If the 2-day sum of the Equipment Index is , the stock rises, showing a 80.5% accuracy in validation.
Validation results showing that rules with high learning support (Rules 3, 5, 11) maintained their predictive power in unseen data.
Critical Insight: The "Reduct"
The paper utilizes Rough Set Theory to perform data reduction. Unlike standard relational databases that focus on storage, Rough Sets focus on finding the Reduct—the minimal set of attributes that preserves the decision-making power of the original dataset. This "compactification" is what allows for the generation of human-readable rules.
Summary & Limitations
This work serves as a foundational bridge between database theory and predictive modeling.
- Limitations: The authors admit that choosing the "discretization range" (e.g., what constitutes "Rising" vs "No Change") is still a matter of trial and error (Heuristics).
- Takeaway: Attribute transformation is not just "pre-processing"; it is the act of defining the search space for intelligence. Linear transformations like averaging are powerful, but the future of the field lies in discovering non-linear manifolds through automated polynomial search.
