Augmented Lineage: Bridging the Gap Between Data Provenance and AI Explainability
Augmented Lineage: Traceability of Data Analysis Including Complex UDFs
The paper introduces "Augmented Lineage," a novel framework for data traceability in complex pipelines that integrate relational operators with User-Defined Functions (UDFs) for AI/ML tasks. By extending traditional tuple-level lineage with "Reasoning Lineage" (RL), the method captures the decision basis of AI models (e.g., image bounding boxes or feature importance) alongside source data origin.
TL;DR
As data analysis evolves from simple SQL queries to complex AI-driven pipelines, knowing which data was used is no longer enough—we need to know why the model made its choice. This paper introduces Augmented Lineage, a framework that captures both the physical data path and the logical reasoning of ML models (UDFs). By utilizing a smart Function Materialization strategy, the authors achieve high-speed traceability with 90% less storage overhead than traditional methods.
The Motivation: When "Source" is Not Enough
In traditional database systems, Lineage (or Provenance) answers the question: "Which rows in the database led to this result?"
However, imagine a bank using an ML model to reject a loan. The traditional lineage tells you which customer records were accessed. But it doesn't tell you that the model rejected the loan specifically because the "Debt-to-Income" ratio was the deciding factor. In the age of AI and unstructured data (images, video), we need the Reason—the specific pixels or features that triggered an output.
The Core Concept: Augmented Lineage (AL)
The authors redefine lineage as a pair: .
- Source Lineage (SL): The set of original tuples from source tables.
- Reasoning Lineage (RL): The logic provided by the UDF (e.g., "The model looked at these coordinates in Figure A").
Figure: The data analysis is modeled as an operator tree, where complex AI logic is encapsulated in Function Operators ().
Methodology: Balancing Speed and Storage
Traceability usually faces a "trilemma" between:
- Lazy (Rerun): No storage cost, but slow because you have to re-compute expensive AI models.
- Eager (Full Materialization): Instant tracing, but massive storage waste for every intermediate step.
- Function Materialization (FM) [Proposed]: The "Goldilocks" solution. It only stores the results of the expensive UDFs (like deep learning models) and reruns the cheap relational operators (like Joins and Filters).
The Algorithm
The process involves two main steps:
- Canonicalization: Reordering the operator tree (pulling up selections/projections) to minimize the work needed during tracing.
- Tracing Query: Using semi-joins () to navigate backward from the output to the source, capturing "Reason" tags whenever a UDF node is hit.
Experimental Evidence
The authors tested their method using a Person Recognition task (Expensive UDF) vs. a String Processing task (Cheap UDF).
Figure: Performance Comparison. For expensive ML models, FM provides a massive speedup over Rerunning while keeping storage manageable.
Key Metrics:
- Storage Efficiency: FM materialized only tuples compared to for Full Materialization (a ~91% reduction).
- Speed: In "Medium" datasets, FM was over 500x faster than the Rerun approach for AI tasks.
Critical Analysis & Conclusion
The value of this work lies in its pragmatism. By acknowledging that AI models are the bottleneck in data analysis, the authors focused optimization where it matters most: the UDF transition points.
Strengths:
- Synthesizes database theory with XAI requirements.
- Introduces a cost-effective storage strategy (FM).
Limitations & Future Work:
- The framework currently assumes UDFs are designed to output "Reasons." In reality, many legacy models are black boxes that don't natively support this.
- Future research should look into automated reason generation (e.g., automatically applying LIME or Grad-CAM during the lineage capture phase) and scaling this to real-time stream processing where data moves too fast for traditional materialization.
Final Takeaway: Augmented Lineage represents the next step in "Responsible AI," ensuring that every automated decision remains auditable, explainable, and traceable back to its roots.
