Augmented Lineage: Bridging the Gap Between Data Provenance and AI Explainability

Augmented Lineage: Traceability of Data Analysis Including Complex UDFs

2021-01-01
Masaya Yamada, Hiroyuki Kitagawa, Toshiyuki Amagasa, Akiyoshi Matono
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces "Augmented Lineage," a novel framework for data traceability in complex pipelines that integrate relational operators with User-Defined Functions (UDFs) for AI/ML tasks. By extending traditional tuple-level lineage with "Reasoning Lineage" (RL), the method captures the decision basis of AI models (e.g., image bounding boxes or feature importance) alongside source data origin.

TL;DR

As data analysis evolves from simple SQL queries to complex AI-driven pipelines, knowing which data was used is no longer enough—we need to know why the model made its choice. This paper introduces Augmented Lineage, a framework that captures both the physical data path and the logical reasoning of ML models (UDFs). By utilizing a smart Function Materialization strategy, the authors achieve high-speed traceability with 90% less storage overhead than traditional methods.

The Motivation: When "Source" is Not Enough

In traditional database systems, Lineage (or Provenance) answers the question: "Which rows in the database led to this result?"

However, imagine a bank using an ML model to reject a loan. The traditional lineage tells you which customer records were accessed. But it doesn't tell you that the model rejected the loan specifically because the "Debt-to-Income" ratio was the deciding factor. In the age of AI and unstructured data (images, video), we need the Reason—the specific pixels or features that triggered an output.

The Core Concept: Augmented Lineage (AL)

The authors redefine lineage as a pair: .

  1. Source Lineage (SL): The set of original tuples from source tables.
  2. Reasoning Lineage (RL): The logic provided by the UDF (e.g., "The model looked at these coordinates in Figure A").

Model Architecture and Segments Figure: The data analysis is modeled as an operator tree, where complex AI logic is encapsulated in Function Operators ().

Methodology: Balancing Speed and Storage

Traceability usually faces a "trilemma" between:

  • Lazy (Rerun): No storage cost, but slow because you have to re-compute expensive AI models.
  • Eager (Full Materialization): Instant tracing, but massive storage waste for every intermediate step.
  • Function Materialization (FM) [Proposed]: The "Goldilocks" solution. It only stores the results of the expensive UDFs (like deep learning models) and reruns the cheap relational operators (like Joins and Filters).

The Algorithm

The process involves two main steps:

  1. Canonicalization: Reordering the operator tree (pulling up selections/projections) to minimize the work needed during tracing.
  2. Tracing Query: Using semi-joins () to navigate backward from the output to the source, capturing "Reason" tags whenever a UDF node is hit.

Experimental Evidence

The authors tested their method using a Person Recognition task (Expensive UDF) vs. a String Processing task (Cheap UDF).

Experimental Performance Comparison Figure: Performance Comparison. For expensive ML models, FM provides a massive speedup over Rerunning while keeping storage manageable.

Key Metrics:

  • Storage Efficiency: FM materialized only tuples compared to for Full Materialization (a ~91% reduction).
  • Speed: In "Medium" datasets, FM was over 500x faster than the Rerun approach for AI tasks.

Critical Analysis & Conclusion

The value of this work lies in its pragmatism. By acknowledging that AI models are the bottleneck in data analysis, the authors focused optimization where it matters most: the UDF transition points.

Strengths:

  • Synthesizes database theory with XAI requirements.
  • Introduces a cost-effective storage strategy (FM).

Limitations & Future Work:

  • The framework currently assumes UDFs are designed to output "Reasons." In reality, many legacy models are black boxes that don't natively support this.
  • Future research should look into automated reason generation (e.g., automatically applying LIME or Grad-CAM during the lineage capture phase) and scaling this to real-time stream processing where data moves too fast for traditional materialization.

Final Takeaway: Augmented Lineage represents the next step in "Responsible AI," ensuring that every automated decision remains auditable, explainable, and traceable back to its roots.

Find Similar Papers

Try Our Examples

  • Find other recent papers that integrate Explainable AI (XAI) techniques, such as SHAP or Grad-CAM, directly into SQL-based data provenance systems.
  • What are the foundational papers on "fine-grained provenance" and how does the concept of augmented lineage differ from traditional "Why-provenance" in relational databases?
  • Investigate how the Function Materialization (FM) strategy could be applied to distributed stream processing frameworks like Apache Flink or Spark Streaming for real-time model auditing.
Contents
Augmented Lineage: Bridging the Gap Between Data Provenance and AI Explainability
1. TL;DR
2. The Motivation: When "Source" is Not Enough
3. The Core Concept: Augmented Lineage (AL)
4. Methodology: Balancing Speed and Storage
4.1. The Algorithm
5. Experimental Evidence
6. Critical Analysis & Conclusion