Trust in the Machine: Why Applied Data Science is More About People Than Algorithms

Trust in Data Science: Collaboration, Translation, and Accountability in Corporate Data Science Projects

2020-02-09
Samir Passi, Steven J. Jackson
Summary
Problem
Method
Results
Takeaways
Abstract

This ethnographic study investigates how trust is established in corporate data science through 6 months of fieldwork at a major e-commerce firm. It introduces the concepts of "algorithmic witnessing" and "deliberative accountability" to describe how human translation—not just math—validates SOTA models in real-world business settings.

TL;DR

In corporate boardrooms, "Accuracy" is a plastic number, and "Black Boxes" are tolerated only if they tell a good story. This paper by Passi and Jackson (2018) moves beyond the "clean room" view of data science, revealing through 6 months of ethnographic fieldwork that trust is a collaborative performance, negotiated between skeptical business analysts and data scientists through "Deliberative Accountability."

The "Blind Spot" of Applied AI

While academic papers obsess over 0.1% gains in F1-scores, the corporate world struggles with a more fundamental problem: Why should a manager trust a model that contradicts their 20 years of experience?

Existing research often treats data science as a purely technical pipeline. However, this paper argues there is a "blind spot" in AI research regarding how systems are actually managed in-use. The core challenge isn't just about making models more accurate—it's about managing the friction between data-led objectivity and human interpretive frameworks.

The Four Tensions of Data Science Work

The authors identify four structural tensions that arise when "math meets the market":

  1. (Un)equivocal Numbers: Metrics like "80% Accuracy" don't mean the same thing to a Data Scientist (technical performance) as they do to a Business Analyst (absolute confidence).
  2. (Counter)intuitive Knowledge: If a model confirms what we know, it’s "obvious." If it contradicts our gut, it’s either "broken" or a "novel discovery." Deciding which is which requires intense social negotiation.
  3. (In)credible Data: Data is never "raw." It is often sparse, inconsistent (e.g., credit scores from different vendors), and requires discretionary "repair work" that is rarely documented.
  4. (In)scrutable Models: SOTA models like Neural Networks are black boxes. In corporate settings, this opacity is often bypassed by "Implicit Trust" backed by "Explicit Verification" (pilot testing).

需替换为四种张力示意图

Deep Dive: The Churn Prediction Case

In one of the two featured case studies, a data science team struggled with low accuracy (20-30%) for predicting customer churn. In a research setting, this would be a failure. In the corporate setting of DeepNetwork, it was a negotiation:

  • Decomposing the Error: Instead of one "Accuracy" score, the team broke it down into True Positives versus False Positives.
  • Shifting Goals: They moved from binary "Yes/No" labels to "Likely-to-Churn Probabilities." This allowed the business team to act on a prioritized list rather than demanding a "silver bullet" correct answer.
  • The Power of Gut Feeling: When the model identified "feature importances" that matched the manager's intuition, trust plummeted into the system—even if the actual accuracy was low.

Algorithmic Witnessing vs. Deliberative Accountability

One of the paper's most profound contributions is the distinction between two ways of holding AI accountable:

1. Algorithmic Witnessing

This is the technical side. It’s what data scientists do when they cross-validate, check for overfitting, and reproduce results. It is the "mechanical objectivity" of the field.

2. Deliberative Accountability

This is the social side. It involves stakeholders "seeing" meaning in the results via narratives. It’s about justifying why the rules are the rules through collaborative meetings and pilot tests.

需替换为实验结果/协作流程图

Critical Insight: The "Alien" Nature of AI

The authors argue that models aren't just "opaque"; they are alien. Their logic is inherently foreign to human reasoning. To fix this, we don't just need a "Right to Explanation"; we need a "Right to Justification."

The job of a data scientist, therefore, is not just to "crunch numbers," but to act as a translator. They must narrativize the "mess" of data into a story that aligns with (or convincingly challenges) organizational reality.

Takeaways for the Industry

  • Curriculum Shift: Data science education needs to move beyond Python and Math to include "Interactional Expertise" (how to talk to people who sign the checks).
  • Documentation 2.0: We must document discretionary decisions (e.g., "Why did we delete these 5,000 rows?"), not just the final algorithm.
  • Embrace the Mess: Trust doesn't come from perfection; it comes from a "reasoned practicality."

Conclusion

Trust in Data Science is not a software feature you can download from GitHub; it is a deeply collaborative accomplishment. As the field matures, our ability to translate and negotiate will be just as important as our ability to optimize a loss function.

Find Similar Papers

Try Our Examples

  • Find recent Computer-Supported Cooperative Work (CSCW) papers that investigate the role of "narrativization" in making AI results explainable to non-technical stakeholders.
  • What are the latest SOTA frameworks for "Human-Centered Data Science" that specifically address the "Right to Explanation" under GDPR in corporate settings?
  • Which studies bridge the gap between "contributory expertise" (data science) and "interactional expertise" (business management) within AI development teams?
Contents
Trust in the Machine: Why Applied Data Science is More About People Than Algorithms
1. TL;DR
2. The "Blind Spot" of Applied AI
3. The Four Tensions of Data Science Work
4. Deep Dive: The Churn Prediction Case
5. Algorithmic Witnessing vs. Deliberative Accountability
5.1. 1. Algorithmic Witnessing
5.2. 2. Deliberative Accountability
6. Critical Insight: The "Alien" Nature of AI
7. Takeaways for the Industry
8. Conclusion