Beyond Ad-Hoc Robustness: Building the Generalized S-Divergence Family via Model Adequacy

A New Family of Divergences Originating From Model Adequacy Tests and Application to Robust Statistical Inference

2018-01-17
Abhik Ghosh, Ayanendranath Basu
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces the Generalized S-Divergence (GSD) family, a three-parameter () superfamily of statistical divergences derived from tubular model adequacy tests (MATs). The GSD provides a unified framework that encompasses key prior families including the Power Divergence, Density Power Divergence, and Generalized Kullback-Leibler (GKL) families, achieving state-of-the-art trade-offs between asymptotic efficiency and outlier robustness.

TL;DR

Statistical inference often faces a tug-of-war between efficiency (doing well on clean data) and robustness (not breaking on dirty data). Abhik Ghosh and Ayanendranath Basu introduce the Generalized S-Divergence (GSD) family, a powerful three-parameter toolkit derived systematically from Model Adequacy Tests. This new family proves that the best tools for robust estimation often lie outside traditional, well-studied boundaries like the Density Power Divergence (DPD).

The "Why": The Failure of Classical Intuition

For decades, the Maximum Likelihood Estimator (MLE) has been the gold standard for efficiency. However, a single outlier can destroy an MLE-based model. Researchers created "divergences"—mathematical measures of distance between distributions—to build more resilient models. Families like the Cressie-Read Power Divergence (PD) and Density Power Divergence (DPD) were hits, but their creation was often motivated by empirical success rather than a unified theory.

Furthermore, the paper highlights a critical flaw in academic "Standard Operating Procedures": First-order Influence Function (IF) analysis. Many robust estimators sharing the same parameter look identical under IF analysis, yet perform wildly differently in the real world. We need a more systematic way to generate and distinguish these tools.

The "How": Model Adequacy as a Generator

The authors pivot from "guessing" new divergences to "testing" for them. The core idea is the Model Adequacy Test (MAT). Instead of asking "Does the data perfectly fit the model?", MATs ask: "Is the model adequate within a certain tolerance level ?"

By framing the search for an estimator as a constrained optimization problem:

  1. Minimize a divergence between the true density and a surrogate.
  2. Subject to the surrogate being "adequate" for the model.

Through this Lagrange-multiplier-style logic, the authors derive the GSD family:

Need to replace with Figure 1 showing IF for Poisson and Normal models Figure 1: Comparison of Influence Functions. Note how multiple curves overlap at the model, proving that first-order analysis cannot distinguish between varied robustness profiles.

Key Results: Finding the "Sweet Spot"

The paper doesn't just theorize; it executes. Through simulations on Poisson data contaminated with 10% outliers, the authors reveal that:

  • MLE fails spectacularly under contamination.
  • Subfamily members (like DPD) are better, but not optimal.
  • The "winners" are often GSD members where is small () but and are tuned to specific non-zero values.

Experimental Results Comparison Table 3: Bias and MSE under 10% contamination. The green-shaded regions in the author's analysis highlight that the GSD family provides lower Bias and MSE than traditional DPD or PD estimators.

Critical Insight: The Limitation of First-Order Robustness

One of the most profound takeaways is the warning regarding First-order IF. The paper demonstrates that two estimators can have the same bounded IF but vastly different finite-sample robustness. This occurs because the second-order terms in the Taylor expansion (Von Mises expansion) of the functional can dominate the behavior. The GSD family allows researchers to tune parameters specifically to manage these higher-order sensitivities.

Conclusion & Future Work

The GSD family represents a significant step toward a "Unified Field Theory" of density-based divergences. While the mathematical overhead is higher, the reward is an estimator that retains high efficiency while becoming nearly impervious to typical data contamination.

Future Outlook: The next frontier is applying GSD to continuous high-dimensional data and exploring the topological properties of the parameter space to identify exactly how many "unique" divergences this superfamily actually contains.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize second-order influence function analysis to evaluate the B-robustness of M-estimators and minimum divergence methods.
  • Which paper first introduced the concept of tubular model adequacy tests for multinomial models, and how does the current work generalize it to continuous density functions?
  • Explore if the Generalized S-Divergence (GSD) framework has been applied to high-dimensional machine learning tasks such as robust neural network training or GAN objective functions.
Contents
Beyond Ad-Hoc Robustness: Building the Generalized S-Divergence Family via Model Adequacy
1. TL;DR
2. The "Why": The Failure of Classical Intuition
3. The "How": Model Adequacy as a Generator
4. Key Results: Finding the "Sweet Spot"
5. Critical Insight: The Limitation of First-Order Robustness
6. Conclusion & Future Work