Integrating Domain Expertise: Hyperplane Classifiers with Prior Knowledge

Journal of Computational and Applied Mathematics

2022-01-01
X. Lei, Tongxiang Gu, S. Graillat, Hao Jiang, Jin Qi
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a family of Knowledge-Incorporated Multiple Criteria Linear Programming (MCLP) classifiers. It proposes a novel optimization framework that integrates both classified training samples and human expert prior knowledge—represented as linear or nonlinear logical implications—into MCLP and Kernel-based MCLP (KMCLP) models.

TL;DR

While most machine learning models are "black boxes" that learn only from raw data, this paper explores how to teach models to "listen" to experts. By transforming logical rules into linear programming constraints, the authors develop Knowledge-Incorporated MCLP, a classifier that balances empirical training data with human-derived expertise, leading to more accurate and interpretative results in complex tasks like cancer prognosis.

Background: The Limits of Empirical Learning

The majority of modern classifiers—Support Vector Machines (SVM), Neural Networks, and Decision Trees—operate on a purely empirical principle: they minimize error on a training set. However, this "data-only" approach has two major flaws:

  1. Data Scarcity: In fields like medicine or aerospace, high-quality labeled samples are rare and expensive.
  2. Noise Sensitivity: A few outliers can significantly tilt a discriminant hyperplane if the model has no "common sense" to override them.

Multiple Criteria Linear Programming (MCLP) is a robust optimization-based classifier that seeks a balance between maximizing distances from a boundary (MMD) and minimizing deviations (MSD). This paper takes MCLP to the next level by incorporating Prior Knowledge.

The Core Challenge: How to Speak "Math" to Logic?

The central innovation of this paper is the mathematical translation of expert rules.

  • The Rule: "If a tumor is large () and lymph nodes are affected (), then the cancer is likely to recur."
  • The Translation: This rule defines a region in space. Using the Farkas Theorem, the authors convert this logical implication into a set of linear inequalities that act as constraints in the optimization problem.

Architecture: From Linear to Kernel Space

For data that isn't linearly separable, the authors employ the Kernel Trick. By projecting data into a higher-dimensional space, they can incorporate not just linear rules, but also complex, nonlinear regions (like ellipsoids or spheres) as part of the classification logic.

Model Architecture and Geometric Logic In the figure above, Line (a) represents the original MCLP result, while Line (b) shows how the boundary shifts to respect the expert-defined knowledge sets (the triangle and rectangle).

Experimental Evidence

The authors validated their approach using three distinct scenarios:

  1. Synthetic Data: Proved that the decision boundary shifts predictably when knowledge sets are introduced.
  2. Checkerboard Data: Demonstrated that the parameter (the weight given to knowledge) allows the model to create much sharper, more accurate separation curves in nonlinear "xor-like" problems.
  3. Wisconsin Breast Cancer (WPBC): This real-world application is the highlight. By incorporating three nonlinear geometric rules (ellipses and triangles in feature space), the model improved its prediction of cancer recurrence significantly.

Real World Application - WPBC Data Visualization of nonlinear knowledge regions in the Breast Cancer dataset.

Key Performance Metrics

Attribute CombinationMCLP (No Knowledge)Knowledge-Incorporated MCLPImprovement
F1, F3, and F459.78%66.30%+6.52%
F3 and F457.61%63.04%+5.43%

Critical Insight: Why Does This Work?

The "magic" resides in the Objective Function. By adding a penalty term , the model is forced to minimize the violation of expert rules. If the data suggests a boundary that contradicts the expert, the parameter determines whether the model follows the data or the human. This provides a tunable Inductive Bias that is grounded in domain reality rather than just statistical randomness.

Conclusion & Future Outlook

This paper provides a rigorous framework for "Knowledge-Incorporated" learning. While modern Deep Learning often ignores these symbolic approaches, the methodologies presented here are crucial for Explainable AI (XAI). Future work could involve automatically extracting these rules from large-scale unlabelled data or using Large Language Models (LLMs) to generate the initial logical implications that feed into these linear programming solvers.

The primary limitation remains the manual definition of knowledge regions; however, as demonstrated, even imperfect expert rules can steer a model toward a much more robust solution than data alone.

Find Similar Papers

Try Our Examples

  • Search for recent papers that extend Multiple Criteria Linear Programming (MCLP) using deep learning architectures or neural networks.
  • What are the primary theoretical differences between Knowledge-Based Support Vector Machines (KBSVM) and the Knowledge-Incorporated KMCLP proposed in this paper?
  • Find studies that apply prior-knowledge-integrated classifiers to sparse data scenarios in medical diagnostics or financial fraud detection published after 2020.
Contents
Integrating Domain Expertise: Hyperplane Classifiers with Prior Knowledge
1. TL;DR
2. Background: The Limits of Empirical Learning
3. The Core Challenge: How to Speak "Math" to Logic?
3.1. Architecture: From Linear to Kernel Space
4. Experimental Evidence
4.1. Key Performance Metrics
5. Critical Insight: Why Does This Work?
6. Conclusion & Future Outlook