Are Unsupervised Neural Networks Ignorant? Sizing the Effect of Environmental Distributions
Are unsupervised neural networks ignorant? Sizing the effect of environmental distributions on unsupervised learning Action editor: Axel Cleeremans
This paper investigates the sensitivity of unsupervised neural networks to environmental distributions, specifically exploring how Recurrent Associative Memories (RAMs) and Hard Competitive Networks (HCNs) handle category and exemplar frequencies. The study demonstrates that while RAMs (like NDRAM) effectively mirror environmental prior odds, HCNs (like ART1) remain "ignorant" of these biases, often failing to optimize performance based on statistical distribution.
TL;DR
In a world of biases—where you're more likely to see a dog than a wolf in a city—rational agents must use frequency information to make decisions. This paper tests whether unsupervised neural networks are "ignorant" of these environmental biases. The verdict? Recurrent Associative Memories (RAMs) are statistically savvy, while Hard Competitive Networks are effectively "blind" to the frequency of the events they experience.
The "Ignorance" Problem: Why Frequency Matters
Laplace once defined "ignorance" as having uniform priors—treating every possibility as equally likely regardless of past experience. In machine learning, we strive for the opposite: Bayesian optimality, where prior odds help resolve ambiguity.
Psychological research confirms that humans are not "ignorant"; we are sensitive to how often categories and specific examples appear. If we use Artificial Neural Networks (ANNs) to model human cognition or to build autonomous AI, these models must be able to absorb environmental distributions. Yet, for decades, the two pillars of unsupervised learning—Competitive Networks and Associative Memories—have rarely been compared on this specific ability.
Methodology: The Orthonormal and the Complex
The researchers tested two families of networks:
- Recurrent Associative Memories (RAMs): Such as the Brain-State-in-a-Box or the newer NDRAM, which use Hebbian-style learning.
- Hard Competitive Networks (HCNs): Such as ART1 or Rumelhart’s competitive learners, which use a "winner-take-all" mechanism.
The Experiment Design
The authors subjected these networks to four environmental distributions:
- Uniform: Everything is equal (Control).
- Bimodal/Multimodal: Some specific examples appear more often, but categories are balanced.
- Step: One category is more frequent than the others.
- Exponential: Specific examples and specific categories are both highly biased.
Figure 1: (a) RAM architecture (b) Competitive Network architecture.
The Core Insight: How RAMs See Frequency
The study found that RAMs encode environmental information in two distinct, elegant ways:
- Category Frequency Eigenvalues: If Category A appears more often than Category B, the eigenvalue associated with A’s weight matrix grows larger. This effectively "stretches" the attractor basin for that category.
- Exemplar Frequency Eigenvectors: If a specific version of a "7" is shown more often, the attractor (eigenvector) for that category rotates toward that specific stimulus.
Equation: Bayesian posterior probability serves as the normative benchmark for these models.
The Failure of Hard Competition
The results for Hard Competitive Networks (like ART1) were startling. Despite being exposed to highly biased data, these networks consistently performed as if the environment were uniform.
Why the failure? The authors point to the "Min/Max" decision function. By only updating the winning unit and discarding the magnitude of the activation or the relative competition between units, the network "compresses" the output space. All information regarding the frequency of input—the "dilatations" of weight vectors—is lost.
Table 1: Classification of random vectors. Note how RAMs shift their categorization toward the biased 'A' category in the Step and Exponential conditions, while Competitive Networks remain near a 50/50 split.
Critical Analysis & Conclusion
This paper delivers a significant blow to the use of hard competitive learning for cognitive modeling. If a model cannot reflect that "Category A occurs 5 times more often than B," it cannot simulate human categorization behavior or perform optimal machine inference.
Main Takeaways:
- RAMs (NDRAM) are biologically and mathematically superior for tasks requiring sensitivity to statistical distributions.
- Hard Competitive Networks suffer from a structural "blindness" caused by their winner-take-all logic.
- Future Direction: To fix competitive networks, researchers should look toward Soft Competitive Learning, where multiple units are updated, preserving the "topology" and "density" of the input space.
In the quest for models that aren't "ignorant," the Associative Memory family currently holds a significant lead.
