Cosine Loss: The Secret Weapon for Small Data Deep Learning

Deep learning on small datasets without pre-training using cosine loss

Björn Barz, Joachim Denzler
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces the Cosine Loss as a superior alternative to categorical cross-entropy for training deep neural networks from scratch on small datasets (typically <100 samples per class). By maximizing cosine similarity between feature vectors and class embeddings, the authors achieve State-of-the-Art performance on several fine-grained benchmarks without relying on ImageNet pre-training.

TL;DR

The long-standing dogma in Deep Learning is that "Large Data + Cross-Entropy = Success." This paper challenges that status quo, proving that for small datasets (under 100 samples per class), the Cosine Loss function is significantly more effective than the standard Softmax+Cross-Entropy combo. By simply focusing on the angle of features rather than their magnitude, the authors achieved a staggering 30% accuracy boost on the CUB-200 dataset without any ImageNet pre-training.

The Problem: The Overfitting Trap of Cross-Entropy

In the academic world, we often take ImageNet pre-training for granted. However, in industrial and specialized medical fields, pre-training is a minefield:

  • Legal Risks: ImageNet's license is gray for commercial use.
  • Domain Shift: Pre-training on "cats and dogs" doesn't help much with "satellite infrared imagery" or "ultrasound scans."
  • Mathematical Vulnerability: Categorical cross-entropy rewards the model for pushing feature magnitudes to infinity to minimize loss, which is a one-way ticket to Overfitting when data is scarce.

Methodology: High-Dimensional Geometry as a Regularizer

The core insight is beautifully simple: L2 Normalization.

While Cross-Entropy operates in an Euclidean prediction space (or probability space after Softmax), the Cosine Loss projects features onto a Unit Hypersphere.

Why does this work?

  1. Invariance to Scale: In high-dimensional spaces, the "magnitude" of a vector often represents noise or training artifacts. By L2-normalizing, the model only learns the direction (semantic meaning).
  2. Boundedness: Unlike Cross-Entropy, which can be infinitely large, Cosine Loss is strictly bounded . This prevents outliers or mislabeled samples from generating massive gradients that "explode" the model's weights.
  3. Hierarchy Integration: It allows for "Semantic Embeddings" where class prototypes aren't just one-hot vectors, but vectors whose angles reflect real-world relationships (e.g., a "Sparrow" vector is closer to a "Finch" than a "Car").

Model Architecture and Loss Comparison Figure 1: Comparison of loss landscapes. Note how Cosine Loss (c) provides a more uniform gradient landscape compared to the steep, narrow optima of Cross-Entropy (a).

Experiments: Breaking the "No Pre-training" Barrier

The authors tested their hypothesis across Fine-Grained Visual Categorization (FGVC) tasks and even Text Classification.

Key Breakthroughs:

  • CUB-200 (Birds): Jumped from 51.9% (Cross-Entropy) to 67.6% (Cosine Loss).
  • AG News (Text): Significant gains when training on only 10-25 samples per class.
  • The "Small Data" Threshold: The experiments show that Cosine Loss is superior until you hit roughly 200 samples per class, at which point the flexibility of Cross-Entropy begins to catch up.

Performance Comparison Table Table 1: The performance gap is most evident in "from scratch" scenarios where pre-trained weights aren't used.

Dataset Size Effect Figure 2: The trend is clear: the fewer the samples, the larger the "Cosine Advantage."

Critical Insight & Conclusion

This paper serves as a vital reminder that our "standard" tools (like Cross-Entropy) were optimized for the "Big Data" era. When we step into the "Small Data" regime—which is the reality for most specialized industries—we need to rethink the fundamental geometry of our loss functions.

Takeaway for Practitioners: If you are struggling with a dataset of only a few thousand images and cannot use a pre-trained model, stop tweaking your learning rate and try switching to Cosine Loss. It provides a hyperparameter-free regularization that might be the key to convergence.

Limitations: While Cosine Loss is a powerhouse for small data, it loses its edge once datasets become massive (like CIFAR-100 or ImageNet), where the model has enough signal to overcome the noise in feature magnitudes.

Find Similar Papers

Try Our Examples

  • Which recent papers explore "logit scaling" or "temperature tuning" in Cosine-based losses to bridge the gap between small and large dataset performance?
  • Trace the origin of "L2-constrained Softmax" and compare how subsequent works like ArcFace or CosFace differ from the direct Cosine Loss proposed here.
  • Search for studies applying Cosine Loss to non-visual modalities such as time-series sensor data or small-scale medical MRI classification tasks.
Contents
Cosine Loss: The Secret Weapon for Small Data Deep Learning
1. TL;DR
2. The Problem: The Overfitting Trap of Cross-Entropy
3. Methodology: High-Dimensional Geometry as a Regularizer
3.1. Why does this work?
4. Experiments: Breaking the "No Pre-training" Barrier
4.1. Key Breakthroughs:
5. Critical Insight & Conclusion