GRMAC: Winning the "AI Meets Beauty" Challenge with Flexible Attention

Beauty Product Retrieval Based on Regional Maximum Activation of Convolutions with Generalized Aention

Jun Yu
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces Generalized-attention Regional Maximum Activation of Convolutions (GRMAC), a novel image descriptor tailored for beauty and personal care product retrieval. By integrating a flexible attention mechanism with multi-model feature fusion, the method achieved 1st place in the ACM Multimedia 2019 "AI Meets Beauty" Grand Challenge with a MAP@7 score of 0.4086.

TL;DR

Determining the "right" product in a database of half a million beauty items is notoriously difficult due to cluttered backgrounds and subtle packaging differences. The USTC NELSLIP team secured 1st place at ACM MM 2019 by introducing GRMAC, a descriptor that uses a tunable hyperparameter to "squeeze" out background noise, paired with a strategic fusion of global and regional deep features.

Background: Why Beauty Product Retrieval is Hard

In traditional image retrieval (like landmarks), the geometry is often rigid. In beauty products, however, a bottle of lotion might look entirely different depending on the camera angle, or it might be surrounded by dozens of other items on a shelf.

Previous SOTA methods like RMAC (Regional Maximum Activation of Convolutions) extract features from various local regions and sum them up. The flaw? It treats the chaotic background and the actual product with the same level of importance.

Methodology: The Power of the Generalized Attention Mask

The core innovation is the Generalized-attention Regional Maximum Activation of Convolutions (GRMAC).

1. Beyond Mean Thresholding

Earlier attention-based descriptors used the "mean value" of activations to decide which regions to keep. However, mean values are sensitive to outliers. GRMAC introduces a hyperparameter to calculate a more robust threshold : By tuning , researchers can control how "strict" the filter is. As increases, the mask becomes more selective, focusing only on the most salient parts of the product.

2. The Framework

The team utilized a dual-backbone approach:

  • DenseNet201: Known for its feature reuse capabilities.
  • SE-ResNet152: Leveraging Squeeze-and-Excitation blocks to weigh channel-wise importance.

Model Architecture Figure 1: The GRMAC pipeline—from raw image to fused feature vector.

Experiments & Results

The authors tested their method on the Perfect-500K dataset.

The Impact of Hyperparameter

A critical discovery was that the optimal value of varies by model. For SE-ResNet architectures, a specific effectively "cleans up" the mask, whereas a (equivalent to the older RAMAC method) often includes too much background noise.

Mask Visualization Figure 2: Visual proof—as p increases (left to right), the attention mask focuses more tightly on the product.

Feature Fusion: The Winning Formula

The team found that Feature Fusion is the secret sauce. Specifically, combining a Global descriptor (MAC) with a Regional descriptor (GRMAC) yielded the best results. This suggests that while regional features capture fine-grained details, global features provide necessary context that prevents the model from "over-focusing" on small parts of the packaging.

Descriptor PairResult (MAP)
DenseNet201 (GRMAC) + SE-ResNet152 (MAC)0.3877 (Validation)
Final Competition Score (Test)0.4086

Critical Analysis & Conclusion

Takeaway

GRMAC proves that attention in image retrieval doesn't have to be a complicated "black box" transformer layer. Sometimes, a mathematically sound, tunable thresholding mechanism applied to convolutional feature maps is more effective and computationally efficient for large-scale indexing.

Limitations & Future Work

While the method won the competition, the hyperparameter currently requires manual tuning (Ablation). A logical next step would be an Adaptive GRMAC where is learned per image or per category using a small policy network, allowing the model to decide how much background suppression is needed based on the image's complexity.


Author Insight: This work highlights a shift in CBIR (Content-Based Image Retrieval)—moving away from simple global pooling toward "Smarter Aggregation" that understands object-background relationships.

Find Similar Papers

Try Our Examples

  • Find recent papers that improve the RMAC or GeM pooling descriptors for fine-grained product retrieval tasks beyond 2020.
  • Which study first introduced the concept of Regional Maximum Activation of Convolutions (RMAC), and how does the attention mechanism in this paper specifically modify that original architecture?
  • Explore how dynamic attention mask techniques similar to GRMAC are being applied in other vision tasks like object re-identification or medical image indexing.
Contents
GRMAC: Winning the "AI Meets Beauty" Challenge with Flexible Attention
1. TL;DR
2. Background: Why Beauty Product Retrieval is Hard
3. Methodology: The Power of the Generalized Attention Mask
3.1. 1. Beyond Mean Thresholding
3.2. 2. The Framework
4. Experiments & Results
4.1. The Impact of Hyperparameter $p$
4.2. Feature Fusion: The Winning Formula
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work