SCNN: Decoding the Silent Language of Children's Drawings via Shallow Learning

Children's Drawing Psychological Analysis using Shallow Convolutional Neural Network

2020-11-01
Yue Yuan, Jing Huang, Xiang Ma, Ke Yan
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a specialized Shallow Convolutional Neural Network (SCNN) designed for the psychological analysis of children's tree drawings. By employing multi-label classification with a Sigmoid activation layer, the model identifies six distinct mental health issues (e.g., anxiety, insecurity) with an accuracy of 70.4%.

TL;DR

Psychologists have used tree drawings to peer into the human subconscious for nearly a century. This paper bridges psychology and AI by introducing a Shallow Convolutional Neural Network (SCNN) specifically tuned for the abstract nature of children's sketches. By identifying multiple psychological markers like anxiety and aggressiveness simultaneously, the model achieves a 70.4% accuracy rate, offering a promising scalable tool for early mental health intervention.

Problem & Motivation: The "Sparse Data" Challenge

In modern psychology, the "Tree Drawing Test" is a classic projective tool. Why trees? Because they allow children to project deep emotions with fewer defense mechanisms than verbal communication. However, translating this to AI presents a unique paradox:

  1. Feature Scarcity: Unlike ImageNet photos, children's drawings are mostly "blanks" with no texture or complex color gradients.
  2. Multi-Dimensionality: A child isn't just "anxious"; they might be "anxious," "insecure," and "under pressure" all at once. This requires a multi-label classification approach rather than traditional one-hot categorization.

Existing deep networks (like VGG or ResNet) often over-compress these sparse images or lose critical structural nuances in their deep pooling layers.

Methodology: Why "Shallow" is Better

The authors' core insight is that for abstract sketches, depth can be the enemy of detail. They proposed a Shallow CNN (SCNN) that prioritizes structural integrity over high-level abstraction.

1. Architectural Innovation

To combat feature loss, the team made several non-traditional design choices:

  • Large Kernels (7x7 and 5x5): Larger receptive fields are used early on to capture the global structure of the tree (roots, trunk, canopy) which carries the psychological weight.
  • Ditching the Pooling Layers: Standard max-pooling discards spatial information. In SCNN, pooling is disabled to keep every stroke's relative position intact.
  • High Dropout (0.5): Given the small dataset size (300+ samples), a heavy dropout is used to prevent the model from simply memorizing specific drawing styles.

Model Architecture Fig 1: The SCNN architecture featuring specialized convolution layers and a multi-label Sigmoid output.

2. Multi-Label Logic

Instead of using a Softmax layer (which assumes only one "winner" label), the model uses a Sigmoid activation in the final layer. This treats each of the 6 psychological categories as an independent probability, allowing the AI to flag multiple concerns for a single child.

Experiments & Results: Human-in-the-loop Validation

The dataset was sourced from a primary school in Hangzhou, with labels provided by professional psychoanalysts. The researchers categorized issues into: Anxiety, Pressure, Aggressiveness, Diffidence, Insecurity, and Lack of Learning Ability.

Performance Highlights:

  • Accuracy: The model hit a peak of 70.35% on the test set. While this might seem lower than SOTA on MNIST, psychological diagnosis is hit-or-miss even for humans; a 70% automated screening tool is a massive efficiency gain for clinics.
  • Ablation Study: The team tested various kernel sizes. Interestingly, combining a 7x7 first layer with a 5x5 second layer proved most effective, validating the need for a wide receptive field at the start.

Experimental Results Table 1: Hyper-parameter tuning showing the impact of kernel sizes and dropout on Test Accuracy.

Critical Analysis & Conclusion

The "Takeaway"

This research proves that AI doesn't always need "deeper" networks to solve "deeper" human problems. By tailoring the architecture to the specific constraints of the medium (sketches), the SCNN acts as a reliable filter to help psychologists prioritize high-risk cases.

Limitations & Future Work

The current model works on static images. However, the process of drawing (the order of strokes, the pressure applied to the pen) often contains more psychological data than the final result. The authors suggest that moving to sequence-based monitoring (using touch screens/tablets) will be the next frontier in AI-assisted psychoanalysis.

This work establishes a foundational baseline for using simple yet robust CNNs to interpret the complex, abstract world of child psychology.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize Deep Learning for the House-Tree-Person (HTP) psychological test or other projective drawing techniques.
  • Which paper first proposed the "Sketch-a-Net" architecture, and how does the current SCNN's approach to sparse feature extraction differ from it?
  • Explore how Graph Neural Networks (GNNs) or Vision Transformers (ViTs) are being applied to classify non-photorealistic hand-drawn sketches in clinical settings.
Contents
SCNN: Decoding the Silent Language of Children's Drawings via Shallow Learning
1. TL;DR
2. Problem & Motivation: The "Sparse Data" Challenge
3. Methodology: Why "Shallow" is Better
3.1. 1. Architectural Innovation
3.2. 2. Multi-Label Logic
4. Experiments & Results: Human-in-the-loop Validation
4.1. Performance Highlights:
5. Critical Analysis & Conclusion
5.1. The "Takeaway"
5.2. Limitations & Future Work