CGL-GAN: Efficiency Meets Precision in Text-to-Image Synthesis

Exploring Global and Local Linguistic Representations for Text-to-Image Synthesis

2020-02-11
Ruifan Li, Ning Wang, Fangxiang Feng, Guangwei Zhang, Xiaojie Wang
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces CGL-GAN, a cross-modal text-to-image synthesis framework that leverages both global and local linguistic representations within a single-pair Generative Adversarial Network. It achieves high-resolution image generation (256x256) comparable to multi-GAN baselines while significantly reducing model complexity and parameter count.

TL;DR

CGL-GAN breaks the trend of "bigger is better" in text-to-image synthesis by proving that a single, well-architected GAN can match the performance of complex multi-stage models. By introducing a Cross-modal Projection Block (CPB) that aligns word and sentence embeddings with corresponding layers of the visual hierarchy, the authors achieve high-resolution 256x256 image generation with up to 88% fewer parameters than contemporary SOTA models like Obj-GAN.


1. The Multi-GAN Dilemma

Generating a 256x256 image from a single sentence is a "one-to-many" mapping problem of extreme complexity. Early models like GAN-INT-CLS struggled with sparsity; a single vector representing a whole sentence isn't enough to tell a generator where to place a "red beak" or how to texture "wet feathers."

To fix this, the community moved toward Stacked GANs (e.g., StackGAN++, AttnGAN). These models generate images in stages: first a low-res sketch, then a refined 128x128 image, then finally 256x256. While effective, this creates a "parameter explosion" and makes end-to-end training notoriously difficult. CGL-GAN asks a critical question: Can we achieve this detail using only one generator and one discriminator?


2. Methodology: Structural Alignment

The "Secret Sauce" of CGL-GAN is the Cross-modal Projection Block (CPB) within the discriminator. Instead of just concatenating text and image features, the authors realized that vision and language have matching hierarchical structures.

The Hierarchical Intuition:

  • Local Alignment: Word-level hidden states (from a Bi-LSTM) contain fine-grained spatial cues. These are projected onto the mid-level feature maps () of the discriminator, which represent object parts.
  • Global Alignment: The entire sentence embedding represents the "gist" of the scene. This is mapped to the final high-level feature map (), which captures global semantics.

Architecture of CGL-GAN Figure 1: The CGL-GAN framework. Note how the discriminator splits the linguistic input into Global and Local streams to meet the image features at various depths.

By aligning granularity (word-to-part and sentence-to-object), the discriminator provides much more informative gradients to the generator, allowing it to learn complex details without needing a multi-stage pipeline.


3. Experimental Breakdown: Doing More with Less

The results displayed in the paper are a masterclass in efficiency.

Performance vs. Parameters

While models like Obj-GAN and MirrorGAN have pushed Inception Scores (IS) slightly higher on MS-COCO, they do so at the cost of nearly 200 million parameters. CGL-GAN holds its own with only 23.3 million.

MethodParameters (M)IS (COCO)
StackGAN107.88.45
AttnGAN169.425.89
CGL-GAN (Ours)23.313.62

Note: While AttnGAN’s IS is numerically higher, the researchers point out that IS is an exponential metric; the perceptual gap is often smaller than the numbers suggest.

The Evolution of Detail

Visual evidence shows that CGL-GAN significantly outperforms earlier single-GAN baselines. By the end of 104K iterations (shown below), the model evolves from "color blobs" to highly structured bird species with distinct background separation.

Training Iterations Figure 2: Evolution of the generative process. As training progresses, the local linguistic representations help "carve out" the specific features like the yellow breast or black wings.


4. Critical Insight: The CPB Variants

The authors conducted a brilliant ablation study (Table III in the paper) on four variants of the projection block:

  1. CGL-GAN-OG: Global info only.
  2. CGL-GAN-OL: Local info only.
  3. CGL-GAN-GL: Mismatched alignment (Local-to-High, Global-to-Low).
  4. CGL-GAN-LG: Natural Alignment (The Winner).

The "Natural Alignment" version (Local features to Low-level visual maps) resulted in a 65% reduction in FID compared to the "Local-only" version. This proves that high-resolution synthesis is not just about having the data, but about routing the data to the right level of the neural hierarchy.


5. Conclusion & Future Outlook

CGL-GAN is a reminder that architectural elegance can often substitute for raw computational power. By respecting the inherent hierarchy of language and vision, the authors stabilized a single-stage GAN for high-fidelity tasks.

Future Work: The authors suggest that this CGL-mode could itself be "stacked" in the future. Imagine a multi-stage GAN where every stage is as efficient as CGL-GAN—we could potentially see 1024x1024 synthesis with the footprint of a mobile-friendly model.

Takeaway for Practitioners: If your GAN is unstable, don't just add more layers; look at how you are feeding your conditional data. Are you asking a high-level layer to understand low-level details? Alignment might be the key.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize cross-modal alignment or contrastive learning (like CLIP) to stabilize single-stage text-to-image GANs without using multiple discriminators.
  • Identify the seminal paper that first introduced word-level local representations in conditional GANs for fine-grained image generation and how CGL-GAN's projection method differs.
  • Explore research that applies hierarchical linguistic-to-visual projection blocks to other generative tasks, such as text-to-video synthesis or 3D object generation.
Contents
CGL-GAN: Efficiency Meets Precision in Text-to-Image Synthesis
1. TL;DR
2. 1. The Multi-GAN Dilemma
3. 2. Methodology: Structural Alignment
3.1. The Hierarchical Intuition:
4. 3. Experimental Breakdown: Doing More with Less
4.1. Performance vs. Parameters
4.2. The Evolution of Detail
5. 4. Critical Insight: The CPB Variants
6. 5. Conclusion & Future Outlook