"Holy Starches Batman!!": Crowdsourcing the Invisible Art of Comics

"Holy Starches Batman!! We are Getting Walloped!": Crowdsourcing Comic Book Transcriptions

2016-10-20
Christine Samson, Casey Fiesler, Shaun K. Kane, Shaun K. Kane
Summary
Problem
Method
Results
Takeaways
Abstract

This paper explores the feasibility of using crowdsourcing via Amazon Mechanical Turk to generate accessible text transcriptions for comic books. The researchers conducted a pilot study with 60 participants to evaluate how instructional context (knowing the reader is blind) and fan-based domain knowledge influence the quality and detail of image descriptions.

TL;DR

Comics are a blind spot in digital accessibility due to their unique blend of visual and textual storytelling. This pilot study by Samson et al. demonstrates that crowdsourcing can bridge this gap, proving that simply telling a worker they are helping a blind person significantly increases the richness of the descriptions produced.

Contextualizing the Problem: Why Comics Aren't Just "Images"

Traditional accessibility focuses on "Alt-text" for static images or "Audio description" for linear video. Comics exist in a "third space"—the sequence and layout (the "gutter") are just as important as the dialogue.

Current limitations in the field include:

  • Tactile Failure: Braille overlays on comic pages often become too bulky or obscure the art for low-vision users.
  • AI Gap: While models like ImageNet are great at identifying "a cat," they struggle with "Batman being walloped by a starch-based weapon," which requires character recognition and narrative context.

Methodology: Testing the Human Element

The researchers recruited 60 Turkers and presented them with a two-page spread from Batman '66. They tested two primary variables:

  1. Instructional Context: Does knowing the end-user is blind change the worker's effort?
  2. Domain Expertise (Fandom): Does a Batman fan write a better description than a casual observer?

Comic Panel Used in Study Figure 1: The Batman '66 spread used to test worker descriptions.

Key Findings: Purpose Drives Detail

The study found a clear "empathy effect." When workers were aware of the accessibility goal, the mean word count jumped by nearly 37%. While word count is a proxy for detail rather than a direct metric of quality, it indicates a significantly higher cognitive investment.

The Fandom Factor

The data also suggested a "Fandom Trend." Participants who identified as comic readers produced descriptions that were, on average, 78 words longer than non-readers. This highlights the importance of Inductive Bias in human transcribers—fans already know who Bruce Wayne is, allowing them to focus on describing action rather than trying to figure out "who the guy in the cape is."

Critical Analysis & The Future of Fan-Sourcing

This paper serves as a vital "first step" but leaves open the question of Quality vs. Quantity. A long description isn't necessarily a good one if it's poorly structured.

The Strategic Shift: The authors suggest moving away from paid "micro-task" platforms like Mechanical Turk and toward Fan Communities. Platforms like Archive of Our Own (AO3) have already shown that fans are willing to perform immense amounts of "invisible labor" (tagging, transcribing, translating) for free out of a sense of community and social justice.

Conclusion

Samson et al. remind us that accessibility is not just a technical challenge—it's a social one. By framing transcription as a narrative act performed for a specific audience, we can unlock higher-quality data that could eventually be used to train specialized Vision-Language Models (VLMs) for the "invisible art" of comics.

Sample Action Sequence Figure 2: The complex interplay of action panels that requires human-level narrative synthesis.

Find Similar Papers

Try Our Examples

  • Find recent papers or SOTA methods that utilize "fan-sourcing" or community-driven altruism for large-scale accessibility transcription tasks.
  • Which study first formally defined the "complex interplay between words and images" in comics (e.g., McCloud 1993), and how have modern computer vision models attempted to encode this sequential logic?
  • Explore research that applies the "context-aware" instruction methodology used in this study to modern LLM-based image captioning for the visually impaired.
Contents
"Holy Starches Batman!!": Crowdsourcing the Invisible Art of Comics
1. TL;DR
2. Contextualizing the Problem: Why Comics Aren't Just "Images"
3. Methodology: Testing the Human Element
4. Key Findings: Purpose Drives Detail
4.1. The Fandom Factor
5. Critical Analysis & The Future of Fan-Sourcing
5.1. Conclusion