Vibe Researching as Wolf Coming: The Dawn of the AI Agent in Social Science

Vibe Researching as Wolf Coming: Can AI Agents with Skills Replace or Augment Social Scientists?

Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces "vibe researching," an AI-driven paradigm for social science where autonomous agents like "scholar-skill" (a 21-skill plugin for Claude Code) execute entire research pipelines from hypothesis generation to manuscript submission. It proposes a cognitive task framework to define the boundary between human judgment and AI automation based on codifiability and tacit knowledge.

TL;DR

The "Wolf" of autonomous AI research has arrived. This paper introduces "vibe researching"—a shift where AI agents move beyond simple chatbots to execute entire research pipelines. Using the scholar-skill system as a case study, the author argues that while AI can replicate the structure of top-tier research at 10x speed, the soul of social science—theoretical originality and tacit field knowledge—remains an irreducibly human domain.

Problem & Motivation: Beyond the Chatbot

For decades, social scientists used computers as "fancy calculators" (Wave 1) or "efficient scrapers" (Wave 2). Even the recent use of LLMs for text classification (Wave 3) kept the human firmly in the driver's seat for reasoning.

We are now entering Wave 4: Agentic AI. Unlike a chatbot that answers a prompt, an AI agent holds state, uses tools, and executes multi-step workflows. The motivation for this paper is to address the existential dread appearing in academia: if an AI can write the code, run the lit review, and simulate the reviewers, what is left for the PhD?

Methodology: The Scholar-Skill System

To ground this theory, the author presents scholar-skill, a plugin for Claude Code that organizes 21 specialist skills.

Table of AI Skills

The system doesn't just "talk"; it acts:

  • Literature Synthesis: Accesses a local Zotero library of 20,000+ items to find gaps.
  • Causal Identification: Constructs Directed Acyclic Graphs (DAGs) and writes R/Stata code for complex estimators like "staggered Difference-in-Differences."
  • Reviewer Simulation: Spawns multiple AI "agents" to act as harsh peer reviewers, flagging major and minor concerns.

The Cognitive Task Framework

The core contribution of this paper is the Codifiability vs. Tacit Knowledge matrix.

Cognitive Task Framework

The author argues the boundary for delegation isn't a "stage" (e.g., "I'll let AI do the data cleaning, and I'll do the writing"). Instead, it's cognitive:

  1. Type A (Formulation): High tacit knowledge (field politics, intuition). KEEP HUMAN.
  2. Type B & D (Planning/Communication): The "dangerous middle zone." AI drafts; humans must judge.
  3. Type C (Execution): High codifiability. DELEGATE TO AI.

Experiments & Results: Speed vs. Comprehension

The results from the scholar-skill system highlight a massive productivity premium:

  • Speed: Theoretical themes extracted in 10 minutes from thousands of PDFs.
  • Scaffolding: Researchers can now use advanced methods (e.g., Causal Forests) that they might not have the coding proficiency to implement from scratch.

However, more is not always better. The paper identifies a "Jagged Frontier": AI-generated theory often looks "well-structured and plausible" but is ultimately a recombination of existing ideas. It lacks the "metaknowledge"—the understanding of what the scientific community currently finds "live" or "settled"—which is only gained via years of human interaction at conferences and in the field.

Critical Analysis & Conclusion

The Pedagogical Crisis

The most striking takeaway is the warning for graduate education. If we only teach students "how to run regressions" (execution), we are training them in a skill with zero market value in the age of AI. Instead, we must teach them Evaluative Expertise: the ability to judge if the AI's output is actually right.

The 5 Principles for Responsible "Vibe Researching"

  1. Disclose: Normalize reporting AI use in "Methods" sections.
  2. Verify: If errors are published in your name, they are your errors.
  3. Maintain Skills: You cannot judge what you cannot do yourself.
  4. Protect Originality: The research question must remain human-driven.
  5. Design for Access: Ensure AI productivity gains don't just benefit elite, English-speaking institutions.

Final Thought: The "Wolf" is already in the room. Our task is not to hide, but to use these agents to amplify our reach while fiercely protecting the "theoretical imagination" that makes social science human.

Find Similar Papers

Try Our Examples

  • Search for recent studies or benchmarks evaluating the "jagged technological frontier" regarding LLM performance in specialized academic reasoning tasks.
  • Which papers first established the concept of "metaknowledge" or "tacit knowledge" in scientific communities, and how do they inform the limits of AI-generated theories?
  • Explore how AI agent frameworks like AutoGen or Open-Ended Scientific Discovery (The AI Scientist) are being adapted specifically for qualitative versus quantitative research workflows.
Contents
Vibe Researching as Wolf Coming: The Dawn of the AI Agent in Social Science
1. TL;DR
2. Problem & Motivation: Beyond the Chatbot
3. Methodology: The Scholar-Skill System
3.1. The Cognitive Task Framework
4. Experiments & Results: Speed vs. Comprehension
5. Critical Analysis & Conclusion
5.1. The Pedagogical Crisis
5.2. The 5 Principles for Responsible "Vibe Researching"