WisPaper
WisPaper
Search
Assistant
Pricing
TrueCite

Could tool-using language models reshape AI research over the next decade?

Tool-using language models show real promise in research, but evidence reveals a gap between best-case potential and current typical use.

Direct answer

Yes, tool-using language models are poised to reshape AI research over the next decade, but the transformation will be uneven and require careful governance. The strongest evidence here comes from a 2023 study that found GPT-3 could score 12,100 essays with accuracy comparable to human raters, suggesting that for structured, repetitive tasks, these models already deliver reliable results [2]. However, a 2024 survey of 226 researchers across 59 countries found that while 87.6% were aware of these tools, only 18.7% had actually used them in publications, and most of those uses were limited to grammar and formatting [1]. Across the studies here, the larger and more quantitative investigations consistently show that tool-using models excel at narrow, well-defined tasks—like toxicity prediction [4] or automated scoring [2]—but the survey data reveals a wide gap between what's possible and what's actually being adopted in practice.

5sources cited

This article was generated with WisPaper-powered search and paper analysis.

What can tool-using language models actually do in research?

The most concrete evidence comes from tasks that are repetitive, rule-governed, and data-rich. A 2023 study used GPT-3 to automatically score all 12,100 essays in the TOEFL11 corpus—a large, standardized dataset of non-native English writing—and found that the model's scores matched human benchmarks with 'a certain level of accuracy and reliability' [2]. This means that for grading essays, screening abstracts, or formatting citations, these models can already handle the workload at scale, freeing researchers for higher-level thinking. Similarly, a 2025 study compared GPT-4 and GPT-4o to traditional machine-learning models for predicting molecular toxicity and found that the language models were 'comparable' to deep-learning approaches in predicting bone, neuro, and reproductive toxicity, and could even be combined with molecular docking to identify cardiotoxic compounds in common herbs [4]. The key takeaway: when the task is well-defined and the data is structured, tool-using models perform at or near the level of specialized algorithms—without requiring the user to write code.

Why is there a gap between what's possible and what's actually happening?

Despite the technical promise, real-world adoption lags far behind. A 2024 survey of 226 medical and paramedical researchers from 59 countries—all trained in a Harvard Medical School program—found that while 87.6% were aware of large language models, only 18.7% had used them in a publication [1]. Even among those users, the majority (64.9%) applied them only for grammar and formatting, not for core analytical tasks [1]. This suggests that awareness is high but practical integration into research workflows is still shallow. The same survey found that 40.5% of users did not acknowledge the AI's use in their papers, pointing to ethical and transparency concerns that may slow adoption [1]. A 2023 editorial on using AI for management research echoed this, noting that the main disadvantages lie in the architecture of current general-purpose models—they can produce plausible-sounding but incorrect outputs, and researchers need to understand these limitations to use them safely [3]. So the bottleneck isn't capability; it's trust, training, and guidelines.

Will the reshape be uniform across all fields?

No—the evidence points to an uneven transformation. Fields with standardized, high-volume tasks (like essay scoring in language testing [2] or toxicity screening in drug development [4]) are likely to see rapid integration, because the models can directly replace or augment human effort. In contrast, fields that rely on nuanced interpretation, such as qualitative research or complex literature synthesis, face bigger hurdles. A 2023 study on using AI for systematic literature reviews found that while AI can speed up screening and data extraction, the 'most substantial disadvantages' come from the models' inability to truly understand context, meaning human oversight remains essential [3]. The 2024 survey of clinical researchers reinforces this: 52% believed LLMs would have a major impact on writing and editing, but only 32.6% were sure of the broader scope [1]. The reshape will be real, but it will happen fastest in the most routine, data-intensive corners of research.

About These Sources

This answer is built on 5 peer-reviewed studies — published from 2023 to 2025, 3 from 2024 or later, 2 in Q1 journals, collectively cited 724 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 57 papers retrieved from a database of over 500 million.

Sources used in this answer

1

Use of large language models as artificial intelligence tools in academic research and publishing among global clinical researchers.

In a cross-sectional survey of 226 medical researchers from 59 countries, 87.6% were aware of LLMs, but only 18.7% had used them in publications, mostly for grammar/formatting (64.9%); 40.5% did not acknowledge AI use, and 52% predicted major impact on writing/editing tasks.

2

Exploring the potential of using an AI language model for automated essay scoring

Using GPT-3 to automatically score all 12,100 essays in the TOEFL11 corpus, the study found accuracy and reliability comparable to human raters, with potential to enhance scoring via linguistic features.

3

On the use of AI-based tools like ChatGPT to support management research

This editorial on AI in management research provides guidelines for using AI in systematic literature reviews, noting advantages in objectivity/repeatability but major disadvantages in current models' architecture (e.g., lack of true understanding), requiring human oversight.

4

Large Language Models as Tools for Molecular Toxicity Prediction: AI Insights into Cardiotoxicity.

Comparing GPT-4 and GPT-4o with traditional ML models for molecular toxicity prediction, GPT-4 performed comparably on bone, neuro, and reproductive toxicity; combined with molecular docking, it identified cardiotoxic compounds in common herbs.

5

Generative AI and AI Tools in English Language Teaching and Learning: An Exploratory Research

Through semi-structured interviews with English language teachers, this exploratory study found that GenAI tools (ChatGPT, Gemini, Copilot) positively influenced teaching efficiency, student engagement, and writing skills, supporting personalized learning.