What can tool-using language models actually do in research?
The most concrete evidence comes from tasks that are repetitive, rule-governed, and data-rich. A 2023 study used GPT-3 to automatically score all 12,100 essays in the TOEFL11 corpus—a large, standardized dataset of non-native English writing—and found that the model's scores matched human benchmarks with 'a certain level of accuracy and reliability' [2]. This means that for grading essays, screening abstracts, or formatting citations, these models can already handle the workload at scale, freeing researchers for higher-level thinking. Similarly, a 2025 study compared GPT-4 and GPT-4o to traditional machine-learning models for predicting molecular toxicity and found that the language models were 'comparable' to deep-learning approaches in predicting bone, neuro, and reproductive toxicity, and could even be combined with molecular docking to identify cardiotoxic compounds in common herbs [4]. The key takeaway: when the task is well-defined and the data is structured, tool-using models perform at or near the level of specialized algorithms—without requiring the user to write code.
Why is there a gap between what's possible and what's actually happening?
Despite the technical promise, real-world adoption lags far behind. A 2024 survey of 226 medical and paramedical researchers from 59 countries—all trained in a Harvard Medical School program—found that while 87.6% were aware of large language models, only 18.7% had used them in a publication [1]. Even among those users, the majority (64.9%) applied them only for grammar and formatting, not for core analytical tasks [1]. This suggests that awareness is high but practical integration into research workflows is still shallow. The same survey found that 40.5% of users did not acknowledge the AI's use in their papers, pointing to ethical and transparency concerns that may slow adoption [1]. A 2023 editorial on using AI for management research echoed this, noting that the main disadvantages lie in the architecture of current general-purpose models—they can produce plausible-sounding but incorrect outputs, and researchers need to understand these limitations to use them safely [3]. So the bottleneck isn't capability; it's trust, training, and guidelines.
Will the reshape be uniform across all fields?
No—the evidence points to an uneven transformation. Fields with standardized, high-volume tasks (like essay scoring in language testing [2] or toxicity screening in drug development [4]) are likely to see rapid integration, because the models can directly replace or augment human effort. In contrast, fields that rely on nuanced interpretation, such as qualitative research or complex literature synthesis, face bigger hurdles. A 2023 study on using AI for systematic literature reviews found that while AI can speed up screening and data extraction, the 'most substantial disadvantages' come from the models' inability to truly understand context, meaning human oversight remains essential [3]. The 2024 survey of clinical researchers reinforces this: 52% believed LLMs would have a major impact on writing and editing, but only 32.6% were sure of the broader scope [1]. The reshape will be real, but it will happen fastest in the most routine, data-intensive corners of research.
About These Sources
This answer is built on 5 peer-reviewed studies — published from 2023 to 2025, 3 from 2024 or later, 2 in Q1 journals, collectively cited 724 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 57 papers retrieved from a database of over 500 million.
Sources used in this answer
Use of large language models as artificial intelligence tools in academic research and publishing among global clinical researchers.
In a cross-sectional survey of 226 medical researchers from 59 countries, 87.6% were aware of LLMs, but only 18.7% had used them in publications, mostly for grammar/formatting (64.9%); 40.5% did not acknowledge AI use, and 52% predicted major impact on writing/editing tasks.
Exploring the potential of using an AI language model for automated essay scoring
Using GPT-3 to automatically score all 12,100 essays in the TOEFL11 corpus, the study found accuracy and reliability comparable to human raters, with potential to enhance scoring via linguistic features.
On the use of AI-based tools like ChatGPT to support management research
This editorial on AI in management research provides guidelines for using AI in systematic literature reviews, noting advantages in objectivity/repeatability but major disadvantages in current models' architecture (e.g., lack of true understanding), requiring human oversight.
Large Language Models as Tools for Molecular Toxicity Prediction: AI Insights into Cardiotoxicity.
Comparing GPT-4 and GPT-4o with traditional ML models for molecular toxicity prediction, GPT-4 performed comparably on bone, neuro, and reproductive toxicity; combined with molecular docking, it identified cardiotoxic compounds in common herbs.
Generative AI and AI Tools in English Language Teaching and Learning: An Exploratory Research
Through semi-structured interviews with English language teachers, this exploratory study found that GenAI tools (ChatGPT, Gemini, Copilot) positively influenced teaching efficiency, student engagement, and writing skills, supporting personalized learning.
