Where do tool-using language models already work well in science?
In specialized, data-rich domains, tool-using language models are already performing at a level that is practically useful. For instance, a deep learning framework called CCLMoff, which incorporates a pretrained RNA language model, was trained on a comprehensive dataset to predict off-target effects of CRISPR/Cas9 gene editing. It demonstrated strong generalization across diverse next-generation sequencing datasets, meaning it could accurately identify unintended edits even for sequences it had never seen before [1]. This is a concrete, practical tool for designing safer gene therapies.
Similarly, in emergency ophthalmology, large language models (LLMs) like GPT-4 and Llama-3-70b were tested against certified ophthalmologists on 73 real-world emergency cases. Human experts scored an average of 3.72 out of 4, while GPT-4 scored 3.52 and Llama-3-70b scored 3.48. The difference between human and GPT-4 or Llama-3-70b was not statistically significant, meaning these LLMs performed comparably to human doctors in diagnosing and planning treatment for eye emergencies [2]. This suggests that for narrow, well-defined tasks with clear decision criteria, tool-using models can be reliable decision-support tools.
What are the main limitations keeping them from broader scientific use?
Despite successes in specific tasks, experts broadly agree that current tool-using language models have significant limitations that prevent them from being trusted as general scientific tools. In a 2025 opinion piece, four groups of scientists debated the role of LLMs in science. One group argued that LLMs are often misused and overhyped, and that their limitations—such as a tendency to generate plausible-sounding but incorrect information—warrant a focus on more specialized, interpretable tools [3]. Another group emphasized that humans must retain responsibility for setting the scientific roadmap, not the models [3].
Performance is also inconsistent across models and tasks. In the same ophthalmology study, GPT-4o scored significantly lower (3.20) than human experts (3.72), showing that even within the same family of models, newer versions do not always improve [2]. A 2023 study on ChatGPT in economics found it could generate useful code for economic models, but the authors also highlighted limitations and recommended careful use [4]. A 2026 review of agentic tool use in LLMs noted that existing studies are fragmented across tasks and tool types, lacking a unified view of how to make these models reliably effective [5]. This fragmentation means that a model that works well for one scientific task may fail unpredictably on another.
The central trade-off: specialized tools versus general-purpose assistants
The evidence points to a clear trade-off: tool-using language models are most useful when they are specialized for a particular scientific task, but they are often marketed or imagined as general-purpose scientific assistants. The CRISPR off-target predictor [1] and the emergency ophthalmology decision-support tool [2] both succeeded because they were applied to narrow, well-defined problems with high-quality training data and clear evaluation criteria. In contrast, the opinion piece [3] and the review of agentic tool use [5] both caution that general-purpose use of LLMs in science is risky because the models lack true understanding and can produce confident errors.
This trade-off is not a failure of the technology but a guide for how to use it. The strongest evidence here suggests that tool-using language models are already practical for specific scientific subfields—like CRISPR design or emergency diagnosis—where the cost of a mistake is high but the model's performance is validated against human experts. For broader scientific tasks, such as generating hypotheses or designing experiments, the models are not yet reliable enough to be used without close human oversight [3][5]. The path forward, as suggested by the experts, is to develop more specialized, interpretable tools rather than expecting a single model to do everything [3].
About These Sources
This answer is built on 5 peer-reviewed studies — published from 2023 to 2026, 4 from 2024 or later, 2 in Q1 journals, collectively cited 59 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 48 papers retrieved from a database of over 500 million.
Sources used in this answer
A versatile CRISPR/Cas9 system off-target prediction tool using language model
CCLMoff, a deep learning framework using a pretrained RNA language model, achieved accurate CRISPR/Cas9 off-target prediction with strong generalization across diverse NGS-based detection datasets, laying the foundation for an end-to-end sgRNA design platform.
Using large language models as decision support tools in emergency ophthalmology.
In a prospective comparative study of 73 real-world emergency ophthalmology cases, GPT-4 (mean score 3.52) and Llama-3-70b (3.48) performed comparably to human experts (3.72) on a 4-point scale, with no statistically significant difference, while GPT-4o (3.20) performed significantly worse.
How should the advancement of large language models affect the practice of science?
In a 2025 opinion piece, four groups of scientists debated LLM use in science: one group argued LLMs are overhyped and limited, another emphasized transparent attribution, and two groups argued that humans must retain responsibility for the scientific roadmap.
Large Fourth-Generation Language Models as a New Tool in Scientific Research
A 2023 study demonstrated ChatGPT's ability to generate C# programs for agent-based and computable general equilibrium economic models, highlighting its potential for interdisciplinary research while also analyzing its limitations and providing recommendations for use.
Agentic Tool Use in Large Language Models
A 2026 review organized agentic tool use in LLMs into three paradigms (prompting, supervised learning, reward-driven learning), finding that existing studies remain fragmented across tasks and tool types, lacking a unified view of how tool-use methods differ and evolve.
