WisPaper
WisPaper
Search
Assistant
Pricing
TrueCite

Are the risks of multimodal AI agents being underestimated?

Evidence from 12 studies shows multimodal AI agents face steep accuracy drops, safety blind spots, and regulatory gaps that are often overlooked.

Direct answer

Yes, the risks of multimodal AI agents are being underestimated. Across the studies reviewed, these systems show dramatic performance drops when moved from static tests to real-world sequential tasks—diagnostic accuracy can fall to below a tenth of benchmark scores [1]. Safety compliance systems, while promising, still miss low-risk incidents (sensitivity as low as 0.03) [5], and the EU AI Act's risk classification framework may not fully capture the novel hazards of multimodal and agentic systems [7]. The evidence suggests that current optimism about these agents' reliability and safety outpaces what the data actually supports.

9sources cited

This article was generated with WisPaper-powered search and paper analysis.

How much does accuracy drop when multimodal agents face real-world complexity?

The biggest risk is that multimodal AI agents perform far worse in realistic, sequential tasks than in static benchmarks. In the AgentClinic study, which simulated clinical decision-making with patient interactions and tool use, diagnostic accuracy for the same MedQA problems dropped to below a tenth of the original score [1]. This means an agent that scores 90% on a multiple-choice test might succeed less than 9% of the time when it has to gather information, use tools, and make decisions step by step—a gap that is rarely communicated to users or regulators.

Even in industrial settings, performance is uneven. A construction safety framework achieved perfect helmet detection (F1 = 1.0) but only 0.75 for glasses and 0.03 sensitivity for low-risk incidents, meaning it almost never flagged minor hazards [5]. The same study noted that the system was biased toward severe risks, which could create a false sense of security about everyday dangers. Across the papers, the pattern is consistent: multimodal agents excel in narrow, well-defined tasks but stumble badly when the environment is messy, incomplete, or requires sustained reasoning.

What safety blind spots and biases do these agents have?

Multimodal AI agents inherit and amplify biases from their training data, and several studies found that these biases are not being adequately tested. The AgentClinic study explicitly perturbed agents with biases and found that clinical simulations were vulnerable to them [1]. A talk on trustworthy multimodal AI highlighted that domain-specific safety evaluations—especially in multilingual and distribution-shifted settings—uncover risks that generic red-teaming misses [6]. This suggests that standard safety testing is insufficient for the complex, real-world contexts where these agents are deployed.

The construction safety framework showed a clear bias toward severe risks, with low-risk sensitivity at just 0.03 [5]. In healthcare, a postpartum depression risk prediction system achieved an F1 score of 0.68 on structured data, but performance varied across modalities and the system was tested on a single hospital's data [8]. These blind spots are not just academic: in petrochemical safety, while multimodal AI models showed potential, the study itself noted that traditional methods react to hazards only after they occur, implying that AI systems might introduce new failure modes if not rigorously validated [4].

Are current regulations keeping up with the risks?

The evidence suggests that regulatory frameworks are lagging behind the capabilities of multimodal AI agents. One paper directly examines the EU AI Act and concludes that its risk classification approach—while a step forward—may not adequately address the unique challenges posed by large language and multimodal models [7]. The Act's categories (unacceptable, high, limited, minimal risk) were designed before agentic and multimodal systems became widespread, and the paper implies that these systems could fall through the cracks.

Meanwhile, the technical papers show that multimodal agents are being deployed in high-stakes domains like oil and gas inspection [3], counterfeit detection [2], and clinical decision support [1][8][9], often with claims of improved safety and efficiency. But the same papers also reveal that performance is highly variable and context-dependent. For example, a robotic inspection system saved over 1,000 hours annually in manual inspections [3], yet the same system's anomaly detection relies on cloud-based AI models that may not be transparent or auditable. Without regulatory requirements for continuous monitoring, explainability, and bias testing across diverse real-world conditions, the risks of these systems will likely remain underestimated.

About These Sources

This answer is built on 9 peer-reviewed studies — published from 2023 to 2026, 8 from 2024 or later, 2 in Q1 journals — selected as the most relevant from 12 studies that passed quality screening, drawn from 65 papers retrieved from a database of over 500 million.

Sources used in this answer

1

AgentClinic: a multimodal benchmark for tool-using clinical AI agents

In simulated clinical environments, diagnostic accuracy for the same MedQA problems dropped to below a tenth of the original score when agents had to use tools and make sequential decisions, revealing a massive gap between benchmark and real-world performance.

2

AuthentiCheck: A Multimodal AI Framework for Counterfeit Detection

A multimodal counterfeit detection framework that fuses computer vision, NLP, and metadata significantly outperformed single-channel alternatives, but the study did not test adversarial or real-world distribution shifts.

3

Advancing Operational Integrity Through Multimodal AI-Powered Robotic Inspection

A multimodal AI-powered robotic inspection system for oil and gas facilities saved over 1,000 hours annually in manual inspections and detected thermal, acoustic, and corrosion anomalies, but relies on cloud-based AI models that may lack transparency.

4

Transforming Petrochemical Safety Using a Multimodal AI Visual Analyzer

An evaluation of Gemini 1.5 Pro, GPT-4, and Copilot for petrochemical safety hazard detection found potential but noted that traditional methods react only after hazards occur, implying AI systems may introduce new failure modes.

5

A Multimodal AI Framework for Construction Safety: Compliance Detection and Risk Prediction

A construction safety framework achieved perfect helmet detection (F1=1.0) but only 0.03 sensitivity for low-risk incidents, showing a strong bias toward severe risks and limited ability to flag everyday hazards.

6

Towards Trustworthy Multimodal AI Systems.

A talk on trustworthy multimodal AI highlighted that domain-specific safety evaluations in multilingual and distribution-shifted settings uncover risks missed by generic red-teaming, and that existing explainability techniques are unreliable.

7

… and the implications of their use–are fostered by the EU AI Act, particularly from a risk classification standpoint, amid rapid advances in large language and multimodal …

An analysis of the EU AI Act suggests its risk classification approach may not adequately address the unique challenges posed by large language and multimodal AI systems, implying regulatory gaps.

8

ClinPreAI: An Agentic AI System for Early Postpartum Depression Risk Prediction from Multimodal EHR Data.

An agentic AI system for postpartum depression risk prediction achieved F1=0.68 on structured data, outperforming AutoML, but was tested on a single hospital's data and performance varied across modalities.

9

Multimodal radiology AI

A review of multimodal radiology AI argues that integrating imaging with EMRs and genetic data improves disease risk prediction, but notes major challenges from data heterogeneity and non-intuitive modalities.