Can AI safety benchmarks reduce the gap between expert optimism and public concern?

AI safety benchmarks can help but currently fall short of bridging expert-public divides due to gaps in risk coverage and measurement validity.

Direct answer

AI safety benchmarks can help reduce the gap between expert optimism and public concern, but they are not yet sufficient on their own. A 2024 survey of 18 AI risks found that US voters perceive these risks as both more likely and more impactful than experts do, and they favor slower AI development [3]. However, a review of over 200 safety evaluations reveals systemic gaps: few benchmarks cover non-text modalities, many ethical risks are simply not evaluated, and most tests ignore the real-world context where AI systems operate [2]. Across the studies here, the largest review of safety benchmarks [1] and the survey of public perceptions [3] both point to the same conclusion: current benchmarks are too narrow and disconnected from public concerns to fully bridge the gap.

4sources cited

This article was generated with WisPaper-powered search and paper analysis.

How big is the gap between experts and the public?

A 2024 survey of 18 specific AI risks—ranging from civilizational collapse to misinformation and systemic bias—directly measured the divide. The study surveyed AI experts and a representative sample of US registered voters, finding that voters consistently rated AI risks as both more likely to occur and more impactful if they did occur [3]. Voters also expressed a clear preference for slower AI development, a view not shared by experts. This gap is not minor: it means that even if experts believe a risk is well-managed, the public may feel it is imminent and severe, creating a trust deficit that benchmarks alone cannot fix.

Do current benchmarks address the risks the public worries about?

A separate 2026 review of 210 safety benchmarks reinforces this point, documenting that many benchmarks suffer from technical, epistemic, and sociotechnical shortcomings [1]. The authors argue that benchmarks often fail to follow established risk management principles, do not clearly map what can and cannot be measured, and lack robust probabilistic metrics. In other words, even when a benchmark exists, it may not measure what it claims to measure—or it may measure something trivial while missing the real-world danger.

Can better benchmarks actually reduce public concern?

Yes, but only if they are redesigned to address the specific gaps that fuel public worry. The 2024 survey found that policy interventions may best assuage collective concerns if they attempt to more carefully balance mitigation efforts across all classes of societal-scale risks, effectively nullifying the near-vs-long-term debate over AI risks [3]. This means benchmarks need to cover both immediate harms (like biased outputs) and longer-term existential risks—not just one or the other. The 2026 review offers a concrete roadmap: develop robust probabilistic metrics, apply measurement theory to connect benchmark objectives to real-world outcomes, and use a checklist to ensure benchmarks are epistemologically sound [1]. Early examples like the AI Safety Gridworlds (2022) show that even simple benchmark environments can reveal whether an AI agent is gaming rewards, ignoring side effects, or failing under distributional shift [4]—but these tests are still far from the comprehensive, context-aware evaluations the public expects.

About These Sources

This answer is built on 4 studies (1 peer-reviewed, 3 preprints) — published from 2022 to 2026, 3 from 2024 or later, collectively cited 132 times — selected as the most relevant from 4 studies that passed quality screening, drawn from 46 papers retrieved from a database of over 500 million.

Sources used in this answer

1

How should AI Safety Benchmarks Benchmark Safety?

A 2026 review of 210 safety benchmarks found that most suffer from technical, epistemic, and sociotechnical shortcomings, and recommends following risk management principles, mapping what can(not) be measured, and using robust probabilistic metrics to improve validity.

2

Gaps in the Safety Evaluation of Generative AI

A 2024 empirical review of over 200 safety evaluations for generative AI identified three systemic gaps: a modality gap (few tests for non-text), a risk coverage gap (many ethical risks unevaluated), and a context gap (most tests ignore real-world deployment settings).

3

Implications for Governance in Public Perceptions of Societal-scale AI Risks

A 2024 survey of AI experts and US registered voters across 18 AI risks found that voters perceive risks as more likely and impactful than experts do, and favor slower AI development; the study suggests policy interventions balancing mitigation across all risk classes could reduce public concern.

4

AI Safety Gridworlds

A 2022 study introduced a suite of reinforcement learning environments (AI Safety Gridworlds) that test safety properties like safe interruptibility and reward gaming; it found that two deep RL agents (A2C and Rainbow) could not solve these environments satisfactorily.