What would make users trust process-level evaluation for long-horizon AI R&D agents in AI research and development agents?

Trust in long-horizon AI R&D agents requires process-level evaluation that is auditable, integrity-based, and focused on trajectory drift, not just final scores.

Direct answer

Users would trust process-level evaluation for long-horizon AI R&D agents if it goes beyond final scores to track the agent's evolving trajectory against the user's authorized task, using auditable, deterministic checks rather than opaque end-to-end judgments. Evidence shows that such prefix-level monitoring can detect drift with over 93% F1 while keeping benign coverage above 95.8% [2], and that integrity-based explanations—especially about data biases—foster appropriate trust more often than generic transparency [4]. Across the studies, trust builds when evaluation is transparent about process, not just outcomes [1][5].

5sources cited

This article was generated with WisPaper-powered search and paper analysis.

Why final scores aren't enough for long-horizon agents

For long-horizon AI R&D agents, a single final score hides where progress is gained or lost and whether the agent actually followed the user's intent. A 2026 systematic evaluation of seven frontier models on 36 long-horizon tasks found that current agents behave more like engineering optimizers than autonomous researchers: their performance varies substantially across runs, and their strongest solutions mostly adapt or combine established techniques rather than produce genuine novelty [5]. This means a high final score could come from a lucky run or a clever hack, not from reliable, trustworthy behavior.

The same study showed that similar final outcomes can arise from different process bottlenecks, and that experience reuse can either help or mislead subsequent decisions [5]. So evaluating only the end result gives users no way to know whether the agent is drifting toward an unauthorized objective or making decisions that will degrade later. Trust requires visibility into the process itself.

What makes process evaluation trustworthy: auditable trajectory checks

The key to trust is not just checking each step in isolation but monitoring the entire trajectory against the user's authorized task. A 2026 paper introduces 'ontological trust,' a property of trajectory prefixes that decomposes trust into Role, Goal, and Evidence (RGE). This monitor uses LLMs only to derive structured representations; the trust-state updates and intervention decisions are deterministic, making the output a replayable and auditable trust trajectory rather than a single judge verdict [2]. In tests across OSWorld, FinanceBench, and EICU-AC, RGE outperformed rule-, judge-, and shield-style baselines on detecting prefix-paired drift, exceeding 93% Drift F1 on every benchmark while keeping benign coverage at or above 95.8% [2].

What this means for users: you can see exactly why the agent's trust level changed at each step, and you can replay the evaluation to verify it. This is a stark contrast to opaque end-to-end judges that give a single verdict without explanation. The deterministic nature of the trust updates is crucial—it makes the evaluation reproducible and auditable, which is a foundation for trust.

Integrity-based explanations build appropriate trust

Process evaluation must also communicate the agent's integrity—its honesty, transparency, and fairness—in ways that help users calibrate their trust. A 2023 user study with 160 participants found that an AI agent that explicitly disclosed potential biases in its data or algorithms achieved appropriate trust more often than agents that were merely honest about capabilities or transparent about decision-making processes [4]. Interestingly, honesty-like explanations were better for trust recovery after a mistake [4].

This suggests that process evaluation should not just report what the agent did, but also surface its limitations and biases. Users trust agents that are upfront about their weaknesses, not just their successes. This aligns with the finding from a 2022 medical AI study where 91% of clinicians reported positive expectations and satisfaction when the system provided explanations and functionalities that helped them understand and interact with the AI [1].

User perception and adoption depend on transparent process signals

Trust is not just about technical accuracy; it's about how users perceive and adopt the agent. A 2025 study of 632 participants in China found that trust links users' dual decision paths (heuristic and systematic) to drive adoption behavior [3]. This means that users rely on both quick impressions and deep analysis, and process evaluation must cater to both.

For long-horizon R&D agents, this implies that process evaluation should provide both high-level signals (e.g., 'the agent is on track') and detailed, inspectable logs (e.g., 'here's the evidence for each step'). The 2026 evaluation framework [5] also emphasizes the need for rule-based metrics that characterize within-run behavior, which can help users understand not just what the agent did, but how it did it. When users can see the process and understand the agent's reasoning, they are more likely to trust and adopt it.

About These Sources

This answer is built on 5 studies (3 peer-reviewed, 2 preprints) — published from 2022 to 2026, 3 from 2024 or later, 3 in Q1 journals, collectively cited 151 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 51 papers retrieved from a database of over 500 million.

Sources used in this answer

1

BreastScreening-AI: Evaluating medical intelligent agents for human-AI interactions

In a real-world study with 45 clinicians across nine institutions, adding AI assistance reduced false positives by 27% and false negatives by 4%, and 91% of clinicians reported positive expectations and satisfaction, suggesting that transparent explanations and functionalities are key to trust.

2

Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents

Introduced an online monitor (RGE) that decomposes trust into Role, Goal, and Evidence, using deterministic updates for auditability; it exceeded 93% Drift F1 on prefix-paired drift detection across OSWorld, FinanceBench, and EICU-AC while keeping benign coverage at or above 95.8%.

3

How Do Consumers Trust and Accept AI Agents? An Extended Theoretical Framework and Empirical Evidence

Based on a survey of 632 participants in China, trust mediates the link between heuristic and systematic decision paths and user adoption of AI agents, indicating that both quick cues and detailed information are needed for trust.

4

Integrity-based Explanations for Fostering Appropriate Trust in AI Agents

In a between-subject user study with 160 participants, an AI agent that explicitly disclosed potential data biases achieved appropriate trust more often than agents that were honest about capabilities or transparent about decision-making, and honesty-like explanations aided trust recovery.

5

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

A systematic evaluation of seven frontier models on 36 long-horizon tasks found that final scores hide process bottlenecks and experience reuse issues; agents varied substantially across runs and rarely produced genuine novelty, highlighting the need for process-level metrics.