i
News
News · 2026-10-08

Companies pull back from letting AI agents ship production changes

@neuronium_ai @neuronium_ai

VentureBeat’s August survey found that fewer companies with autonomous AI agents are willing to let them ship changes to production without human approval: the share that already allows it, or is preparing to, fell from 75% in July to 56%. The retreat came as companies expanded their use of evaluation tools, including OpenAI’s, and reported incidents where systems passed internal checks but later caused customer problems. The results suggest a widening gap between automating evaluation and trusting it to make the final release decision.

Cover: Companies pull back from letting AI agents ship production changes

The retreat is not just a change in respondents

The two surveys drew different groups of VentureBeat readers and panel participants, so the overall shift may partly reflect who answered. But it also appeared among respondents who said they make final AI purchasing decisions: within that group, support for releasing changes without review, or preparing to do so, fell from 88% in July to 61% in August.

The August survey included 140 respondents, 53% of whom were decision-makers, compared with 44% in July. The shift therefore cannot be explained simply by a change in the overall respondent mix, though the results are not necessarily representative of the wider market.

The change also preceded Dario Amodei’s September 12 essay, “We Need to Slow Down on AI,” which called for slowing advanced model development and strengthening safety oversight. Sam Altman of OpenAI and Demis Hassabis of Google DeepMind later also argued for greater caution. Since the survey took place in August, it cannot show that those public statements caused companies to change their views.

Among all August respondents, 27% said they already allow some low-risk agents to release changes without human review, while 20% were preparing to do so within the next 12 months. Another 35% did not expect to allow it in the foreseeable future; 16% said their companies did not use autonomous agents, and 2% were unsure.

Among respondents whose companies use autonomous agents, 32% already allowed unreviewed releases of low-risk changes or updates, and 24% were preparing for them. But 42% expected human approval to remain in place for the foreseeable future, up from 20% in July. That increase was statistically significant.

Evaluation tools are growing; release authority is not

Companies’ greater caution about handing over release decisions did not come with a comparable retreat from evaluation infrastructure. The share of respondents whose companies used OpenAI’s built-in evaluation tools rose from 31% in July to 59% in August.

Asked which single reliability or evaluation area their company expected to increase investment in most over the next year, respondents chose:

Human review processes: 30%
Production monitoring tools: 26%
Automated evaluation pipelines: 21%
Safety and policy evaluation: 11%
No increase in reliability spending: 11%

These are respondents’ expectations about the fastest-growing area of investment, not a breakdown of actual budgets. Human review was the most common choice in both months, but its August lead over production monitoring was too small to be statistically significant. The results do not show that companies are cutting spending on automated evaluation.

A separate finding helps explain why the distinction matters. Among respondents whose companies conduct pre-release checks, 61% reported at least one customer-facing problem in the past 12 months involving an AI or large language model feature that had passed internal review. In July, the comparable figure was 53% among 101 respondents; VentureBeat had previously reported 49%, calculated across all 108 July participants.

The August figure was numerically higher, but the difference was not statistically significant. The share reporting exactly one such incident did rise significantly, from 27% in July to 39% in August. These figures count companies that encountered at least one incident, not failures per agent or release; a company deploying hundreds of AI features has more opportunities to encounter a problem than one deploying only a few.

Confidence in automated checks also remains limited, though the survey asked respondents to name only their single biggest concern. Just 9% chose “we trust automated evaluation today,” down from 13% in July and up from 5% in June. The July-to-August change was not statistically significant, and the other 91% should not be read as saying they have no trust at all.

The most frequently selected concern was a mismatch between evaluation results and real-world performance, at 27%. Respondents also cited a lack of explainability, at 24%; bias or inconsistent evaluation, at 16%; data leakage or privacy concerns, at 13%; and immature tools, at 12%.

I think the central distinction is not whether companies want automated evaluation. Their tool use suggests they do. The question is how much authority they are willing to attach to its results. In August, respondents who reported a post-review customer incident were about as likely to be moving toward autonomous releases as those who did not: 59% versus 55%. The difference was too small to establish a meaningful relationship.

Production monitoring may miss wrong answers

Pre-release checks are only one layer of oversight. Among 118 August respondents whose companies use autonomous agents, just 29% said real-time answer-quality checks were their primary form of production monitoring.

The most common method was transaction-trace logging, named by 36%. Another 16% primarily tracked infrastructure through an API gateway; 11% did not know or said monitoring was outside their area, and 8% relied on irregular checks.

36%transaction logs
29%real-time checks
16%API gateway

Trace logging records infrastructure activity, token counts, and inputs and outputs for later debugging, but does not necessarily assess whether an agent’s answer is correct. A confidently delivered but incorrect answer could leave a clean trace and an HTTP 200 status: the request succeeded technically, even if the answer was wrong.

In total, 53% named transaction logging or API-gateway tracking as their primary monitoring method. Both can help identify infrastructure problems, but neither directly checks answer quality as the survey described them.

That gap is visible even among companies already allowing autonomous releases. Of 38 respondents at those companies, only 10 named real-time answer-quality checks as their primary monitoring method—about 26%, close to the 28% among 40 comparable respondents in July.

My guess is that the survey captures two decisions moving at different speeds: companies are building evaluation capacity while remaining reluctant to make evaluation the sole gatekeeper for production changes. It does not establish that other companies lack additional quality controls; the monitoring question allowed only one answer. Nor does it show how problems are usually detected, whether through automated alerts, employees, or customer complaints.

OpenAI gains a bigger place in the stack

OpenAI’s evaluation platform appeared in 59% of technology stacks in August, up from 31% in July, a statistically significant increase. Among the August respondents who answered the platform question, 40% named OpenAI as their primary evaluation platform.

Those figures measure different things. The 59% figure is the share using OpenAI tools; the 40% figure is the share of respondents choosing OpenAI as their main platform. Among companies already using OpenAI’s built-in evaluation tools, about 69% named them as their primary platform.

Other reported usage and primary-platform shares were:

Confident AI: present in 36% of stacks; primary platform for 10%
Braintrust: present in 23% of stacks; primary platform for 12%
Anthropic: present in 16% of stacks; primary platform for 10%
LangSmith: present in 13% of stacks; primary platform for 2%

The August comparison used the same base of 136 respondents for usage and primary-platform shares. Respondents could select several tools they used, but only one primary platform, so usage shares are not mutually exclusive. Confident AI’s usage rose from 27% in July to 36% in August, but that change was not statistically significant. The survey did not compare primary-platform shares across the two months because July respondents could name a primary platform they had not marked as in use.

More than three-fifths of August respondents—62%—planned to implement, add, or replace an evaluation platform within 12 months. More than a third, 35%, planned to do so within three months. Those plans do not necessarily signal dissatisfaction or an intention to leave an existing vendor: some companies may add a tool rather than replace one.

The survey points to a tension, not a verdict on which tools work. Companies are adding evaluation platforms while holding back on unreviewed releases, and many still monitor whether systems function rather than whether their answers are right. As agents gain permission to change more consequential business systems, the hard decision is not whether to automate evaluation, but when its result is reliable enough to replace a human’s final approval.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X