i
News
News · 2026-09-27

Why 95% of AI pilots miss the results companies expect

@neuronium_ai @neuronium_ai

A widely cited estimate says 95% of AI pilots fail to deliver the expected return or results, at least by the measures available. At The Next Endeavor in Mountain View, executives from GAI Insights, Jellyfish, Kloudfuse and Nasik discussed what it takes to turn a pilot into something a company can run in production. Their answers pointed less to a shortage of AI tools than to the work around them: changing roles and processes, measuring results and deciding who is responsible when systems fail.

Cover: Why 95% of AI pilots miss the results companies expect

The work around the models

The panelists described different parts of that problem:

Andrew Lau, co-founder of Jellyfish, said companies often fall short of the productivity gains they expect because AI adoption is moving faster than their ability to reorganize work. Teams, roles and processes all have to adapt.
Pankaj Thakkar, CEO and co-founder of Kloudfuse, said its observability platform covers metrics, logs and traces across infrastructure, applications and now AI agents, with customer data staying in the customer’s environment.
Rahul Todkar, founder and CEO of Nasik, said building agents can take seconds or minutes; operating them in production, especially at large companies, is much harder. Nasik has built an open runtime that brings agents, coding tools, frameworks and other tools into one system.
Paul Baier, co-founder of GAI Insights, said the company holds frequent briefings and discussions with boards and executives because the subject remains confusing and changes quickly.

Todkar also drew a distinction between the pursuit of AGI and what he called IGA: “intelligence that is governed and administered.” Agents without access controls, governance and oversight, he said, can become a serious source of danger.

The pilot is not the whole system

The panelists argued that AI projects need to be judged in context. Lau suggested treating software development separately from broader pilots: instead of looking only at return on investment, companies can compare token costs with results such as faster engineering work. The right measure depends on who uses a tool and how it fits a particular function, whether product development, finance or human resources.

That makes a single ROI figure a poor answer to a more basic question: what changed in the business? Todkar said companies need to discuss more often how they measure results. Lau said practices are still shifting roughly every three months, though patterns are beginning to emerge across industries. Repeated use cases can make productivity gains more visible, but complex deployments are often difficult to assess from the outside.

Thakkar tied that measurement problem to observability, a longstanding discipline for tracking system performance and understanding why a system behaves as it does. In production, companies also need to know who owns a system and who responds when something goes wrong. Todkar warned that production architectures differ sharply from isolated test environments; debugging takes time, and opaque layers can overlap.

What companies still have to prove

I think the panel’s most useful point is also the least glamorous: a successful AI pilot is not just a model performing well in a test. It is a system that can be integrated into real work, monitored and assigned to someone who is accountable for it. The estimate that 95% of pilots miss their expected results sets a stark backdrop, but the discussion offered no common benchmark for success—and no evidence that better observability or governance alone will improve that figure.

The same caution applies to open models. Lau said companies need criteria for comparing models, and that open models are already used in production partly because they can be assessed against those criteria. Thakkar sees them as a logical choice because they are easier to adapt for specific tasks, but said the industry is not yet fully ready for that shift. The unresolved question is not simply which model to choose, but whether companies can measure its value and operate it safely once it leaves the pilot.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X