The work around the models
The panelists described different parts of that problem:
Todkar also drew a distinction between the pursuit of AGI and what he called IGA: “intelligence that is governed and administered.” Agents without access controls, governance and oversight, he said, can become a serious source of danger.
The pilot is not the whole system
The panelists argued that AI projects need to be judged in context. Lau suggested treating software development separately from broader pilots: instead of looking only at return on investment, companies can compare token costs with results such as faster engineering work. The right measure depends on who uses a tool and how it fits a particular function, whether product development, finance or human resources.
That makes a single ROI figure a poor answer to a more basic question: what changed in the business? Todkar said companies need to discuss more often how they measure results. Lau said practices are still shifting roughly every three months, though patterns are beginning to emerge across industries. Repeated use cases can make productivity gains more visible, but complex deployments are often difficult to assess from the outside.
Thakkar tied that measurement problem to observability, a longstanding discipline for tracking system performance and understanding why a system behaves as it does. In production, companies also need to know who owns a system and who responds when something goes wrong. Todkar warned that production architectures differ sharply from isolated test environments; debugging takes time, and opaque layers can overlap.
What companies still have to prove
I think the panel’s most useful point is also the least glamorous: a successful AI pilot is not just a model performing well in a test. It is a system that can be integrated into real work, monitored and assigned to someone who is accountable for it. The estimate that 95% of pilots miss their expected results sets a stark backdrop, but the discussion offered no common benchmark for success—and no evidence that better observability or governance alone will improve that figure.
The same caution applies to open models. Lau said companies need criteria for comparing models, and that open models are already used in production partly because they can be assessed against those criteria. Thakkar sees them as a logical choice because they are easier to adapt for specific tasks, but said the industry is not yet fully ready for that shift. The unresolved question is not simply which model to choose, but whether companies can measure its value and operate it safely once it leaves the pilot.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X