Code that passes without being understood
Jordan Nanos, a technical staff member at independent research firm SemiAnalysis, described the OpenAI example in an interview earlier this year. He said the engineers he reviewed the code with understood the hardware and the system’s principles, but could not explain what the DeepSeek MLA kernel did, even line by line.
The kernel supports work on using available hardware more efficiently to train and run AI models. Nanos said AI understood and tested the code, producing kernels that were both correct and performant. In his account, developers did not need to reason deeply about every line: AI could manage how data moved through the hardware and how its compute elements were used.
That is a plausible change in how engineering work gets done. It is also a change in what code review is supposed to guarantee.
Undo’s analysis points to how often developers encounter the gap:
The last two figures suggest the cost is not limited to the risk of a hidden defect. Engineers may spend much of the workday inspecting code whose immediate output looks right, but whose logic they cannot confidently maintain.
The cost of oversight
Amazon’s 90-day “code safety reset,” launched earlier this year after a series of failures disrupted order processing, shows why that distinction matters. Amazon disputed claims that AI caused the failures. But notes from internal meetings indicate that executives were concerned about the “large blast radius” of changes generated with generative AI.
The concern is straightforward: if human review weakens, errors and hallucinations may go unnoticed. The more code teams accept without understanding, the harder it becomes to judge whether a working result is safe to change or dependable in a different context.
Sergey Kleftsov, an AI researcher, offered a more optimistic interpretation in a recent Medium post. He called the shift a change in the software-development paradigm: developers can focus on architecture and correctness criteria, while AI handles many implementation details. Nanos has also argued that deep human reasoning about the code itself is not always necessary.
I think that argument is strongest when the system’s boundaries and correctness criteria are clear. The figures from Undo point to a less settled reality: engineers are not only delegating implementation, but repeatedly spending time reconstructing what the tools produced. That is not proof the model of development has failed, but it does make productivity harder to measure by output alone.
The question I would want answered is how teams know when generated code is safe to leave opaque. If the next generation of engineers has fewer reasons to learn how software works beneath the interface, the industry may gain speed while losing the people best equipped to diagnose its failures.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X