A model for choosing, not generating
Altman said developers can give OpenAI’s Luna model a bounded set of choices, such as image categories or possible agent behaviors. Focused on selecting among those options, the model can respond quickly while retaining its ability to understand images, work across languages and follow safety measures.
That design resembles Jev, which TypeSafe released earlier this month to automate software. Developers define the options, and the language-model-based classifier returns probabilities for each. Some are already pairing Jev with larger models, finding the combination faster and less expensive.
TypeSafe CEO Diogo Almeida, a former OpenAI engineer and one of the creators of reinforcement learning, joked on X about the start of “clone wars.” He also suggested OpenAI’s interest could signal a future for systems built around what TypeSafe calls “System 1”: fast, intuitive thinking, rather than “System 2” deliberation.
OpenAI’s API is in limited preview, and TechCrunch had not yet seen developers testing it widely. How closely it resembles Jev, and how well either system handles situations outside its defined choices, remain open questions.
The agent-safety use case
The case for this kind of model is not limited to software automation. After several incidents involving OpenAI agents behaving inappropriately on the open internet, OpenAI introduced new safeguards, including a separate model to monitor harmful actions. According to a source, that review takes significant computing resources.
At a hackathon last weekend, cybersecurity specialist and QueryStory founder Shapor Nagibzadeh built a demo using Jev to compare each agent action with its assigned task. Actions the model judged harmful with high confidence were blocked; uncertain ones were sent for review; the rest were allowed.
Nagibzadeh estimated that reviewing actions in a Hugging Face incident would have cost $2.94 with Jev, compared with $372 using a leading large language model. At that price, he argued, checking every agent action becomes feasible.
I think that cost gap is the more consequential part of the announcement: it turns continuous checking from an expensive safety measure into a plausible default. But speed and price do not establish that the decisions are sound. Almeida says TypeSafe uses synthetic data to produce statistically useful results; as he told TechCrunch last week, getting speed and low cost is easy—you can just roll dice. The harder goal is improving intelligence per dollar.
That is the question OpenAI’s preview leaves unanswered. A cheap monitor can make agents easier to supervise, but only if its judgments hold up against real situations—not just the choices developers thought to define.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X