i
DATAIST
News · 2026-09-15

Salesforce and Nvidia built Koa to cut Agentforce's token bill

@neuronium_ai @neuronium_ai

Salesforce and Nvidia have built Koa, a reasoning model trained for sales and customer support, and it will sit inside Agentforce next to the models Salesforce already pays for. Until now, when an Agentforce agent hit a long or multi-step reasoning task, Salesforce's AI gateway routed the prompt out to a frontier model — Claude or ChatGPT. Koa is the in-house answer to that, and the pitch is not that it is smarter. It is that it should finish the same task on fewer tokens.

Cover: Salesforce and Nvidia built Koa to cut Agentforce's token bill

Salesforce and Nvidia have built Koa, a reasoning model trained for sales and customer support, and it will sit inside Agentforce next to the models Salesforce already pays for. Until now, when an Agentforce agent hit a long or multi-step reasoning task, Salesforce's AI gateway routed the prompt out to a frontier model — Claude or ChatGPT. Koa is the in-house answer to that, and the pitch is not that it is smarter. It is that it should finish the same task on fewer tokens.

Agentforce is the platform where Salesforce customers assemble AI agents for routine work: answering customer questions, booking meetings. Jayesh Govindarajan, executive vice president of Salesforce AI, told TechCrunch the company had already built a stack of small language models for individual tasks inside that set. Reasoning was the gap. For anything requiring it, Salesforce had to go to the frontier labs.

The sales case for Koa runs to six points. It is an open alternative to the closed frontier models. It is trained for specific work tasks rather than for solving mathematics problems nobody can solve. No real customer data went into the training set, so it cannot surface one company's material to another. It uses fewer tokens on the same job, and therefore costs less to run. The AI gateway routes to it automatically depending on what a request needs. And it complies with customers' data requirements and the security rules already built into Salesforce.

Govindarajan's account of why the base model is Nvidia's Nemotron is the sharpest thing in the announcement. Salesforce had wanted to train its own enterprise-grade frontier model for some time and could not find a foundation to start from. Before Nemotron, he said, there was no available sovereign American pretrained model that was simultaneously usable, at the current state of the art, and clear about where its training data came from. His example of unclear provenance was Qwen, Alibaba's popular open model from China.

That is a three-way filter, and it is worth being exact about what it screens out. Chinese open weights fail on provenance, in Govindarajan's telling. Closed American frontier models fail on availability — you cannot post-train what you are not allowed to hold. What survives is the short list of American open-weight models near the state of the art, and Govindarajan named none of them, except the one he was rejecting. Nvidia gets something from that framing it does not get from selling chips: a claim on the base layer itself.

The privacy argument comes from how the post-training was done. Salesforce and Nvidia used no real Salesforce customer data. They generated synthetic material instead — a simulated support environment with a virtual specialist fielding calls from irritated customers, and scenarios with a sales rep trying to close a deal. A general-purpose model came out of that specialised in the two jobs Salesforce sells software for. Koa is meant to suit those customers' work better than Claude or ChatGPT and to spend fewer tokens doing it.

Kari Briski, Nvidia's vice president for enterprise generative AI software, told TechCrunch that Nemotron uses an architecture designed for economical inference, and named three things that matter for a system like this: sovereign AI, a short time to first token, and efficient reasoning, which is what determines the cost of the tokens.

Here is what I think is actually going on. The model is not the asset — the gateway is. Agentforce's router decides which model handles which request, and Salesforce owns the router. Until Koa, that router pointed at Anthropic and OpenAI for anything that required reasoning, and every enterprise reasoning task passing through it was revenue for someone else. Koa does not have to beat Claude to change that. It has to be adequate on the narrow band of work Agentforce actually runs — a support question, a scheduling request — so the router can be tuned to try it first. Salesforce writes the model, writes the router, and grades the outcome. The labs keep whatever is too hard for Koa, which is the slowest-growing part of the workload.

Salesforce is not walking away from either lab. It announced a partnership with Anthropic called ClaudeForce, letting companies use Claude as the interface to AI while data stays in Salesforce's system of record, protected by Salesforce's own infrastructure. Read the two announcements together and the shape is clear: Claude is welcome as the face, Salesforce keeps the records, and the reasoning that used to be metered by someone else moves in-house wherever it can. Naming a partnership after a partner's product while shipping a model aimed at the work that partner was doing is a particular kind of confidence.

What is absent from all of this is arithmetic. Koa should use fewer tokens than Claude or ChatGPT on the same request — should, in the future tense, with no benchmark attached, no measured ratio, no price. Suiting a customer's work better is not a number either. Enterprise buyers are being asked to take an efficiency claim on the word of the company that saves the money and the chip vendor that supplied the base, with the company that saves the money also controlling the router that decides when the claim gets tested.

The problem for the labs is not Koa in particular. It is that Salesforce has published the recipe: take an American open-weight base, post-train it on synthetic reconstructions of the work your customers actually do, and stop paying by the token for the routine tasks that never needed a frontier model at all. Salesforce has the distribution, it knows which tasks its customers run, and it now has a base model that arrives without another lab's terms attached. Every enterprise software company with an agent platform and a frontier model bill is looking at the same three ingredients. The labs' best customers are the ones with the most reason to leave.